Daemon Corporation

Global Leader in Excellence.

Research Report № 005

Capability Mismatch: How Ideas, Knowledge, Technology, and Institutions Co-evolve

Central finding

Conceptual, epistemic, technological, and institutional capabilities often develop together, but not as a synchronized system and not according to a general law. An advance in one domain can expose a constraint in another: faster transportation can overwhelm existing coordination practices; identification systems can reveal taxpayers without creating the ability to collect from them; accurate computation can optimize the wrong institutional objective; and technically capable AI can increase rather than reduce work when review, context, and accountability remain unresolved.

These are genuine capability mismatches only under demanding conditions. Uneven development by itself is not enough. A consequential mismatch exists when:

  1. a particular outcome requires capabilities from more than one domain;
  2. one identifiable capability is inadequate relative to that task;
  3. failure or delay occurs at the interface between otherwise usable capabilities;
  4. adding the missing complement improves both the interface and the outcome; and
  5. competing explanations—scarcity, incentives, power, execution, or external shock—do not explain the same pattern as well.

The best evidence comes from interventions that isolate complements. In Liberia, neither taxpayer identification nor information about penalties increased property-tax payment on its own, but their combination raised payment from 2.2 to 9.7 percent. A separate experiment showed that signaling actual enforcement capacity produced an additional effect.1 In Indian welfare programs, broadly similar biometric technology improved payment delivery when combined with local intermediaries, cash infrastructure, manual overrides, and gradual implementation in Andhra Pradesh, but increased exclusion and transaction costs when imposed rapidly with weak fallbacks in Jharkhand.2 These cases demonstrate more than the coexistence of several capabilities: they identify particular complements and show that outcomes changed when those complements were supplied.

But the framework has sharp limits. Large asymmetries can remain harmless when an adequate substitute, slack, or bounded task absorbs them. Antebellum railroads operated for years without integrating available telegraphy because timetables and time-interval spacing remained adequate at prevailing traffic levels.3 Complete-looking capability bundles can also persist without producing an expected transformation: double-entry bookkeeping spread centuries before industrial capitalism without apparently serving the profit-calculation function later theories attributed to it.4 And many apparent mismatches are better understood as conflicts over rents, authority, legitimacy, or distribution. Containerization, for example, required real technical complements, but regulatory protection and labor conflict—not an absent capability—were major barriers to adoption.5

The evidence therefore supports capability mismatch as a conditional causal diagnosis, not as a general theory of historical change. Its most important discipline is simple:

Specify the outcome, specify the timescale, locate the failing interface, and do not treat technical resolution as equivalent to institutional recognition or social improvement.


1. From uneven development to consequential mismatch

The four-domain model is useful because each domain can, at least provisionally, impose a different constraint:

  • Conceptual capability determines how a problem is represented: the distinctions, categories, objectives, models, and vocabularies available for understanding it.
  • Epistemic capability determines what can be observed or established: measurement, evidence, scientific knowledge, and methods of validation.
  • Technological capability determines what can materially be done: tools, infrastructure, computation, communications, and physical techniques.
  • Institutional capability determines what can be coordinated and sustained: authority, organizational form, governance, incentives, rules, norms, and legal arrangements.

Co-evolution occurs because a change in one domain alters what is required from the others. A new technology may make operations faster while increasing the burden on managerial coordination. A measurement system may render people or transactions legible while creating new demands for appeals, enforcement, privacy protection, and exception handling. An institutional reform may increase access to finance without immediately reducing its price. A new conceptual category may make a problem visible but still lack evidence, tools, or authority sufficient to act on it.

This suggests a bottleneck model. For a particular task, performance may be limited less by the total stock of capability than by the least adequate necessary complement. If technological capacity is abundant but institutional authorization is absent, more technology may have little value. If an agency can identify obligations but cannot credibly enforce them, better records may not generate compliance. If an AI system can draft code but lacks tacit repository context, reliable evaluation, or cheap human review, faster generation may merely move work downstream.

That model produces testable implications:

  • Investment in an already abundant domain should have low returns while the bottleneck persists.
  • The value of a proposed complement should be greatest where its absence is most severe.
  • Supplying two interdependent capabilities together should yield more than the independent contribution of either one.
  • Process evidence should show queues, handoff failures, overrides, delays, or errors concentrated at the predicted interface.
  • Once the bottleneck is relaxed, the old interface problem should diminish—even if a new constraint subsequently appears.

These implications resemble the logic of organizational complementarity developed in research on technology, workplace design, skills, and incentives.6 Yet most historical evidence cannot establish interaction effects this cleanly. The diagnostic model should therefore be treated as a proposed instrument, not a validated general measure.

Distinguishing mismatch from its rivals

Coordination overload, workaround proliferation, “glue work,” institutional brittleness, and failure to exploit a technology are possible symptoms of mismatch. None is diagnostic by itself.

Candidate explanation Evidence that would favor it
Capability mismatch Failure clusters at a specific cross-domain interface; supplying the predicted complement produces disproportionate gains; additional investment in nonbinding domains has little effect.
Resource scarcity Broad increases in money, staff, time, equipment, or infrastructure improve performance without requiring a distinctive interaction among capabilities.
Incentive conflict Behavior changes when compensation, sanctions, monitoring, or ownership changes, even though technical and epistemic capacities remain constant.
Power or distributive struggle Obstruction is selective, tracks threatened rents or authority, and persists despite demonstrated feasibility; bargaining or coercive settlement matters more than further capability development.
Poor execution Failures are local or idiosyncratic, decline with routine learning, and do not recur systematically at the same interface across competent implementations.
External shock Outcomes track exposure to war, crisis, price movements, or other external timing rather than the presence or absence of a proposed complement.
Ordinary failure No stable bottleneck, interaction, or repeatable mechanism can be identified; the outcome is compatible with noise, bad judgment, or multiple unrelated errors.

The crucial test is not whether a system exhibits stress, but whether a specific missing complement predicts where that stress appears and whether adding it changes the result.


2. What the historical cases actually show

Railroads: organizational invention before technological adoption

The nineteenth-century railroad illustrates two different co-evolutionary processes that are easily collapsed into a single story.

At the New York and Erie Railroad, Daniel McCallum diagnosed the difficulty of managing a large, geographically distributed enterprise as a failure of organizational system. His 1856 account emphasized divided responsibility, verification, personal accountability, and a structure adapted to the scale of a long railroad. His organizational chart made reporting relationships visible; the associated managerial system helped establish a template later used in large industrial enterprises.7

This is a case of conceptual and institutional innovation applied to an existing technological system. The relevant advance was not a new communications tool. It was a representation of organizational structure combined with explicit authority, reporting, and accountability. The evidence establishes a contemporaneous diagnosis and an influential response, although it does not provide quantitative operating data sufficient to measure the exact effect on accidents, delays, or costs.

Telegraph adoption followed a different path. Antebellum managers often saw little reason to integrate telegraphy because regimented timetables and time-interval train spacing provided acceptable safety and efficiency at prevailing traffic levels. The technology was available, but the existing substitute was good enough. Only after the telegraph demonstrated its value during the Civil War did railroads and telegraph companies form durable operating relationships.3

The contrast matters:

  • The Erie coordination problem was addressed through organizational design without requiring a new technological complement.
  • Telegraph availability did not automatically create a mismatch or compel adoption.
  • A demonstration under wartime pressure changed beliefs and legitimacy more than it changed the underlying technology.
  • McCallum himself connects the episodes: after his organizational work at Erie, he later led the U.S. Military Railroads during the war that demonstrated telegraphy’s value.

This is not a simple sequence from technological advance to institutional adaptation. It is evidence that organizations may solve an acute problem through one domain while leaving another available capability unused until scale, pressure, or perceived reliability changes.

Electrification: real complementarity embedded in a multi-causal lag

Factory electrification is often presented as the canonical example of a general-purpose technology awaiting organizational redesign. Electric motors accounted for less than 5 percent of U.S. factory mechanical-drive capacity in 1899 and only slightly more than half by the early 1920s—roughly four decades after the first central power station opened.8

Part of the delay did reflect complementarity. Replacing a central steam engine with electric motors yielded limited value if factories retained layouts designed around shafts, belts, and centralized power. The larger gains required unit drive, new floor plans, redesigned production flow, and personnel able to build and operate the new system.

But Paul David’s account is explicitly multi-causal. It also identifies deficiencies in conventional productivity measurement, the slow accumulation of experienced factory architects and electrical engineers, sunk-cost behavior, and the timing of regulated electricity-price reductions after 1914–17.8 The lag therefore cannot be reduced to one missing organizational capability. It combined learning, prices, regulation, measurement, capital replacement, and design.

Electrification supports a modest claim: technologies may deliver their largest gains only when complementary skills and organizational forms change with them. It does not establish that every delayed return to a technology is a mismatch, nor that one institutional complement was the sole binding constraint.

Longitude: solving a technical problem did not settle institutional recognition

The British Board of Longitude funded competing approaches to determining longitude at sea, especially lunar-distance calculations and marine chronometers. These should not be described as jointly necessary capabilities: mariners could use them together or as alternatives. Parallel funding was an institutional hedge under uncertainty over which approach would succeed.9

John Harrison’s H4 chronometer met the Board’s accuracy criterion. Yet the formal prize was never awarded to Harrison—or to anyone. Technical resolution and institutional recognition diverged.

A specific governance mechanism helps explain the divergence. Nevil Maskelyne, an advocate of the competing lunar-distance method, became Astronomer Royal and sat on the Board judging Harrison’s device. He returned a negative report contesting the watch’s performance. Harrison eventually appealed outside the normal adjudicative process: after intervention by George III, Parliament granted him £8,750 in 1773. This was a separate parliamentary payment, not the balance of the formal prize.10

The case shows why “the capability arrived” is an incomplete causal statement. The chronometer’s technical performance did not automatically produce recognition, reward, or adoption. Adjudication was shaped by institutional authority and conflict of interest. No additional improvement in timekeeping could, by itself, repair that governance defect.

Credible commitment: one reform, two outcomes, two timescales

England’s post-1688 financial development supplies the clearest warning against treating an institutional effect as a single outcome.

North and Weingast argued that constitutional changes made government commitments to creditors more credible. Government borrowing increased by more than an order of magnitude between 1688 and 1697, which they interpreted as evidence of a rapid change in lenders’ willingness to supply funds. They also argued that long-term interest rates fell from approximately 14 percent in the early 1690s to 6–8 percent by the end of that decade and continued downward thereafter.11

Sussman and Yafeh found a different pattern for the price of credit: interest rates remained high and volatile for four decades after the Glorious Revolution and responded strongly to wars and domestic instability. Their analysis challenges the claim that institutional reform immediately reduced the cost of capital, but it does not rebut the documented increase in borrowing volume.12

The apparent contradiction dissolves once the outcome is disaggregated:

  • Quantity of credit: government borrowing capacity expanded quickly.
  • Price of credit: borrowing costs remained elevated and volatile for decades.

An institutional change could therefore relax one constraint—the amount lenders were willing to provide—without immediately overcoming war risk or other determinants of price. Whether the case confirms “institutional capability” depends on which result is specified and over what period.

Containerization: complements existed, but conflict was not a capability deficit

The standardized shipping container was more than a box. Effective containerization required specialized ships, mechanized terminals, compatible handling systems, work-rule changes, and dimensional standards. The creation of ISO Technical Committee 104 in 1961 and publication of ISO 668 in 1968 were genuine coordination advances.5

Yet technological incompleteness does not explain the entire delay. A 1931 Interstate Commerce Commission ruling required containers to be charged according to the most expensive item inside, suppressing their use on railroads. Longshore unions resisted a system that threatened employment. These were distributive and regulatory conflicts, not failures to understand, measure, or engineer container transport. Labor accommodation required bargaining over the gains from mechanization, including the 1960 Pacific Coast Mechanization and Modernization Agreement.5

Later military estimates suggested that earlier adoption during the Vietnam War could have produced very large savings. That estimate demonstrates perceived economic value; it does not establish that Vietnam itself broke the political deadlock.

Containerization thus contains two simultaneous processes:

  1. a real bundle of technological and coordination complements had to be constructed; and
  2. actors with threatened rents could obstruct the use of a technically workable system.

Calling the entire delay a capability mismatch would erase the second process. Calling it entirely political would erase the first.

Bookkeeping: capabilities can exist without triggering transformation

The history of double-entry bookkeeping challenges the idea that assembling a capable representation necessarily induces institutional change. Werner Sombart treated double-entry accounting as a necessary basis for rational capitalist enterprise. Basil Yamey’s examination of merchant records, as characterized in the available secondary literature, found that bookkeeping was generally used for narrower record-keeping and balance-striking purposes rather than systematic calculation of profit or capital.4

The technique diffused across European commerce for centuries before industrial capitalism. Its presence did not produce the transformation attributed to it. Later interpretations relocate its importance to other mechanisms, such as constituting money and credit in bank ledgers. That possibility may be valid, but it also illustrates a risk: a flexible co-evolution framework can be rescued from disconfirmation by changing the causal channel after the expected outcome fails to occur.

Because Yamey’s original 1949 article was not directly available for this assessment, this counterexample rests on secondary characterization and should be given less weight than the experimental cases. Its methodological warning remains important: a complete-looking capability bundle is not evidence that the expected transformation will follow.


3. The strongest contemporary evidence: interfaces and implementation bundles

Legibility is not enforcement: taxation in Liberia

The Liberian tax experiments provide the clearest direct evidence of complementarity.

The first experiment independently varied taxpayer-identification information and information about legal penalties. Neither intervention alone significantly raised property-tax payment. Together, they increased payment from 2.2 to 9.7 percent, with the combined effect statistically distinguishable from either single treatment.1

A separate experiment tested collection capacity rather than merely penalty information. Previously paying but delinquent owners were informed that their properties would be included in a forthcoming enforcement batch. Payment increased by 7 percentage points from a 21 percent control baseline—a 33 percent relative increase.1

These results distinguish three capacities often collapsed into “state capacity”:

  • identifying who owes tax;
  • communicating liability and legal consequences;
  • credibly collecting or enforcing payment.

The first experiment shows an interaction between identification and penalty information. The second shows that credible enforcement is a separate institutional capability. Better data alone did not guarantee compliance; neither did abstract legal authority.

Identity is not delivery: welfare biometrics in India

Biometric identification produced very different outcomes under different implementation regimes.

In Andhra Pradesh, a randomized rollout covering 157 subdistricts and approximately 19 million people combined biometric smartcards with local payment intermediaries, physical cash-delivery infrastructure, manual overrides, and gradual implementation. Beneficiaries received payments more quickly, employment-program earnings increased without higher government expenditure, and estimated leakage declined.2

In Jharkhand, rapid mandatory biometric authentication with weak fallback arrangements produced no average improvement in beneficiary receipts or leakage, raised transaction costs, and increased the probability that beneficiaries received nothing—especially where credentials had not yet been issued. A later inventory-reconciliation rule reduced leakage partly by reducing legitimate benefit receipt as well.2

The comparison is not a controlled replication: the states, programs, and timing differed. It nevertheless shows that biometric identification is neither sufficient nor uniformly beneficial. Identity verification has to be integrated with enrollment, payment delivery, staffing, exception handling, fallback rights, and transition rules.

The case also reveals a normative limit. Reducing leakage is not automatically institutional progress if legitimate beneficiaries are excluded. Capability analysis can identify whether an arrangement functions according to a target; it cannot decide whether the target, distribution, or procedure is just.

Accuracy is not validity: the health-risk algorithm

A widely used health-management algorithm predicted future medical spending as a proxy for health need. At the same risk score, Black patients were substantially sicker than White patients because unequal access to care meant that spending did not reliably measure illness across groups. Replacing the proxy would have increased the proportion of Black patients selected for additional assistance from 17.7 to 46.5 percent.13

The system was not simply inaccurate at its encoded task. It predicted cost while the institution cared about need. More computation or more observations of the same biased spending process would not necessarily solve the problem.

This failure crosses the conceptual–epistemic boundary:

  • deciding that “need” should be represented by expected spending is conceptual;
  • testing whether spending validly measures need is epistemic;
  • implementing the model is technological;
  • deciding who receives care-management resources is institutional.

The domains are analytically distinguishable here, but not ontologically separate. The important mismatch lies in their interface.

Better participation without a new institution: Brazilian electronic voting

Brazil’s electronic-voting reform provides a counterexample to the claim that multiple capabilities must arrive together to create institutional change. A machine-availability threshold in the 1998 election produced a sharp comparison between jurisdictions. Electronic voting increased valid votes, especially in places with higher illiteracy, without a corresponding discontinuity in turnout or registration. It was associated with changes in representation, health spending, and infant health among less-educated populations.14

The relevant institution—representative elections—already existed. Technology reduced the difficulty of casting a valid ballot. This was an important improvement, but it is better classified as making participation easier within an existing institution than as creating a new institutional form.


4. When do capability bundles make institutions newly feasible?

The evidence supports a weaker and more useful formulation than the claim that institutions become possible only when all four domains “arrive together.”

Several complementary capabilities can create a feasibility discontinuity: an arrangement that was conceivable or occasionally achievable becomes reliable, affordable, fast, or scalable enough for routine use. This is different from logical novelty.

Digital government illustrates the distinction. Shared identity, interoperable registries, common standards, legal authorization, funding, and data-exchange infrastructure support low-latency “once-only” public services in which citizens need not repeatedly provide the same information to separate agencies. Estonia’s X-tee infrastructure is a prominent example.15 Governments coordinated records before such systems through paper, telephone, and bilateral databases. What is new is the practical scale, speed, and repeatability—not inter-agency coordination itself.

The stronger claim is therefore best expressed conditionally:

New institutional arrangements become practically feasible when a bundle of capabilities lowers previously binding costs of observation, communication, verification, coordination, or enforcement below the thresholds required for reliable operation.

That proposition is supported in particular settings:

  • Liberia’s tax authority needed legibility, penalty information, and credible collection capacity.
  • Andhra Pradesh’s payment reform needed identity verification plus delivery infrastructure and exception handling.
  • Interoperable digital services require identity, data standards, infrastructure, authority, and funding.
  • Machine-scale oversight of digital agents requires automated monitoring, logging, escalation rules, and human review capacity.

But three qualifications are essential.

First, “newly practical” does not mean “previously inconceivable.” Second, supplying a capability bundle does not guarantee adoption if power, rents, or legitimacy intervene. Third, feasibility does not imply desirability. An integrated administrative system may be efficient while undermining privacy, due process, distributional fairness, or democratic control.


5. AI and autonomous systems: new agency, cheaper assistance, and unresolved authority

AI is a useful current test because technical capabilities are changing faster than surrounding arrangements can be evaluated. The evidence supports three categories.

Genuinely newly practical

Frontier AI systems can now complete some well-specified, tool-mediated digital tasks over multiple steps. One benchmark estimated a roughly 110-minute task horizon at 50 percent success, although horizons at higher reliability were much shorter and performance worsened on less structured tasks.16 The important novelty is not text generation alone, but bounded digital agency: systems can take sequences of actions through software tools without each action being individually specified by a human.

Automated monitoring of agent behavior is also newly practical at machine scale. OpenAI reports monitoring tens of millions of internal coding-agent interactions over five months, with less than 0.1 percent outside coverage and approximately 1,000 cases escalated for asynchronous human review.17 This is an existence proof that review need not scale one-for-one with agent actions. It is not proof that monitoring prevents harm: the figures are first-party and unaudited, and the process described is largely corrective rather than preventive.

These capabilities make plausible new forms of bounded delegation—for example, giving agents authority to perform reversible, logged, low-stakes digital operations subject to automated checks and human escalation. That is a design possibility, not evidence that such arrangements are already superior.

Cheaper or easier, not institutionally novel

In a study of 5,179 customer-support agents, generative-AI assistance increased productivity by about 14 percent on average, with larger benefits for less-experienced workers.18 The institution and workflow remained recognizable: humans handled customer service inside an existing managerial process, with AI providing guidance.

A professional-services experiment found that AI users completed more tasks, more quickly and at higher quality, when assignments fell inside the system’s demonstrated capability frontier. On tasks outside that frontier, users were less likely to reach correct answers.19 AI altered the cost and quality of work, but did not remove the need for human task selection and judgment about when assistance was reliable.

These cases resemble electronic voting more than the creation of a novel institution. Technology changes the interface and production function; it does not by itself create a new form of authority or coordination.

Still blocked

A randomized study of experienced open-source developers working in familiar repositories found that access to early-2025 AI tools made them 19 percent slower, even though they expected and perceived a speed improvement.20 Review burden, missing tacit context, and low acceptance of generated patches moved the bottleneck from production to evaluation and integration.

For high-stakes autonomous authority, the remaining barriers are broader:

  • reliability over long or messy tasks;
  • correlated and propagated failures in multi-agent systems;
  • auditability and evidence sufficient to reconstruct decisions;
  • responsibility and liability when an autonomous system causes harm;
  • legitimate delegation of public, fiduciary, or managerial authority;
  • incentives to disclose failure rather than conceal it;
  • control over tools, resources, and permissions;
  • protection against concentrating power in system owners or operators.21

These are not merely engineering problems, although technical advances could reduce some of them. They are also questions about who may act, who bears risk, whose objectives are encoded, who can appeal, and who has authority to stop the system.

What AI does not yet show

The AI evidence is too recent to establish a completed co-evolutionary cycle. It shows capability advances, early integration experiments, new review burdens, and governance gaps. It does not yet show durable stabilization followed by a new mismatch, nor does it demonstrate that autonomous governance of firms or public institutions is effective or legitimate.


6. The recurring cycle: useful pattern, unreliable law

The proposed sequence—

advance → mismatch → stress → experimentation → complementary development → stabilization → new mismatch

—describes some cases well.

It fits the Erie Railroad at the level of organizational crisis and redesign. It fits Liberia, where interventions isolated missing complements. It fits Andhra Pradesh, where biometric technology became useful as part of a delivery system rather than as a stand-alone device. Electrification loosely follows the pattern, with eventual redesign and skill accumulation, although prices, measurement, and sunk capital complicate the story.

Elsewhere the cycle breaks:

  • No stress: Telegraph non-adoption persisted because railroad timetables remained adequate.
  • No transformation: Double-entry bookkeeping existed for centuries without producing the institutional change attributed to it.
  • Conflict rather than mismatch: Containerization was obstructed partly by regulation and distributional struggle.
  • Technical success without institutional recognition: Harrison’s chronometer met the technical criterion, but governance blocked the expected reward.
  • Different stabilization timescales: England’s borrowing volume expanded quickly while its borrowing costs remained high and volatile.
  • Interface improvement without new institutional form: Electronic voting improved participation in an existing electoral institution.

The sequence is most plausible when actors can identify and control the missing complement. McCallum could redesign reporting and authority; a tax agency could change notices and enforcement signals; a payment program could add intermediaries and fallback procedures.

It performs poorly when the obstacle is a conflict over rents, authority, or legitimacy. In those cases, adaptation may require bargaining, coalition change, external intervention, or coercion rather than another capability advance. Power can suppress a workable alternative; incumbents can preserve an inefficient arrangement; and a durable workaround can remove enough pressure that deeper adaptation never occurs.

The present evidence does not include a clear case of complete institutional collapse or enduring lock-in that never yielded to later adaptation. Claims about those branches of the cycle therefore remain hypotheses rather than findings.


7. What the four-domain model explains—and what it obscures

Explanatory value

The model’s strongest contribution is analytical decomposition. It helps separate:

  • technical performance from institutional recognition;
  • identifying an obligation from enforcing it;
  • predicting a proxy from validating the underlying objective;
  • increasing the quantity of a resource from lowering its price;
  • improving an institutional interface from creating a new institution.

Those distinctions materially change the interpretation of the cases. They prevent “technology worked” from standing in for “the institution adapted,” and prevent “the outcome improved” from standing in for “the reform was fair or desirable.”

Diagnostic value

The framework can be diagnostically useful if applied before the outcome is known. Analysts must define:

  1. the task and desired outcome;
  2. the capabilities it requires;
  3. the observable interface at which they interact;
  4. the proposed bottleneck;
  5. the rival explanations;
  6. the change expected if the bottleneck is relaxed.

Without these commitments, “mismatch” risks becoming a retrospective label for any disappointing result. Workarounds, delays, and governance failures can always be redescribed as evidence that some capability was missing. That flexibility makes the framework difficult to falsify.

Predictive value

Predictive support is presently weak. No study in the evidence base operationalizes all four capabilities prospectively and demonstrates that the model predicts outcomes better than scarcity, incentives, power, execution quality, or shocks. Liberia validates a particular interaction, not the entire ontology. The Indian welfare comparison strongly supports implementation complementarity, but is not an identical matched replication across states.

The framework can generate useful predictions, but it has not yet been validated as a general forecasting instrument.

The domains do not partition reality cleanly

Several important capabilities straddle domains:

  • Digital identity is technological infrastructure and an epistemic claim about who someone is.
  • Algorithmic target selection combines a conceptual definition with epistemic validation.
  • Software permissions and workflow gates are simultaneously technical code and institutional rules.
  • Measurement systems are often inseparable from the authority that determines what is measured and how classifications affect people.
  • Enforcement is an institutional arrangement but may depend on coercive resources that deserve separate treatment.

The four domains should therefore be treated as diagnostic lenses, not natural kinds.

Important missing or cross-cutting factors

The cases repeatedly identify capabilities not comfortably contained in the model:

  • Human skill and tacit knowledge: decisive in electrification, workplace computing, and AI-assisted programming.
  • Attention and review capacity: a major constraint when automation increases the volume of outputs requiring validation.
  • Coercive and enforcement capacity: distinct from administrative legibility or formal legal authority.
  • Legitimacy and trust: inputs to adoption, possible products of reliable performance, and possible casualties of exclusion.
  • Material resources: money, personnel, infrastructure, and physical access may constrain every domain.
  • Power and political will: determine whether technically workable alternatives are adopted, blocked, or imposed.
  • Normative objectives: rights, fairness, distribution, privacy, and contestability determine whether increased capability constitutes progress.

Research on state capacity likewise finds administrative, extractive, and coercive dimensions to be interrelated rather than cleanly separable.22

The available evidence does not permit a substantive comparison with all the broader theories named in the question—cumulative cultural evolution, technological paradigms, sociotechnical transitions, the adjacent possible, or related accounts of institutional evolution. Complementary-assets reasoning is directly supported in several cases, and path dependence is descriptively plausible where infrastructure, regulation, and vested interests preserve existing arrangements. Stronger claims about how the four-domain model maps onto those literatures would require additional evidence.


8. Implications for institutional design and experimentation

The framework is most useful as a way to design tests rather than to label failures after the fact.

1. Define the outcome at the right level

A reform should not be scored by one aggregate indicator if different margins can move differently.

  • Measure both the quantity and price of finance.
  • Measure both leakage and legitimate benefit receipt.
  • Measure both task completion and human review time.
  • Measure both technical accuracy and validity relative to the institution’s actual objective.
  • Measure both service speed and appeal, exclusion, and error-correction rates.

2. Instrument the interfaces

Organizations should measure where capabilities meet:

  • queue lengths and handoff delays;
  • override and exception rates;
  • failed authentications;
  • repeated data entry;
  • review time per automated output;
  • rates of appeal and successful reversal;
  • unresolved ownership of decisions;
  • workarounds and shadow systems;
  • differences across units with otherwise similar resources.

These data can distinguish “not enough capacity” from “the capacities do not connect.”

3. Test complements factorially where possible

The Liberian design is unusually informative because it separates components. Similar experiments could vary:

  • digital identity with and without usable fallback processes;
  • AI tools with and without workflow redesign or dedicated review capacity;
  • interoperable data infrastructure with and without legal authority to share data;
  • autonomous agents with different permission boundaries, logging, and escalation rules;
  • measurement systems with and without credible enforcement;
  • technical decision aids with and without appeal and contestability.

The central question should be whether the combination produces an interaction beyond the independent effects of its parts.

4. Compare matched implementations

The same technology should be studied under different institutional conditions. Especially valuable comparisons would include jurisdictions or firms that share a technical system but differ in:

  • funding;
  • authority;
  • staffing and skill;
  • appeals;
  • privacy protections;
  • fallback procedures;
  • incentives;
  • implementation speed;
  • labor participation;
  • allocation of liability.

This is more informative than comparing a successful technology adopter with a wholly different nonadopter.

5. Preserve bounded substitutes and reversibility

Adequate substitutes can be valuable rather than backward. Manual fallback protected beneficiaries in Andhra Pradesh; timetables made delayed telegraph adoption tolerable. Redundancy, local autonomy, and reversible deployment can absorb capability asymmetry while a new system is being validated.

For AI, this suggests beginning with reversible, logged, low-stakes delegation rather than broad authority. The relevant institutional innovation may be not “AI autonomy” in general but carefully bounded agency with explicit permissions, monitoring, escalation, and human responsibility.

6. Treat power and legitimacy as rival mechanisms, not residual categories

If technically feasible reforms are selectively obstructed by actors who bear losses, further investment in technical capability may accomplish little. A valid diagnosis may point toward bargaining, compensation, representation, conflict-of-interest safeguards, or redistribution—not another tool.

7. Separate feasibility from progress

A capability bundle can make an arrangement workable without making it legitimate. Institutional evaluation should independently specify:

  • who benefits and who bears risk;
  • whose objective is optimized;
  • who can contest a decision;
  • what rights persist under automation;
  • who is accountable for errors;
  • whether participation and exit are meaningful.

This prevents technological feasibility from being smuggled into the analysis as institutional desirability.


9. Uncertainties and unresolved questions

The evidence supports several robust conclusions: specific capability complementarities exist; technology alone is frequently insufficient; and asymmetry is not dysfunctional when substitutes, slack, or bounded tasks remain adequate. But important limits remain.

  • The historical cases are deeply examined illustrations, not a matched comparative design with capabilities measured before outcomes.
  • The proposed diagnostic framework has not been validated prospectively against rival explanations.
  • The bookkeeping counterexample rests on secondary characterization because the central primary article was not directly obtained.
  • The antebellum railroad-telegraph evidence is available as a complete conference abstract rather than the full paper.
  • The evidence does not document a clear case of permanent institutional collapse or durable lock-in.
  • The AI record is too short to establish a full cycle from advance through mismatch, adaptation, stabilization, and renewed mismatch.
  • Several proposed domains overlap, while skill, attention, coercion, legitimacy, material resources, power, and normative objectives may require separate treatment.
  • The evidence identifies arrangements that have become practical, but does not establish that they are institutionally superior.

The framework’s near-term predictive opportunity is consequently modest but real. It may help identify where a newly abundant capability is likely to encounter a bottleneck—AI generation confronting review capacity, digital identity confronting exception handling, interoperability confronting legal authority—provided those predictions are specified before implementation and tested against alternatives.

Conclusion

Capabilities co-evolve through reciprocal constraint, not synchronized progress. New ideas can reveal phenomena that cannot yet be measured; measurement can expose obligations that cannot be enforced; technology can accelerate activity beyond an organization’s capacity to coordinate it; institutions can authorize action while blocking recognition, appeal, or legitimate use.

A consequential mismatch is not simply a gap between what is technically possible and what institutions currently do. It is a task-specific interaction in which an identifiable missing complement constrains performance, leaves observable traces at an interface, and responds to targeted intervention. The Liberian tax and Indian welfare cases show that such mismatches are real. Railroads, electronic voting, bookkeeping, longitude, credible commitment, and containerization show why they cannot be assumed.

The recurring cycle of advance, mismatch, adaptation, and renewed mismatch is therefore a useful heuristic, strongest where actors can identify and control the missing complement. It is not a law of institutional evolution. Adequate substitutes may prevent stress; power may block adoption; a capability bundle may remain unused; technical success may go unrecognized; and a high-performing system may still pursue the wrong objective.

The model’s greatest value is not that it predicts institutional progress from technological progress. It is that, used carefully, it prevents that inference. It directs attention to the complements, interfaces, human capacities, authority structures, and normative judgments that determine whether a possibility becomes workable—and whether a workable arrangement should be adopted at all.


Footnotes

  1. Oyebola Okunogbe, “Becoming Legible to the State: The Necessary but Not Sufficient Role of Identification in Taxation,” American Economic Review, forthcoming, DOI 10.1257/aer.20230586. The identification-plus-penalty-information result and the enforcement-signaling result come from separate experiments. ↩ ↩2 ↩3

  2. Karthik Muralidharan, Paul Niehaus, and Sandip Sukhtankar, “Building State Capacity: Evidence from Biometric Smartcards in India,” American Economic Review 106, no. 10 (2016): 2895–2929, DOI 10.1257/aer.20141346; and Muralidharan, Niehaus, and Sukhtankar, “Identity Verification Standards in Welfare Programs: Experimental Evidence from India,” Review of Economics and Statistics 107, no. 2 (2025): 372–392; working-paper version at NBER. ↩ ↩2 ↩3

  3. Benjamin Schwantes, “Disinterest and Distrust: The Ambivalent Relationship between Railroad Managers and Telegraph Entrepreneurs in Antebellum America,” Business History Conference, 2007, https://thebhc.org/node/1216. Evidence was available as the complete conference abstract; the full paper was not obtained. ↩ ↩2

  4. Basil S. Yamey, “Scientific Bookkeeping and the Rise of Capitalism,” The Economic History Review 1, no. 2/3 (1949), DOI 10.1111/j.1468-0289.1949.tb00108.x, characterized in “Basil Yamey,” Wikipedia, and Matringe, “Accounting Played a Central Role in the Rise of Financial Capitalism as We Know It Today,” LSE Business Review, December 7, 2017, https://blogs.lse.ac.uk/businessreview/2017/12/07/accounting-played-a-central-role-in-rise-of-financial-capitalism-as-we-know-today/. Yamey’s original article was not directly inspected. ↩ ↩2

  5. William Sjostrom, review of Marc Levinson, The Box: How the Shipping Container Made the World Smaller and the World Economy Bigger (Princeton University Press, 2006), EH.Net, https://eh.net/book_reviews/the-box-how-the-shipping-container-made-the-world-smaller-and-the-world-economy-bigger/; “A History of ISO Containers,” S Jones Containers, https://www.sjonescontainers.co.uk/containerpedia/a-history-of-iso-containers/. ↩ ↩2 ↩3

  6. Timothy F. Bresnahan, Erik Brynjolfsson, and Lorin M. Hitt, “Information Technology, Workplace Organization, and the Demand for Skilled Labor: Firm-Level Evidence,” NBER Working Paper 7136, https://www.nber.org/papers/w7136; Ann P. Bartel, Casey Ichniowski, and Kathryn L. Shaw, “How Does Information Technology Affect Productivity? Plant-Level Comparisons of Product Innovation, Process Improvement, and Worker Skills,” Quarterly Journal of Economics 122, no. 4 (2007): 1721–1758; Sinan Aral, Erik Brynjolfsson, and Lynn Wu, “Three-Way Complementarities: Performance Pay, Human Resource Analytics, and Information Technology,” Management Science 58, no. 5 (2012): 913–931. These studies show clustering consistent with complementarity but are primarily observational and do not eliminate reverse causality or omitted managerial quality. ↩

  7. Daniel McCallum’s 1856 report is reproduced in “Daniel McCallum,” Wikipedia, https://en.wikipedia.org/wiki/Daniel_McCallum. On later diffusion, see Max Olson, “Book Notes: The Visible Hand,” FutureBlind, October 8, 2012, https://futureblind.com/2012/10/08/book-notes-the-visible-hand/, characterizing Alfred D. Chandler Jr., The Visible Hand (Harvard University Press, 1977). Chandler’s original book was not directly consulted for this assessment. ↩

  8. Paul A. David, “The Dynamo and the Computer: An Historical Perspective on the Modern Productivity Paradox,” American Economic Review 80, no. 2 (1990): 355–361. ↩ ↩2

  9. “Board of Longitude,” Wikipedia, https://en.wikipedia.org/wiki/Board_of_Longitude. The source describes lunar distance and chronometers as methods used “in conjunction with or instead of” one another, supporting parallel experimentation but not joint necessity. ↩

  10. “John Harrison,” Wikipedia, https://en.wikipedia.org/wiki/John_Harrison. The account supports the Maskelyne conflict-of-interest mechanism and the separate £8,750 parliamentary grant, but a specialist archival or historical source was not obtained. ↩

  11. Douglass C. North and Barry R. Weingast, “Constitutions and Commitment: The Evolution of Institutions Governing Public Choice in Seventeenth-Century England,” The Journal of Economic History 49, no. 4 (1989): 803–832. ↩

  12. Nathan Sussman and Yishay Yafeh, “Constitutions and Commitment: Evidence on the Relation Between Institutions and the Cost of Capital,” The Journal of Economic History 66 (2006). ↩

  13. Ziad Obermeyer, Brian Powers, Christine Vogeli, and Sendhil Mullainathan, “Dissecting Racial Bias in an Algorithm Used to Manage the Health of Populations,” Science 366, no. 6464 (2019): 447–453, DOI 10.1126/science.aax2342. ↩

  14. Thomas Fujiwara, “Voting Technology, Political Responsiveness, and Infant Health: Evidence from Brazil,” Econometrica 83, no. 2 (2015): 423–464, DOI 10.3982/ECTA11520. ↩

  15. Organisation for Economic Co-operation and Development, Digital Government Outlook 2026, DOI 10.1787/0496b2bc-en. ↩

  16. Thomas Kwa, Ben West, Joel Becker, et al., “Measuring AI Ability to Complete Long Software Tasks,” Advances in Neural Information Processing Systems (2025), arXiv:2503.14499. ↩

  17. OpenAI, “How We Monitor Internal Coding Agents for Misalignment,” 2026, https://openai.com/index/how-we-monitor-internal-coding-agents-misalignment/. The deployment figures are first-party and unaudited. ↩

  18. Erik Brynjolfsson, Danielle Li, and Lindsey R. Raymond, “Generative AI at Work,” NBER Working Paper 31161, April 2023, revised October 2023. The figures used here—5,179 agents and a 14 percent average productivity gain—come from the retained working-paper version, not the subsequently published article. ↩

  19. Fabrizio Dell’Acqua, Edward McFowland III, Ethan R. Mollick, et al., “Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality,” Organization Science (2026), DOI 10.1287/orsc.2025.21838. ↩

  20. Joel Becker, Nate Rush, Beth Barnes, and David Rein, “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity,” METR, 2025, arXiv:2507.09089. ↩

  21. International AI Safety Report 2026, chaired by Yoshua Bengio, UK Department for Science, Innovation and Technology, DSIT 2026/001, arXiv:2602.21012. ↩

  22. Jonathan K. Hanson and Rachel Sigman, “Leviathan’s Latent Dimensions: Measuring State Capacity for Comparative Political Research,” Journal of Politics 83, no. 4 (2021): 1495–1510, DOI 10.1086/715066. ↩