Daemon Corporation

Global Leader in Excellence.

Research Report № 009

AI and the Architecture of Collective Action

Cheaper representation is real; inspectable, adaptive organization is not yet demonstrated

Central finding

Contemporary AI does not yet constitute a qualitatively different general form of organization. The stronger description is narrower: AI adds a flexible interpretive layer to familiar organizational architectures and creates a distinctive local operating mode in well-instrumented technical environments. It can translate unstructured requests, documents, conversations, and code into actions; consult procedures during work; route bounded tasks; and let less-experienced workers draw on codified expertise. But the surrounding requirements—authoritative state, deterministic permissions, testing, monitoring, ownership, appeals, and rollback—largely come from older disciplines such as workflow management, policy-as-code, platform engineering, security engineering, and site reliability engineering.

The historical comparison changes the question. Explicit, inspectable, modular, versioned, and partly executable coordination already exists without contemporary AI—in Wikipedia’s governance system, Debian’s constitution, GitLab’s handbook-first practices, public-administration decision systems, workflow engines, and policy infrastructure. The best longitudinal evidence does not show that these properties reliably produce adaptability. Wikipedia’s explicit and partly machine-enforced rule system instead became progressively harder for newer participants to revise. Expert systems remained operational while becoming increasingly difficult to change. AI-era instruction files exhibit the same tendency to accumulate faster than they are safely deleted.123

The most evidence-supported synthesis is therefore:

AI reduces the cost of representing and consulting coordination more clearly than it reduces the cost of governing, revising, or legitimately contesting it. The likely bottleneck is not writing rules but safely removing them—and deciding who has authority to write, interpret, override, and retire them.

Three practical answers follow:

  1. New organizational form or better automation?
    Primarily better automation and interpretation inside familiar forms. A genuinely distinctive local mode is emerging in software and data work, but no unrelated set of organizations yet shows sustained replacement of substantial managerial mediation with explicit, versioned AI coordination infrastructure plus demonstrated gains in diagnosis, bounded autonomy, or safe reconfiguration.

  2. Is there a common architectural kernel?
    Yes as a set of recurring capabilities, but not as a transferable blueprint. Its strongest components are boundaries, authority and permissions, change governance, observation, assurance, and appeals. These predate AI. Their implementations must remain locally congruent with the work, its risks, and its sources of legitimacy.

  3. What study could discriminate the hypothesis?
    A four-arm comparison of one consequential, exception-bearing cross-functional process: conventional operation; AI as an individual assistant; explicit versioned coordination with AI; and the same explicit architecture without AI. It must measure not only speed and manager interventions but rule deletion, amendment power, exception quality, appeals, maintenance effort, and recovery from a seeded policy change.


1. What kind of object is changing?

The hypothesis becomes misleading if every prompt, skill, workflow, or agent configuration is treated as an organizational mechanism. These are different kinds of objects:

  • A function is a coordination problem to be solved: routing work, preserving memory, adjudicating exceptions, enforcing a handoff, or legitimating a decision.
  • A mechanism is the device that produces the effect: retrieval, contextual inference, sandboxing, deterministic validation, logging, peer review, graduated sanctions, or a supermajority threshold.
  • An interface is the boundary across which actors interact: an API, pull request, message queue, permission scope, or issue tracker.
  • A pattern is a recurring solution shape: human escalation, central constraints with local discretion, peer approval, or test-before-deploy.
  • A representation or implementation is a concrete artifact: an AGENTS.md file, prompt, skill, custom GPT, policy bundle, workflow definition, schema, or handbook page.
  • An architecture is a composition of functions, mechanisms, interfaces, representations, owners, feedback loops, and change rights.
  • Explicit coordination infrastructure consists of artifacts and mechanisms deliberately made available for consultation, execution, inspection, enforcement, or revision.
  • Tacit coordination consists of situated judgment, social interpretation, relationships, informal authority, embodied expertise, and practical accommodations that are not fully captured in those artifacts.

This distinction is substantive, not semantic. Organizational-routines theory warns that designing an artifact while expecting it to determine a pattern of collective action is a form of technological determinism: routines are generative systems of interdependent action, not things stored in documents.4 The contemporary evidence repeatedly reproduces that artifact–action gap:

  • instructions can be followed without improving task success;
  • a natural-language security rule can exist without a matching enforcement control;
  • a workflow record can remain compliant while workers coordinate outside it;
  • rules can be explicit and machine-enforced while becoming harder to revise;
  • an instruction file can grow until its owners replace it with a smaller map and a more elaborate maintenance system.

Accordingly, 20,000 custom GPTs are 20,000 representations, not an organizational architecture. An architecture would additionally reveal who owns them, what permissions they exercise, how conflicts are adjudicated, how effects are evaluated, and how obsolete components are retired.


2. What AI actually makes cheaper

The evidence supports substantial but bounded reductions in coordination cost.

Action-time consultation and onboarding

The strongest identified result comes from 5,179 customer-support agents and roughly three million chats. AI assistance increased issues resolved per hour by 14% on average and by 34% for novice and lower-skilled workers, while producing little benefit for experienced high performers. AI therefore can make previously learned practice available during action and shorten part of the experience curve.5

This supports the explicitness hypothesis in a specific form: knowledge need not first be converted into a fully formal decision tree to become usable. A model can interpret examples, natural language, and heterogeneous records at the moment of work.

Individual work more than coordinated work

A six-month randomized experiment across 66 firms and 7,137 knowledge workers found that frequent users spent 3.6 fewer hours per week on email, with a 1.3-hour intent-to-treat reduction, and completed some document work faster. Meeting time did not change significantly. The authors explicitly distinguished activities an individual could change unilaterally from those requiring coordinated organizational change.6

This is important counterevidence. Making individual work faster does not, by itself, change authority, handoffs, shared priorities, meeting structures, or managerial mediation.

Bounded execution at scale

Large sustained deployments exist. Klarna reported 31 million AI-handled conversations since launch, 80% of service chats during 2025, and $59 million in estimated 2025 cost savings. Its comparison of two-minute AI resolution with 12-minute human resolution was measured in September 2024, not as a general 2025 result, and the filing preserves access to human support.7 BBVA has reported broad employee adoption and decentralized creation of specialized GPTs, although the public evidence does not reveal their operative governance. These cases establish operational scale, not architectural transformation.

In software, an OpenAI engineering team reported approximately one million lines of agent-generated code and about 1,500 merged pull requests over five months. Agents reproduced failures, implemented and validated fixes, opened and revised pull requests, handled build failures, and merged changes inside a versioned environment of documentation, tests, linters, evaluations, deployment tools, and cleanup agents.8 This is the strongest AI-era example of a composed operating architecture, but it is greenfield, vendor-authored, only five months old, and began with three engineers rather than with a managerial layer that could be displaced.

The cost ledger

Coordination cost Best-supported effect of AI Qualification
Writing an explicit instruction Strongly reduced Cheap addition can accelerate rule proliferation
Consulting procedure at action time Substantially reduced Gains concentrate among less-experienced workers
Interpreting unstructured intent Reduced; AI’s clearest novelty Reliability varies with context and task boundaries
Routing and executing bounded technical work Reduced Requires instrumented tools, authoritative state, and external controls
Preserving organizational memory Easier to capture and retrieve Staleness, provenance, access, and deletion remain unresolved
Onboarding Partly reduced Relational and developmental aspects may be displaced
Enforcing handoffs or permissions Little inherent reduction Enforcement still depends on deterministic platform mechanisms
Resolving conflicts among rules Unchanged or worsened Declared natural-language precedence is unreliable
Detecting exceptions Supported in bounded technical cases Recognizing that a case lies outside the system’s competence remains difficult
Observing failures Improved only when deliberately instrumented AI does not automatically create usable traces or causal diagnosis
Revising procedures Locally easier for technical artifacts No comparative evidence of safer organizational revision
Safely deleting obsolete rules Not shown to improve Cheap addition makes this constraint more important
Supporting local variation Technically feasible It may produce either bounded autonomy or ungoverned sprawl
Protecting integrity of instructions and memory Potentially worsened Prompt injection and larger tool-access surfaces remain open risks

The asymmetry is central. AI makes it easier to add representations but has not shown a corresponding ability to determine when a rule is obsolete, what depends on it, whether its original rationale still applies, or who has authority to remove it.


3. The deletion problem: explicitness can produce calcification

The strongest challenge to the adaptability hypothesis is not a failed AI prototype. It is the recurrence of the same maintenance problem across expert systems, online communities, and current agent instructions.

Expert systems worked—and became difficult to change

Digital Equipment Corporation’s XCON grew to 6,200 rules over seven years, with approximately half changing annually. Its performance remained satisfactory, but its engineers reported that it was becoming increasingly difficult to modify. The system was not “defeated” by rule-based reasoning; it continued to work while becoming progressively harder to revise safely.2

That distinction matters. The problem was not whether explicit rules could coordinate production. It was whether a large, interdependent rule base could remain adaptable.

AI-era files reproduce the accumulation pattern

A study of 247,694 instruction lifetimes across 1,867 repositories found agent instructions growing by 226% over their lifetimes, with additions outpacing deletion and the probability of deletion declining as instructions aged. The proposed explanation is lost rationale: once maintainers no longer know why a rule exists or what depends on it, removal becomes risky.3

In synthetic, verifiable environments, recording rationale removed 99.3% of excess instructions and improved instruction-following. That is promising engineering evidence, not yet organizational evidence. Real organizations may omit or distort rationale precisely when rules embody bargains, liability concerns, political compromises, or informal power.

Wikipedia is the decisive longitudinal case

Wikipedia combines most of the hypothesized properties without contemporary AI:

  • explicit natural-language coordination rules;
  • public revision history;
  • modular distinctions among policies, guidelines, and essays;
  • formal amendment procedures;
  • inspectability;
  • enforcement partly delegated to bots and semi-automated tools.

Over time, however, its formal governance became less revisable. Policy acceptance fell sharply after 2005, and analysis of 120,535 registered-editor contributions to norm pages found rejection increasing over time and newer editors’ rule contributions being rejected more often. The authors describe the result as calcification: rules became less open to revision by the people they governed.1

Informal coordination did not disappear. Editors increasingly created essays, including essays that began as unsuccessful policy proposals. But essays were not enforceable and were overridden by rules embedded in tools. Informal work therefore expanded without restoring equivalent influence over the operative architecture.

Automated enforcement was genuinely useful: the authors judged that removing it would require much more volunteer effort and probably reduce quality. But tool-based rejection of desirable newcomers’ first-session contributions rose from 0% in 2006 to 40% in 2010, and tool reversion was a significant negative predictor of newcomer survival. The case is therefore not “automation failed.” It is that enforcement capacity, quality protection, participation, and revisability moved in different directions.

Implication

Evidence-supported implication: the binding constraint may be less the cost of representing coordination than the cost of safely un-representing it.

Cheap rule production without symmetric deletion creates predictable pressure toward:

  • longer and more contradictory instruction sets;
  • dependence on a small set of maintainers;
  • strictness drift;
  • loss of rationale;
  • fear of removing old controls;
  • informal workarounds around the formal system;
  • declining access to rule-making for newcomers.

A genuinely adaptive architecture therefore needs a theory of retirement, not merely versioning. A history of additions is not the same as an ability to remove.


4. Explicitness is not inspectability

Explicit artifacts can support inspection, but only when observation and enforcement are separately designed.

Written rules do not reveal whether they are enforced

In a study of 481 repositories containing 4,661 instruction segments, only about 4–16% of security rules written in CLAUDE.md files had a matching built-in control; the strict estimate was 4.4%. This does not mean the remaining rules had no enforcement—hooks and other mechanisms were outside the measurement—but it does show that the text itself does not tell an author what is mechanically enforced.9

The same plain-text form can contain:

  • a rule enforced by a sandbox;
  • a rule checked by a linter;
  • a rule evaluated by the model;
  • a suggestion;
  • an aspiration;
  • or an unenforced prohibition.

A natural-language instruction channel can therefore be explicit yet effectively “write-only” with respect to enforcement state.

Conformance is not effectiveness

Across 160 rule-evolution events in AI coding tools, updated rules raised average artifact compliance from 49.14% to 72.13%.10 But a controlled evaluation of repository-level context files found no statistically significant average improvement in task success and more than 20% additional inference cost. Developer-authored files produced a positive 2.4% estimate that was not statistically significant, while outperforming LLM-generated files by a statistically significant 7%.11

This result is more informative than a simple null. It suggests:

  • the presence of an explicit artifact is not enough;
  • automatically generated explicitness may be harmful;
  • human authorship and local knowledge carry much of the value;
  • following instructions can increase without improving the underlying outcome.

The unresolved question is whether conformance is an intermediate good whose benefits appear later or a legibility metric that substitutes for measuring effectiveness.

Formal records can conceal enacted practice

The PRINTFLOW PF2 workflow system in a printing operation could not represent urgent jobs begun before job numbers existed, concurrent work, collaboration around a single machine, or handovers across breaks. One site preserved the formal system through costly workarounds. Other sites kept paper dockets and had an administrator reconstruct activity retrospectively, preserving the accountability record while forfeiting real-time coordination and monitoring.12

The system remained formally legible while actual coordination moved elsewhere. This is the danger of a legible façade: a record can demonstrate conformity to an external principal while concealing how work is accomplished.

External accountability can drive explicitness independently of coordination benefit. PF2 was contractually required. Contemporary policy-as-code adoption is similarly associated with security, compliance, and regulatory demands.13 Growth in explicit artifacts therefore cannot by itself be interpreted as evidence that the artifacts improve coordination.

Inspectability must be designed

Diagnosis is supported where systems deliberately provide:

  • agent-attributed audit events;
  • retained execution traces;
  • source links and provenance;
  • inspectable intermediate results;
  • decision records;
  • structured incident histories;
  • accessible version histories.

OpenAI’s internal data agent, for example, combines organizational context, reusable workflows, pass-through access control, human-authored evaluations, and inspectable query results.14 GitHub’s agent control plane records the human principal on whose behalf an agent acts and protects canonical agent-definition paths; some tool allowlisting features were still in public preview at the time of the evidence.15

These are implemented observation and control mechanisms. They are not automatic consequences of using a model.

The net inspectability effect of shifting conversations from colleagues to models remains undetermined. Model interactions may be more traceable than hallway conversations—or less available to ordinary employees—depending on retention, access, search, and privacy policy. The evidence does not establish the direction. GitLab’s pre-AI handbook practice offers one explicit design response: retain chat for only 90 days so that durable decisions must move into universally accessible issues, merge requests, and handbook pages.16


5. Modularity works when boundaries are real

The modularity hypothesis receives qualified support, but primarily from externally enforced boundaries.

Examples include:

  • operating-system sandboxing that restricts writable locations and network access;
  • pass-through permissions that inherit an employee’s existing data rights;
  • structural tests that prohibit invalid dependencies;
  • role-based access control;
  • protected configuration paths;
  • audit identity linking an agent action to a human principal;
  • signed policy bundles and scoped ownership;
  • peer-reviewed changes and staged deployment.

Anthropic reports that sandboxing reduced permission prompts by 84%, while approximately 93% of prompts had previously been approved—evidence that repeated click-through approval was becoming ritual rather than meaningful oversight.17

The contrast is sharp:

  • “The agent may not write outside this directory,” enforced by an operating-system sandbox, is a mechanism.
  • The same sentence in a prompt is only a representation of a desired boundary.

Natural-language instruction hierarchies do not reliably resolve even simple conflicts. Controlled testing across six models found that declared system/user precedence was unstable. More surprisingly, hierarchy framed through social authority, expertise, or consensus sometimes influenced model behavior more strongly than formal message position; for one model, adherence rose from 47.5% under system/user framing to 77.8% under societal framing.18

This has two implications. First, organizational authority cannot safely be implemented by instruction wording alone. Second, part of an agent’s effective authority structure may come from patterns learned in pretraining rather than from any versioned artifact that the organization can inspect or roll back.

Bounded autonomy is therefore most credible when:

  1. permissions are enforced outside the model;
  2. authoritative state is separated from generated interpretation;
  3. actions are attributable;
  4. invalid outputs fail closed;
  5. escalation is possible;
  6. consequences are reversible.

Evidence for this arrangement is strongest in software and data systems. Transfer to socially contested decisions remains unproven.


6. Tacit coordination is displaced more often than eliminated

The strongest form of the tacit-limit hypothesis is not that judgment, trust, or political negotiation are metaphysically impossible to formalize. The evidence does not support that claim. It supports a more operational one:

Formalization often relocates tacit coordination into less visible work: encoding, exception classification, workarounds, data curation, evaluation design, escalation, and informal interpretation of the formal system.

Discretion moves to encoders

Research on automated public administration documents discretion shifting from front-line case managers to programmers and data analysts who translate legislation into software. Yet formal authority did not shift with it. Later work in the same tradition concludes that discretion may be concealed rather than eliminated.19

The mechanism is organizational. Software artifacts can replace participatory boundary practices among operations, legal interpretation, rule drafting, management, and appeals. The resulting rules may look more objective while the real interpretive choices move upstream into categories, code, data, and exception definitions.

A comparative study across 14 countries found that web-based service delivery did not produce greater accountability. Detailed regulation required for automated decisions could also narrow the discretion of the judge reviewing an appeal. Explicitness did not necessarily enlarge contestability; it could constrain the appeal route itself.19

Workers develop expertise about the system

Algorithmic management uses mechanisms such as recording, rating, recommending, restricting, replacing, and rewarding. Workers respond by developing practical knowledge about how assignments, ratings, and penalties behave and by coordinating informally around the system.20 This is tacit expertise about how to work the formal architecture.

The same pattern appeared in Wikipedia’s unenforceable essays and in PF2’s paper dockets. Formalization did not simply reduce tacit coordination; it generated a second layer of tacit coordination about the formal layer.

Relational and emotional work remains difficult

The customer-support evidence shows that AI is most helpful where workers lack accumulated expertise. That does not mean experienced judgment is eliminated; it means part of its output becomes available to novices.

A randomized Alibaba field experiment found that agentic AI shortened customer-service interactions but reduced customer ratings in AI-eligible conversations. Human intervention preserved quality more successfully in technical escalations than in emotional ones; in emotional cases, workers intervened less proactively and gathered less information.21 This evidence was available only at abstract or partial-text level, so it should be treated as directional rather than definitive.

Interviews with 24 professionals at a large AI-first technology company also found a two-sided pattern: smoother peer collaboration alongside risks to mentoring, feedback seeking, informal leadership, and the interactions through which careers develop.22 A separate internal Anthropic study found that about half of respondents reported unchanged collaboration. One interviewee said that most routine questions had moved to the model but emphasized that the remaining questions were crucial and still went to colleagues. This is evidence of triage, not wholesale social substitution.

What remains unresolved

No current evidence establishes that AI systems can legitimately assume responsibility for:

  • deciding organizational purpose;
  • resolving contested values;
  • bargaining over rules;
  • allocating voice and standing;
  • repairing trust;
  • sustaining identity;
  • exercising informal authority;
  • determining who may rewrite the architecture.

These may not be permanently non-delegable. They are, however, currently binding limits, and current deployments offer almost no evidence on appeals, bargaining, worker influence over instructions, or oversight of architecture maintainers.


7. A common kernel exists—but it is older than AI

A review of 91 empirical studies of common-pool-resource governance found substantial support for Ostrom-derived design principles, while warning against treating them as an institutional panacea and leaving transfer across scale unresolved.23 Those principles map onto nine of the twelve capabilities proposed in the research question.

Debian’s ratified constitution supplies another pre-AI implementation. It distinguishes constitutional decision-making from goals and operating policy; enumerates authority among six bodies; specifies a 3:1 amendment threshold; allows developers to override the Technical Committee at 2:1; provides an override route for delegate abuse; and assigns procedural assurance to a Project Secretary.24

The mapping suggests two different architectures have been bundled together:

  1. a governance architecture for authority, observation, amendment, assurance, and contestation;
  2. an operating architecture for capability selection, execution, interfaces, learning, and rollback.
Candidate capability Evidence-qualified status
Purpose and boundaries Strong recurring governance capability; implementation must remain context-specific
Authority and permissions Strong; distinct from textual instruction precedence
Change governance Strong and recurrent; access to amendment is as important as formal versioning
Observation Strong, but social monitoring and resource/system monitoring should remain distinct
Assurance Strong; includes review, testing, sanctions, and procedural integrity
Appeals and conflict resolution Strong and often omitted from AI-era proposals
Strategy and allocation Recurrent but highly contextual and politically contested
Interfaces and handoff contracts Strong technically; only partial evidence as a general social-organizational module
Learning Plausible function, but not clearly a distinct universal architectural component
Capability management Prominent in AI proposals; little independent collective-action evidence
Execution Mature in technical workflow systems; not validated as a general governance capability
Rollback Mature in software release engineering; largely unevidenced for consequential social decisions

The strongest kernel is therefore not a universal agent stack. It is a governance capability set:

Boundaries → authority → observation → assurance → amendment → appeal

These capabilities recur across very different systems, but recurrence of the function does not imply transferability of the implementation. A supermajority threshold, sandbox, workflow definition, evaluation suite, or escalation rule cannot be assumed to travel unchanged from an open-source project to a bank, public agency, customer-support operation, or professional partnership.

The commons literature itself says why: institutions must fit local conditions. The OpenAI harness team independently makes the same point, warning that its autonomy depends on repository-specific structure and investment.8

DAOs clarify what happens when execution outruns adjudication

On-chain DAOs implement authority, voting, change governance, and execution cryptographically. They generally do not implement comparably strong appeals, graduated sanctions, low-cost conflict resolution, local congruence, or nested governance.

A study of 21 DAO governance systems found voting-power Gini coefficients close to 1.0, single-digit Nakamoto coefficients for delegate voting power in all but four projects, and half of the projects requiring no more than three addresses to pass a decision. It also documented substantial useless voting and millions of dollars in direct or indirect governance costs.25

The appropriate conclusion is not that “organizations-as-code failed.” It is narrower and more useful: DAOs implemented much of the enforceable half of the architecture while omitting much of the adjudicative half, and then exhibited the concentration problems those omitted capabilities exist to address.

Any AI-era architecture that emphasizes execution, permissions, and audit while omitting appeals and conflict resolution risks repeating that composition.


8. What operational deployments actually show

A case should count as organizational architecture rather than task automation only if it demonstrates:

  1. explicit, versioned coordination artifacts;
  2. consultation or enforcement during action;
  3. substitution for a human mediation role;
  4. sustained use beyond a single project cycle;
  5. documented revision, failure, and rollback history;
  6. evidence on effects.

No AI-era case in the evidence base clears all six conditions.

Case What became explicit or executable What remained tacit or human Control and revision Evidence-qualified assessment
OpenAI harness Repository documentation, plans, schemas, architectural rules, tests, evals, deployment and cleanup Prioritization, acceptance criteria, user interpretation, architecture, exception judgment Small senior engineering team; repository and raw data not public Strongest AI-era architecture; five months; replaced implementation work more than managerial mediation; effects not independently established
OpenAI data agent Context layers, memory, workflows, permissions, expected-query evals, query provenance Semantic ownership, tool selection, annotations, exception handling Internal owners with pass-through user permissions Production implementation of several capabilities, not evidence of organizational redesign
Klarna Customer-service automation and internal knowledge assistance Human exception service, policy ownership, complex-case judgment Not publicly inspectable Sustained task automation with corporate performance disclosures
BBVA Decentralized employee GPT creation and specialized assistants Approval, maintenance, conflict resolution, retirement, cross-unit governance Security, legal, and compliance oversight reported, but not publicly documented in detail Large-scale adoption; architecture and causal effects remain opaque
Singapore Pair Government-specific assistant in sustained use Operational governance not publicly recoverable Public dashboard does not expose substantive metric values Existence established; quantitative claims unverifiable
Wikipedia Public, versioned policies and machine-assisted enforcement Consensus, interpretation, informal essays, volunteer motivation Formal amendment plus senior-editor influence Best longitudinal case; explicitness and enforcement coexisted with calcification
Debian Ratified constitutional authority, amendment, assurance, and override Voluntary contribution and operating coordination Developers, Technical Committee, Leader, Secretary, delegates Strong implemented governance architecture; no outcome study
GitLab handbook Durable policies, decisions, issues, merge requests, ownership and staleness checks Interpretation, authorship, management, local judgment Handbook editors and code owners Sustained explicit coordination practice; no causal outcome evidence
DAOs On-chain authority, voting, change, and execution Bargaining, legitimacy, appeals, conflict resolution Token holders and delegates, often concentrated Architectural rather than task-level; measured outcomes are cautionary

The strongest AI-era positive case is therefore real but narrow. Its human work—prioritization, architecture, acceptance criteria, evaluation, and exception judgment—resembles senior engineering, product management, platform engineering, and SRE. The evidence does not yet show whether this is a new managerial role or a rebundling of established ones.


9. Predecessors already supply most of the control architecture

Predecessor regime What it achieved Binding constraint What AI changes
Bureaucratic rules Durable, transferable, auditable procedure Rules can lack local meaning or become performative façades AI lowers consultation and interpretation cost, not the legitimacy problem
Cybernetics Faster organizational information flow Evidence on durable outcomes is limited; Cybersyn ended exogenously Not assessable from this evidence
Expert systems Sustained high-quality rule-based production Rule interaction and revisability at scale AI does not yet solve safe deletion
BPM and workflow Executable processes, branching, human tasks, versioning, migration, audit Requires formal input; enacted work can diverge from models AI lowers the cost of translating unstructured input into formal routes
Policy-as-code Distributed evaluation, signed policy bundles, decision logs, scoped ownership Maintenance concentration, strictness drift, compliance-driven scope Little change to enforcement or ownership requirements
Enterprise knowledge systems Centralized organizational memory Staleness, concentrated authorship, distrust, reversion to asking humans AI improves retrieval; ownership and freshness remain necessary
Algorithmic management Large-scale recording, rating, routing, and control Rationale opacity, contestation, worker counter-practices AI may worsen semantic opacity while broadening control
Automated public administration Major transfer of operational discretion into software Formal authority and accountability did not follow discretion AI reproduces rather than resolves the meta-governance problem
DAOs / organizations-as-code Cryptographic authority, voting, amendment, and execution Concentration and weak adjudication AI does not supply legitimacy or appeals
SRE and release engineering Versioning, monitoring, staged deployment, canaries, rollback Domain-specific and dependent on deterministic state Provides the mature technical substrate on which agents can operate
Commons governance Recurring institutional capabilities across diverse settings Local congruence and scale Demonstrates that the kernel is not an AI invention

BPMN, Camunda, Open Policy Agent, and SRE already provide versioned processes, migration, decision telemetry, signed policy distribution, controlled release, and rollback.26 AI’s most defensible novelty is not these capabilities. It is a probabilistic semantic adapter between unstructured human material and formal tools.

Two predecessor literatures named in the research question—information-processing or computational organization theory, and platform organizations—were not substantively assessed in the available evidence. No conclusion about them is warranted.


10. The managerial bottleneck moves—but authority may not

The bottleneck-shift hypothesis is supported as a phenomenon and not supported as an AI-specific effect.

Public administration documented it decades ago: discretion moved from front-line professionals to programmers and analysts responsible for translating law into software. Formal power structures remained nominally unchanged.19 Continuous delivery produced an analogous technical shift before contemporary generative AI. DORA associated external approval bodies with worse delivery performance and no lower change-failure rate, and recommended moving review into peer-reviewed, automated development workflows.27 That is an association and recommendation, not direct observation of managers reallocating their time.

AI-era engineering accounts show humans concentrating on:

  • prioritization;
  • problem framing;
  • acceptance criteria;
  • architecture;
  • permission design;
  • evaluation;
  • exception handling;
  • stewardship of documentation and memory;
  • assurance and change governance.

These are plausible future bottlenecks, but they are also recognizable forms of existing senior technical and managerial work.

The more consequential finding concerns authority:

When discretion moves to the people who encode and maintain the architecture, formal authority to oversee them does not necessarily move with it.

This creates a meta-governance gap. Architecture maintainers may exercise substantial practical power while appearing to perform technical implementation under someone else’s formal authority. Their choices can remain invisible because they are embedded in classifications, evaluation sets, retrieval corpora, permissions, defaults, and exception definitions.


11. Meta-governance is the missing layer

An explicit operating architecture needs rules not only for action but for rule-making. At minimum, it must allocate authority to:

  • write;
  • approve;
  • interpret;
  • prioritize;
  • override;
  • test;
  • deploy;
  • audit;
  • appeal;
  • retire;
  • and roll back instructions and permissions.

Current model constitutions encode precedence, not legitimacy

The OpenAI Model Spec defines a six-level authority lattice, supersession rules, default inaction for irresolvable high-level conflicts, and a “No Authority” category for tool outputs, quoted text, and other untrusted material. Claude’s Constitution gives a broad priority ordering, permits deviation from lower-level guidelines when they conflict with higher ethical considerations, and describes itself as a continuing work in progress.28

These are genuine model-governance artifacts. They make precedence more explicit. They do not provide:

  • separation of powers;
  • participation by affected actors;
  • independent review;
  • public amendment procedures;
  • organizational appeals;
  • bargaining over the encoded rules;
  • oversight of maintainers.

Debian’s constitution is stronger on these dimensions despite predating contemporary AI. It has ratification, enumerated bodies, amendment thresholds, overrides, and procedural assurance. It also states its own limit: nothing in the constitution obliges anyone to do work for the project.24 Formal authority can structure decisions without manufacturing motivation, commitment, identity, or voluntary effort.

The architecture is not fully inside the organization’s artifacts

Instruction-conflict research adds a novel difficulty. If models respond to learned social hierarchies more strongly than to declared system/user precedence, then the operative authority structure is partly embedded in pretraining. It is not fully represented in any policy file the organization can inspect, approve, or roll back.18

This complicates proposals for agent self-modification. No deployed evidence in the available base shows agents safely rewriting their own governing architecture. If an agent modifies visible instructions while its behavioral priorities also depend on opaque pretrained associations, formal version control captures only part of the change.

The empty categories matter

The contemporary evidence contains:

  • no well-documented appeal against an AI-mediated coordination decision;
  • no union or works-council agreement governing shared agent instructions;
  • no observed bargaining process over model-maintained organizational memory;
  • no implemented external oversight mechanism for architecture maintainers;
  • no deployment evidence for agent self-modification;
  • no published case of rollback or abandonment of AI coordination infrastructure.

These are not proof that the capabilities are impossible. They are the missing evidence required before technical governance can be treated as organizational governance.


12. Verdict on the seven hypotheses

Hypothesis Verdict
H1 — Explicitness Supported, with the payoff unproven. AI lowers the cost of creating and consulting explicit representations. Automatically generated representations can be ineffective or harmful, and human authorship still carries much of the observed value.
H2 — Inspectability Supported only as a deliberately engineered capability. Traces, provenance, audit identity, and decision records can improve diagnosis. Explicit text alone does not reveal enforcement, rationale, or enacted practice.
H3 — Modularity Qualified support. Bounded autonomy works when permissions and interfaces are externally enforced. Prompted boundaries and textual precedence are not reliable substitutes.
H4 — Adaptability Locally demonstrated for technical artifacts; unsupported comparatively and contradicted in the strongest longitudinal collective-action case. Mature workflow and SRE systems already provide controlled revision and rollback. Wikipedia shows that explicit, versioned, executable coordination can become less revisable.
H5 — Common architecture Split verdict. A recurring capability set is credible and predates AI. Transfer of implementations across contexts is not supported and is contradicted by local-congruence requirements and deployment warnings.
H6 — Bottleneck shift Confirmed as a phenomenon, not as an AI effect. It occurred in automated public administration and continuous delivery before generative AI. Present AI evidence shows role rebundling more clearly than wholesale organizational redesign.
H7 — Tacit limit Strongly supported in displacement form. Formalization often moves judgment, discretion, relationships, and politics into encoding, stewardship, exceptions, and workarounds. “Currently binding” is better supported than “irreducible.”

13. Design implications

The following are evidence-supported design implications, not demonstrations of superior organizational performance.

1. Separate interpretation from enforcement

Use models to interpret ambiguous material and propose routes. Use deterministic mechanisms for:

  • permissions;
  • financial or legal limits;
  • data access;
  • prohibited actions;
  • dependency constraints;
  • identity and attribution;
  • deployment gates;
  • rollback.

Natural-language instructions should not be treated as the authorization layer.

2. Design for deletion before designing for generation

Every coordination rule should ideally have:

  • an owner;
  • a stated rationale;
  • provenance;
  • affected dependencies;
  • a review or expiry date;
  • evidence of use;
  • a retirement procedure;
  • an appeal route.

Measure deletion rate and rule age, not only the number of instructions created.

3. Distinguish compliance from outcome

Evaluation should separately measure:

  • instruction compliance;
  • task quality;
  • system-level performance;
  • exception handling;
  • rework;
  • effects on participants;
  • distribution of amendment power.

A rise in conformance is not evidence of improved coordination.

4. Make the enforcement state visible

Authors should be able to tell whether a rule is:

  • mechanically enforced;
  • tested;
  • monitored;
  • interpreted by a model;
  • advisory;
  • or unsupported.

Without that distinction, nominally explicit coordination remains operationally opaque.

5. Treat appeals as part of the architecture

Appeals are not an optional “human in the loop.” They require:

  • standing to challenge;
  • access to the decision record;
  • an adjudicator independent enough to review it;
  • authority to grant a remedy;
  • a route for changing the underlying rule;
  • protection against retaliation or exclusion.

AI-era architecture proposals are generally stronger on execution than on adjudication.

6. Govern maintainers, not just agents

Organizations need explicit answers to:

  • Who can alter shared instructions and memories?
  • Who approves changes?
  • Who can see downstream effects?
  • Who audits the maintainers?
  • Can affected workers initiate amendments?
  • Can they contest classifications and evaluation criteria?
  • What happens when technical discretion exceeds formal authority?

This is the central institutional risk revealed by the predecessor evidence.

7. Preserve the relational channels that exceptions need

If routine questions move to agents, the remaining human interactions may become rarer but more consequential. Organizations should not infer that reduced question volume means reduced need for colleagues, mentorship, trust, or escalation capacity.

8. Prefer local congruence over universal templates

Reusable capability names are plausible; universal rule sets are not. The architecture should vary with:

  • consequence severity;
  • reversibility;
  • professional expertise;
  • regulatory obligations;
  • emotional and relational content;
  • sources of legitimate authority;
  • availability of reliable outcome measures.

14. The smallest discriminating real-world study

The most useful next step is not another demonstration or adoption survey. It is a controlled deployment around one consequential, exception-bearing, cross-functional process.

Four arms

  1. Conventional operation
  2. The same AI model used only as an individual assistant
  3. Explicit versioned coordination architecture with AI
  4. The same explicit versioned architecture without AI

The fourth arm is essential. Without it, any observed gain could come from explicitness, workflow redesign, better ownership, or improved measurement rather than from AI interpretation.

Architecture in arms 3 and 4

Both architectural arms should contain:

  • owned instructions;
  • deterministic permissions;
  • explicit interfaces and handoffs;
  • version control;
  • execution traces;
  • outcome evaluations;
  • exception routing;
  • appeals;
  • controlled deployment;
  • rollback;
  • retained rationale.

The difference should be whether an AI model interprets unstructured inputs and selects among routes.

Pre-registered measures

Measure:

  • manager interventions;
  • meeting time;
  • handoff latency;
  • completion time;
  • rework;
  • policy conflicts;
  • unauthorized actions;
  • exception quality;
  • appeals raised and remedies granted;
  • instruction-maintenance time;
  • integration effort;
  • recovery after a seeded policy change.

Retain every artifact revision and decision trace.

Two mandatory governance measures

Deletion and retirement

  • instructions added;
  • instructions removed;
  • mean age at deletion;
  • reason for deletion;
  • dependencies discovered during deletion;
  • obsolete instructions left in place because removal was considered unsafe.

Distribution of amendment power

  • who proposes changes;
  • who approves them;
  • success rates by tenure;
  • success rates by proximity to the work;
  • success rates for actors affected by the rule;
  • whether informal workarounds appear when formal amendments fail.

Discriminating outcomes

The inspectable, adaptive-organization hypothesis would receive meaningful support only if the AI architecture:

  • reduces managerial mediation and integration cost;
  • preserves or improves exception quality;
  • avoids increased unauthorized action;
  • supports faster and safer recovery from the seeded change;
  • does not accumulate instructions faster than it can retire them;
  • and does measurably better than the same explicit architecture without AI.

Interpretation would then be straightforward:

  • Arm 3 better than arm 4: evidence for a distinct contribution from AI interpretation.
  • Arms 3 and 4 both better than conventional operation: benefit from explicit architecture, not necessarily AI.
  • Arm 2 improves individual speed but not coordination: task assistance rather than architectural change.
  • Both architectural arms calcify: recurrence of the predecessor failure.
  • Formal metrics improve while workarounds grow: a legible façade rather than substantive coordination improvement.

Important uncertainties

The conclusion is bounded by several weaknesses in the evidence.

  • The richest AI-era deployment accounts are published by model vendors or deploying companies. The strongest architectural case is a vendor engineering blog, not an independent longitudinal study.
  • Self-reported productivity is especially unreliable here. Experienced open-source maintainers were 19% slower with early-2025 AI tools while believing they were faster; later METR estimates also remained negative in point estimate, though their confidence intervals included zero.29
  • A prominent consultant experiment concerning performance outside the model’s competence frontier could not be substantively verified and has not been used to support this report’s conclusions.
  • Singapore’s Pair establishes the existence of a government-wide assistant in sustained use, but its quantitative dashboard claims could not be recovered from the available page.
  • The evidence is concentrated in software, data, customer support, open-source governance, public administration, and online communities. Healthcare, education, logistics, manufacturing, and military command are absent or weakly represented.
  • No measured base rate was obtained for divergence between formal process models and enacted practice.
  • Information-processing or computational organization theory and platform-organization research were not substantively assessed.
  • Gaming, procedural churn, and dependency problems were named as risks but were not directly measured.
  • No published case documents the rollback or abandonment of AI coordination infrastructure, even though such cases would be especially informative about adaptability.
  • No evidence shows deployed agent self-modification under a complete governance regime.
  • The direction of the net organizational-memory effect remains unknown because transcript retention, discoverability, and employee access are rarely disclosed.

Conclusion

AI has changed the economics of organizational explicitness, but not yet the fundamentals of organizational governance.

It is now cheaper to turn unstructured material into action-time guidance and to let agents execute bounded workflows inside well-instrumented environments. This can compress parts of the experience curve, redistribute implementation capability, and reduce routine coordination work. In technical settings, it can support a powerful composition of interpretation, deterministic permissions, evaluation, version control, and escalation.

What it has not demonstrated is equally important. AI has not removed the need for authoritative state, enforcement, ownership, adjudication, or rollback. It has not shown that consequential organizational procedures become safer to revise. It has not established that prompts or workflows transfer as modules across heterogeneous organizations. It has not made legitimacy, bargaining, trust, or power over rule-making into technical problems. And it may intensify an old asymmetry: rules become easier to add than to understand, contest, or remove.

The common architecture that survives scrutiny is therefore less an “agentic operating system” than a governance kernel surrounding a probabilistic interpreter:

clear boundaries, real permissions, observable action, independent assurance, governed amendment, and accessible appeal—plus an explicit capacity to retire what no longer belongs.

Most of that kernel predates AI. The genuinely new possibility is that models may extend it into domains where coordination could not previously be formalized cheaply. Whether that produces a more adaptive organization or a faster-growing bureaucracy of instructions remains an empirical question. The next decisive evidence will not be another count of agents, prompts, or generated code. It will be a comparative record showing who intervened, who could amend the rules, what happened to exceptions, whether obsolete rules were actually deleted, and whether AI added value beyond explicit architecture itself.


Footnotes

  1. Aaron Halfaker, R. Stuart Geiger, Jonathan T. Morgan, and John Riedl, “The Rise and Decline of an Open Collaboration System: How Wikipedia’s Reaction to Popularity Is Causing Its Decline,” American Behavioral Scientist 57, no. 5 (2013): 664–688, doi:10.1177/0002764212469365. Full author-hosted text: https://stuartgeiger.com/papers/abs-rise-and-decline-wikipedia.pdf. ↩ ↩2

  2. Elliot Soloway, Judy Bachant, and Keith Jensen, “Assessing the Maintainability of XCON-in-RIME: Coping with the Problems of a VERY Large Rule-Base,” Proceedings of AAAI-87, vol. 2 (1987), https://cdn.aaai.org/AAAI/1987/AAAI87-148.pdf. ↩ ↩2

  3. Kushal Chakrabarti, “Why Does CLAUDE.md Keep Growing? Catastrophic Remembering in Agentic Coding,” arXiv:2608.11095 (2026), https://arxiv.org/abs/2608.11095. ↩ ↩2

  4. Brian T. Pentland and Martha S. Feldman, “Designing Routines: On the Folly of Designing Artifacts, While Hoping for Patterns of Action,” Information and Organization 18, no. 4 (2008): 235–250, doi:10.1016/j.infoandorg.2008.08.001. ↩

  5. Erik Brynjolfsson, Danielle Li, and Lindsey R. Raymond, “Generative AI at Work,” Quarterly Journal of Economics 140, no. 2 (2025): 889–942, doi:10.1093/qje/qjae044; working-paper version, doi:10.3386/w31161. ↩

  6. Eleanor W. Dillon, Sonia Jaffe, Nicole Immorlica, and Christopher T. Stanton, “Shifting Work Patterns with Generative AI,” NBER Working Paper 33795 (2025), doi:10.3386/w33795. ↩

  7. Klarna Group plc, Annual Report on Form 20-F for the Year Ended December 31, 2025 (February 26, 2026), SEC CIK 2003292, https://s205.q4cdn.com/644747736/files/doc_financials/2025/q4/Klarna-Group-plc-20-F-2025.pdf. ↩

  8. Ryan Lopopolo, “Harness Engineering: Leveraging Codex in an Agent-First World,” OpenAI Engineering, February 11, 2026, https://openai.com/index/harness-engineering/. ↩ ↩2

  9. Ting Yan, “When ‘Do Not’ Is Not Deny: Security Rules in CLAUDE.md vs. Built-In Controls,” arXiv:2608.23550 (2026), https://arxiv.org/abs/2608.23550. The study measured matching built-in controls, not all possible enforcement mechanisms. ↩

  10. Guangzong Cai, Ruiyin Li, Peng Liang, Zengyang Li, and Mojtaba Shahin, “Rule Taxonomy and Evolution in AI IDEs: A Mining and Survey Study,” arXiv:2606.12231 (2026), https://arxiv.org/abs/2606.12231. ↩

  11. Thibaud Gloaguen, Niels Mündler, Mark Müller, Veselin Raychev, and Martin Vechev, “Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?” arXiv:2602.11988v2 (2026), https://arxiv.org/abs/2602.11988. ↩

  12. John Bowers, Graham Button, and Wes Sharrock, “Workflow From Within and Without: Technology and Cooperative Work on the Print Industry Shopfloor,” in Proceedings of ECSCW ’95 (1995), doi:10.1007/978-94-011-0349-7_4. ↩

  13. Ruben Opdebeeck et al., “An Empirical Study of Policy as Code: Adoption, Purpose, and Maintenance,” Proceedings of the 23rd International Conference on Mining Software Repositories (2026), doi:10.1145/3793302.3793355. ↩

  14. Bonnie Xu, Aravind Suresh, and Emma Tang, “Inside OpenAI’s In-House Data Agent,” OpenAI Engineering, January 29, 2026, https://openai.com/index/inside-our-in-house-data-agent/. ↩

  15. GitHub, “Enterprise AI Controls & Agent Control Plane Now Generally Available,” GitHub Changelog, February 26, 2026, https://github.blog/changelog/2026-02-26-enterprise-ai-controls-agent-control-plane-now-generally-available/. Enterprise MCP allowlists were still identified as public preview in the underlying record. ↩

  16. GitLab, “The Importance of a Handbook-First Approach to Communication,” GitLab Handbook, accessed September 25, 2026, https://handbook.gitlab.com/handbook/company/culture/all-remote/handbook-first/. ↩

  17. Anthropic, “How We Contain Claude Across Products,” Anthropic Engineering, May 25, 2026, https://www.anthropic.com/engineering/how-we-contain-claude. ↩

  18. Yilin Geng et al., “Control Illusion: The Failure of Instruction Hierarchies in Large Language Models,” Proceedings of the AAAI Conference on Artificial Intelligence 40, no. 36 (2026), doi:10.1609/aaai.v40i36.40339; arXiv:2502.15851, https://arxiv.org/abs/2502.15851. ↩ ↩2

  19. Stavros Zouridis, Marlies van Eck, and Mark Bovens, “Automated Discretion,” Leiden University eLaw Working Paper 2019/06 (2019), https://www.universiteitleiden.nl/binaries/content/assets/rechtsgeleerdheid/instituut-voor-metajuridica/elaw-working-paper-series/wp.2019.006._zouridis-van-eck-and-bovens.pdf. ↩ ↩2 ↩3

  20. Katherine C. Kellogg, Melissa A. Valentine, and Angèle Christin, “Algorithms at Work: The New Contested Terrain of Control,” Academy of Management Annals 14, no. 1 (2020): 366–410, doi:10.5465/annals.2018.0174. ↩

  21. Yiwei Wang, Chuan Zhu, Tianjun Feng, Lauren Xiaoyuan Lu, and Bingxin Jia, “Agentic AI and Human-in-the-Loop Interventions: Field Experimental Evidence from Alibaba’s Customer Service Operations,” arXiv:2605.14830 (2026), https://arxiv.org/abs/2605.14830. The evidence base consumed this study at abstract or partial-text level. ↩

  22. Stephanie Rosenthal and Shamsi Iqbal, “Beyond the Org Chart: AI and the Transformation of Invisible Work,” arXiv:2605.22707 (2026), https://arxiv.org/abs/2605.22707. The study comprises 24 interviews at one company and has no pre-adoption baseline. ↩

  23. Michael Cox, Gwen Arnold, and Sergio Villamayor Tomás, “A Review of Design Principles for Community-Based Natural Resource Management,” Ecology and Society 15, no. 4 (2010): 38, doi:10.5751/ES-03704-150438. ↩

  24. The Debian Project, “Debian Constitution,” version 1.8, https://www.debian.org/devel/constitution. This is an operative governance document, not an outcome evaluation. ↩ ↩2

  25. Rainer Feichtinger, Robin Fritsch, Yann Vonlanthen, and Roger Wattenhofer, “The Hidden Shortcomings of (D)AOs—An Empirical Study of On-Chain Governance,” arXiv:2302.12125v2 (2023), https://arxiv.org/abs/2302.12125. ↩

  26. Object Management Group, Business Process Model and Notation, Version 2.0 (2011), https://www.omg.org/spec/BPMN/2.0/PDF; Camunda, “Versioning Process Definitions,” https://docs.camunda.io/docs/components/best-practices/operations/versioning-process-definitions/, “Process Instance Migration,” https://docs.camunda.io/docs/components/concepts/process-instance-migration/, and “Audit Log,” https://docs.camunda.io/docs/components/audit-log/overview/; Open Policy Agent, “Management APIs and Architecture,” https://www.openpolicyagent.org/docs/management-introduction, and “Bundles,” https://www.openpolicyagent.org/docs/management-bundles; Google, “Canarying Releases,” The Site Reliability Workbook, https://sre.google/workbook/canarying-releases/, and Dinah McNutt, “Release Engineering,” Site Reliability Engineering, https://sre.google/sre-book/release-engineering/. These sources document capabilities, not comparative organizational outcomes. ↩

  27. Nicole Forsgren, Dustin Smith, Jez Humble, and Jessie Frazelle, Accelerate: State of DevOps 2019, DORA/Google Cloud, https://dora.dev/research/2019/dora-report/2019-dora-accelerate-state-of-devops-report.pdf. ↩

  28. OpenAI, Model Spec, August 18, 2026, https://model-spec.openai.com/2026-08-18.html; Anthropic, Claude’s Constitution, January 2026, https://www.anthropic.com/constitution. ↩

  29. Joel Becker, Nate Rush, Beth Barnes, and David Rein, “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity,” METR, 2025, arXiv:2507.09089, https://arxiv.org/abs/2507.09089; Joel Becker et al., “We Are Changing Our Developer Productivity Experiment Design,” METR, February 24, 2026, https://metr.org/blog/2026-02-24-uplift-update/. ↩