Thought Leadership

Why Enterprise AI Needs an Outer Harness

In Q3 2024, a sub-agent proposed $4.2M in journal entries with no human review. The model was fine. The architecture was not. Here is how the outer harness prevents that.

In Q3 2024, a Fortune 500 financial services firm discovered that an autonomous AI agent had proposed journal entries totaling $4.2M in revenue adjustments — with no human review. The agent had inherited a controller's credentials, full write access to the general ledger, and an instruction to "reconcile revenue recognition against active contract amendments."

The model did what it was asked. It read billing tables, matched amended contract milestones, calculated the adjustment, and staged the entry. When auditors asked who approved the change, the answer was a system log, not a name. The firm had deployed the agent inside what researchers call the "inner harness" — the model's native reasoning loop, tool-calling manifest, and unstructured context window — without wrapping it in the "outer harness" that would have enforced read-only defaults, required cryptographic human sign-off on material writes, and recorded the decision so it could be replayed.

That is not a model failure. It is an architecture failure.

Frontier large language models score well on isolated benchmarks. They achieve over 70 percent success rates on synthetic code-generation tasks like SWE-Bench Verified. But when confronted with long-horizon execution across evolving enterprise systems — where "the environment" is the company's live financial stack, not a stable code repository — task resolution rates collapse to under 26 percent on SWE-Bench Pro and 21 percent on SWE-EVO. The model's reasoning is intact. What fails is the operational envelope: which systems the model may touch, which changes require approval, and whether the decision is recorded in a form that survives the chat session.

This essay describes the outer harness — the deterministic control layer enterprises must build around autonomous agents — and the bi-temporal lifecycle graph that makes AI-mediated decisions auditable, so the question "who approved this" has a documentary answer, not a log file you cannot parse.

Inner harness versus outer harness: where control lives

Most agent frameworks stop at the inner harness. The inner harness is what the model vendor provides: the reasoning loop that inspects an input, picks a tool from a list, calls that tool, reads the result, and repeats until the model decides it is done. That loop is where the raw capability lives — the pattern-matching, the next-token prediction, the fluency. It is also where enterprises have the least control.

The outer harness is what you build around that loop. It is the deterministic control layer that decides which tools the agent may call, which outputs require human approval before they commit, and what gets recorded so you can reconstruct the decision later. The inner harness is probabilistic. The outer harness is not. It enforces rules the model cannot override: read-only by default, writes blocked until a named person signs, and every proposal logged whether it was approved or refused.

When enterprises deploy agents without an outer harness — letting the model's inner loop talk directly to CRM, ERP, or financial systems — the failure mode is predictable. The agent does what it was instructed, but nobody enforced the constraint that should have been architectural, not conversational. "Ask before you write" is a conversational instruction. "Writes are blocked at the API layer until a cryptographic approval token is present" is an architectural constraint. Only the second one survives when the model misunderstands, or when an attacker tries prompt injection.

The outer harness has three layers:

Feedforward guides: shape the environment before the model reasons

Feedforward controls prevent errors before the model generates a single token. They work by constraining what the agent can see and what tools it can call, based on the job it is working on and the role of the person who invoked it.

Workspace schema. The structured definition of what this job requires: which data sources are in scope, which output format is expected, which approval rules apply. The schema is deterministic. It does not depend on the model interpreting a prompt correctly. It is the environment the model inherits when the session starts.

Standard operating procedures (SOPs). The authoritative wiki, policy doc, or contract clause the agent must follow. Retrieved from the source of truth — not from a chat memory — and injected into context as immutable ground truth. If the model tries to cite a stale policy, the feedforward guide has already constrained the retrieval to return only active versions.

Role-based tool permissions. Not every agent session should have write access. The permissions the agent inherits depend on the role of the person who started the job. A finance analyst's session can read billing tables but cannot post to the general ledger. A controller's session can propose journal entries, but those entries still sit in escrow until the controller signs them. The model does not decide its own permissions. The outer harness does.

Feedback sensors: validate outputs before they commit

Feedback sensors run after the model produces an output but before that output is allowed to leave the enterprise or change a live system. They are deterministic checks, not subjective judgements.

Abstract Syntax Tree (AST) parsers. If the agent generated code, the AST parser verifies the code is syntactically valid before it can be deployed. If the agent generated a JSON payload for an API call, the schema validator checks it against the expected structure.

Type checkers and linters. Static analysis tools that catch errors the model might have introduced — undeclared variables, type mismatches, security anti-patterns. These run automatically, without human review, and block the output if they fail.

Policy assertion checks. If the agent proposed a discount, the policy check verifies the discount is within approved limits. If the agent proposed a contract clause, the policy check confirms the clause does not waive a right the legal team has declared non-negotiable. These are rules you can test in isolation. They do not depend on the model "understanding" the policy. They enforce it.

Database constraint verifiers. If the agent is proposing a write to a database, the constraint verifier checks foreign-key relationships, uniqueness constraints, and required fields before the write is allowed. A model that hallucinates a customer ID will be caught here, not after the write corrupts the production table.

The steering loop: convert failures into permanent fixes

When a feedback sensor catches an error, the outer harness does not simply tell the model "try again." It routes the failure back as structured telemetry — which rule failed, which value was out of bounds — and, if the same class of error appears repeatedly, it converts the failure into a permanent guard: a tighter schema, a new pre-flight check, an additional approval gate.

The core rule of harness engineering, as systems architects have been teaching for years, is this: every operational failure must produce a structural change in the harness, not an ephemeral tweak to a conversational prompt. If the agent keeps proposing journal entries that violate the company's capitalisation threshold, the fix is not "please remember the threshold is $5,000." The fix is a constraint in the outer harness that automatically rejects any journal-entry payload where the debit exceeds $5,000 and no controller signature is attached.

That is the difference between hoping the model behaves and enforcing the behaviour architecturally.

Mechanics of the enterprise write-gate and privileged non-human identities

The critical threat surface of enterprise generative AI is not the generation of inaccurate text, but the unauthorized execution of state changes across systems of record. When a sub-agent transitions from read-only contextual retrieval to initiating database writes, executing automated clearing house (ACH) transfers, deploying compiled binaries, or transmitting external communications, it crosses from an informational tool to an entitled system user.

IBM’s 2025 Cost of a Data Breach report established that 13% of surveyed organizations experienced breaches directly involving an AI model or application. Within those compromised organizations, an overwhelming 97% lacked adequate AI access controls. The attack vectors leading to these incidents were dominated by supply chain vulnerabilities—specifically over-privileged Application Programming Interface (API) connectors, misconfigured application plugins, and unvalidated third-party tool bindings. The same research documented that unauthorized shadow AI featured in 20% of enterprise security incidents, elevating average data breach damages by approximately $670,000 due to extensive exfiltration of intellectual property and personally identifiable information (PII).

To eliminate this vulnerability, the outer harness must enforce an architectural invariant: read-only is the non-negotiable default state of the enterprise AI substrate. An assistant permitted to read records, inspect invoices, and draft reconciliation proposals cannot directly breach the financial integrity of a business. A model wired directly to a payments or customer database via a write-enabled service token represents an unmonitored privileged identity that operates without fatigue, without human hesitation, and without performance reviews. The OWASP Top 10 for Large Language Model Applications identifies indirect prompt injection (LLM01) and excessive agency (LLM06) as primary enterprise vulnerabilities. If an external, untrusted input—such as an inbound vendor invoice containing hidden instructions—can manipulate a sub-agent’s internal reasoning, an over-privileged connector will execute that instruction with the full authorization of the underlying service token.

Enforcing the write-gate requires implementing a rigorous non-human identity governance framework:

  1. Non-human user joiner-mover-leaver (JML) lifecycles. Sub-agents must never run under shared administrative credentials or static developer tokens created to expedite a pilot. Every sub-agent instance operating within an active workstream must receive an ephemeral, cryptographically distinct identity tied to a named operational owner, bound by a deterministic expiration timestamp, and restricted to fine-grained scopes.
  2. Payload staging and multi-tier approval gates. When a sub-agent issues a write-call, the outer harness intercepts the command at the network layer. The payload is placed into transactional escrow, and an execution gate evaluates the blast radius. Modifications below pre-defined risk and spend ceilings that satisfy deterministic schema validation may clear automatically, while actions exceeding financial or operational thresholds require explicit, cryptographically signed human authorization before the packet is dispatched to the enterprise system of record.
  3. Mandatory refusal logging. If an operator denies a sub-agent's proposed state change, or if a feedback sensor flags an unauthorized tool invocation, the outer harness records the refusal payload, the associated context state, and the rejecting authority. A security and audit architecture that logs only successful API transactions is functionally blind; an absence of recorded refusals demonstrates to internal auditors that safety controls are bypassed rather than actively filtering risk.

The legal and financial necessity of this write-gate architecture is demonstrated in established administrative case law. In Moffatt v. Air Canada (2024 BCCRT 149), the Civil Resolution Tribunal rejected the airline's defense that its automated customer-facing chatbot constituted an independent legal entity responsible for its own misstatements regarding bereavement fares. The tribunal ruled that an organization maintains absolute vicarious liability for the representations, automated concessions, and commitments made by its digital agents, regardless of whether the output was an unmonitored hallucination or an approved transmission. When a sub-agent's write action impacts a consumer or an official ledger, the enterprise cannot disclaim the outcome; it must produce the documentary lineage proving how the action was validated, approved, and released.

Architectural dimensionBare model / consumer shadow AIInner-loop framework (e.g., raw LangChain/AutoGen)Enterprise outer harness (governed operating envelope)
Execution boundaryUnmanaged third-party multi-tenant cloud.Local container or virtual machine; direct API dispatch.Isolated, tenant-confined workspace; network-sandboxed.
Identity and privilegeAnonymous user or personal consumer account.Shared developer API token; broad admin rights.Ephemeral non-human user credentials; strict JML lifecycle.
Tool execution policyUnrestricted natural-language browser output.Unconstrained autonomous function-calling loops.Read-only default; staged payload escrow on writes.
Verification and evalsNone; implicit trust in generated text.Subjective model-as-judge prompt wrappers.Deterministic AST parsers, linters, and schema verifiers.
Audit and state lineageEphemeral browser session; unlogged.Flat execution trace logs; transient memory.Cryptographically signed, bi-temporal Lifecycle Graph.
Failure recovery modeManual human re-prompting on terminal error.Indefinite programmatic retry loops; token exhaustion.Feedback sensor routing to steering loop; human intervention.

Why vector search breaks on enterprise policies: the temporal blindness problem

Enterprise AI suffers from a memory problem. Employees use disconnected tools — one chatbot for summarising, another for drafting, a third for financial analysis — and spend hours copying outputs between systems. That fragmentation creates the illusion of speed at the individual level while organisational throughput stalls. Microsoft and LinkedIn's research found that 75 percent of knowledge workers use generative AI, and 78 percent bring their own tools. The official path is too slow or too restrictive, so people route work through personal accounts on phones.

Banning those tools does not solve the problem. After Samsung engineers leaked proprietary code to ChatGPT in 2023, enterprises issued blanket bans. Employees routed around them. The organisation lost visibility into where intellectual property was going, but the leakage continued. A ban without a usable sanctioned substitute just moves the work underground.

To fix fragmentation, enterprise architecture teams deploy Retrieval-Augmented Generation (RAG) — systems that retrieve relevant documents from a vector database and inject them into the model's context. That works well for static corpora. It breaks on evolving enterprise policies because vector search has a structural flaw: temporal blindness.

Vector databases calculate semantic similarity — how close two pieces of text are in meaning — but they do not understand time. If your company has a 2022 travel policy, a 2024 draft revision, and a 2026 active mandate, the vector database sees all three as equally relevant if they use similar language. When a customer-service agent asks "what is our cancellation policy," the retrieval engine might return the 2022 version because its phrasing happens to match the query slightly better than the 2026 text.

The model, given three conflicting policies, synthesizes an answer. That answer might promise a refund the 2026 policy does not allow. The company is liable. The customer has a screenshot. The vector database did not know which document was "the truth."

Graph-based systems like GraphRAG improve on flat vector search by clustering related documents and pre-generating summaries. But pre-summarized graphs still require batch re-indexing when policies change. If a policy changes at 9:00 AM, agents running at 9:01 AM need to see the new rule. Auditors investigating a 8:59 AM transaction need to see the old rule. Static graphs cannot answer both questions because they overwrite history when they update.

What enterprises need is a memory system that tracks what was true when, not just what is true now.

The lifecycle graph: memory that tracks what was true when

To solve the temporal blindness problem, enterprises need a memory system that never deletes history — it just marks when each fact stopped being true and when the next version started. That is bi-temporal modeling, a technique from database theory (Richard Snodgrass, ISO SQL:2011) that tracks two independent timelines:

1. Valid time — when it was true in the real world. A travel policy is valid from January 1, 2026 to December 31, 2026. A customer contract amendment is valid retroactive to September 1. Valid time answers the business question: "what rule governed this transaction on this date?"

2. Transaction time — when the system learned about it. The policy was entered into the database on December 15, 2025. The amendment was signed on September 28 but recorded on October 2. Transaction time answers the audit question: "what did the system know when it made this decision?"

With both timelines, you can answer two critical questions that vector search cannot:

  • Current state (for operations): What is the active policy right now? Query for records where valid time includes today and transaction time is still open.
  • Point-in-time reconstruction (for audits): What policy was active on September 15, and what did the system know about it on that date? Query for records where valid time included September 15 and transaction time included September 15.

The lifecycle graph never overwrites. When a policy changes, the old version is closed (valid-time-end = policy change date, transaction-time-end = system commit timestamp) and a new version is appended with a link showing it supersedes the old one. Nine months later, auditors can replay the exact state of the world when the agent made a decision. That is the documentary answer regulators and insurers need.

The three layers of the graph

The lifecycle graph organizes memory in three tiers:

Episodic memory: Every prompt, tool call, approval, refusal, and API response is logged with exact timestamps. This is the forensic layer. When someone asks "who approved this $4.2M journal entry," the graph has the payload, the controller's cryptographic signature, and the timestamp.

Semantic memory: Business entities, roles, projects, policies, and their relationships. Each edge carries valid-time intervals. When an agent asks "what is the capitalisation threshold," it traverses only edges where valid-time includes today. Stale policies are in the graph but filtered out.

Community memory: Higher-level patterns and interdependencies. "Revenue recognition rules interact with contract amendment workflows, which are blocked during deployment freezes." This layer helps agents understand context without re-deriving it from raw events.

The key architectural rule: old facts are never deleted, only invalidated. That discipline is what makes the system auditable. When a regulator asks what the agent saw on a given date, you can rewind the graph to that exact state. When an operator asks what the current rule is, the graph filters to active facts only.

Operationalizing multi-agent workstreams in complex enterprise workflows

To evaluate how the Outer Harness and Lifecycle Graph operate in production, consider an enterprise executing a cross-functional financial close: the Q3 Financial Close and Revenue Recognition determination governed by ASC 606 under evolving customer contract amendments.

In an unmanaged environment, finance analysts paste contract excerpts into disparate chat windows, attempt manual spreadsheet reconciliations, and copy unverified journal adjustments into the general ledger. If an amendment was signed on September 28 but retroactive to September 1, a standard RAG system pulling contract files risks citing original payment terms rather than amended revenue milestones, leading to inaccurate revenue recognition and subsequent restatements.

Within the governed operating envelope, the entire workflow is managed as a formal workstream where specialized sub-agents and operational state transitions progress through strictly enforced lifecycle gates.

StageTrigger / inputArchitectural mechanismTemporal state and graph operationResulting state and governance gate
1. Workstream initiationController executes charter for Q3 close.Outer Harness workspace orchestrator generates an ephemeral execution sandbox.Instantiates a unique workstream node on the episodic subgraph; records transaction start time.Ephemeral non-human user token issued; permissions restricted to read-only financial data.
2. Context groundingWorkspace orchestrator initiates retrieval.Enterprise wiki and semantic graph engine query active accounting standards.Traverses the semantic graph for active accounting SOPs where valid time contains the close date.Filters out deprecated 2024 revenue guidance; returns only active ASC 606 corporate rules.
3. Contract ingestion and triangulationGrounded brief assigned to legal and finance sub-agents.Legal and finance sub-agents read billing tables and contract amendments.Executes bi-temporal edge traversals to evaluate contract amendments across validity windows.Read-only connectors prevent modifications to CRM, billing records, or customer vaults.
4. Sensor validation checkFinance sub-agent drafts a journal adjustment.Outer Harness feedback sensors execute deterministic mathematical validation.Evaluates proposed debits against credits; validates the cited amendment node against the signature ledger.Sensor flag: unexecuted amendment cited; triggers the steering loop to search for the executed addendum.
5. Staged write proposalSub-agent resolves the citation and prepares the payload.Outer Harness transaction escrow intercepts the proposed general ledger entry ($1.4M).Generates a staging node on the episodic subgraph capturing the proposed general ledger adjustment.Write-gate locks execution: the proposed transaction exceeds the $100,000 automated ceiling.
6. Human-in-the-loop releaseStaging lock pages the corporate controller.Governance console presents the diff, active policy versions, and citations on the Lifecycle Graph.Graph presents the provenance chain linking ERP invoices, the executed addendum, and the SOP.Controller signs the release token with a cryptographic hardware key; stores an immutable approval record.
7. Ledger commit and state appendValidated cryptographic release token received.Scoped write connector executes the transaction against the general ledger API.Appends a new ledger-state entity; invalidates the prior balance edge; records valid time and transaction time.A single atomic write is committed; ephemeral non-human agent credentials are automatically revoked.

This multi-agent workflow demonstrates compounding intelligence. The output of the revenue recognition determination does not vanish into a chat log. It is committed to the Lifecycle Graph as an immutable, interconnected structure. When internal auditors, tax compliance authorities, or FP&A teams conduct analyses in future quarters, they do not review an ungrounded textual summary. They traverse the graph to inspect the exact brief, the active policy version, the evidence retrieved, the sensor validation traces, the controller’s signature, and the resulting ledger transition.

Regulatory reconstruction and evidentiary defensibility under active enforcement

The deployment of enterprise AI has entered an era of direct enforcement, where administrative agencies evaluate concrete operational systems rather than high-level ethical statements.

In March 2024, the U.S. Securities and Exchange Commission (SEC) announced settled charges against investment advisers Delphia (USA) Inc. and Global Predictions Inc., imposing $400,000 in total civil penalties for making false and misleading public statements regarding their deployment of artificial intelligence. The SEC’s orders established that advertising automated capabilities, algorithmically predictive models, or "AI-run" processes when internal operations rely on manual interventions or conventional automation violates basic statutory protections against misleading statements of material fact.

Similarly, the Federal Trade Commission (FTC) warned organizations that claims regarding algorithmic accuracy, automated oversight, and fairness must be substantiated by demonstrable operational evidence before publication. In the European Union, Regulation (EU) 2024/1689 (the EU AI Act) establishes statutory documentation, data governance, and continuous human oversight duties for high-risk systems, mandating that operators possess the architectural capability to interrupt, override, and reconstruct automated decisions throughout their operational lifecycles.

Organizations cannot satisfy regulatory examiners or defense standards with narrative slide decks, vendor marketing claims, or steering committee minutes. Regulatory bodies, external financial auditors, and commercial insurers require a verifiable reconstruction pack for every material automated action:

  • The exact ingestion state. The verbatim input payload, prompt context, and system instructions dispatched to the model runtime, preserved without post-hoc summarization.
  • The active policy version. The specific regulatory rule, corporate standard, or price sheet active at the moment of execution, verified via immutable bi-temporal valid-time intervals rather than current-state documentation.
  • The entitled human sign-off. The cryptographic identity of the human operator who reviewed the staged payload and approved the write-gate release, or the deterministic rules proving why an automated write cleared below a sanctioned risk ceiling.
  • The sensor verification record. The programmatic telemetry demonstrating that the sub-agent output successfully cleared deterministic AST parsers, schema validations, and security linters before commit.
  • The comprehensive refusal log. The historical ledger of actions the system actively blocked or rejected, proving that governance controls function as active physical barriers rather than passive policy advisories.

When enterprise leadership aligns systems architecture with ISO/IEC 42001 (the international management system standard for artificial intelligence) or the NIST AI Risk Management Framework (AI RMF), the Outer Harness and Lifecycle Graph supply the required structural artifacts. Compliance ceases to be an annual, retrospective paper-gathering exercise. It functions as the continuous, mathematical byproduct of how digital work is executed, audited, and preserved across the enterprise.

A call to CTOs and enterprise architects

Scaling autonomous AI is an engineering and governance problem, not a model-evaluation contest. Frontier models have impressive reasoning capability. That capability is worthless if your enterprise cannot constrain where the model points, govern which systems it can change, and reconstruct how a decision was made nine months later when the auditor asks.

Organisations that let agents operate via unstructured inner loops — the model's native tool-calling manifest and conversational memory — without wrapping them in a deterministic outer harness will generate compliance liabilities, security breaches, and untraceable financial errors. The fix is architectural:

1. Enforce read-only by default. Writes are blocked at the API layer until a named human approves the exact payload. That approval is cryptographically signed and logged. Not "the model asked first." The payload waits in escrow.

2. Build feedforward guides and feedback sensors. Shape the environment before the model reasons: which tools it can call, which policies are active, which role it inherits. Validate outputs before they commit: AST parsers, schema checkers, policy assertions. Convert repeated failures into permanent structural guards, not conversational reminders.

3. Deploy a bi-temporal lifecycle graph. Track what was true when (valid time) and when the system learned about it (transaction time). Never overwrite history; invalidate it. That is how you answer "what policy governed this transaction" (operations) and "what did the system know when it decided" (audit) without reconstructing chat logs.

The worked example in this essay — Q3 financial close, revenue recognition, contract amendments, staged writes, controller sign-off — is the operational reality of enterprise AI. The model did its job. The architecture either prevented a $4.2M error with no human review, or it allowed it. That difference is the outer harness.

Sustainable enterprise AI means every autonomous action occurs inside a deterministic envelope that enforces access, approval, and auditability. The inner harness is what the model vendor gives you. The outer harness is what you build to make it safe. Build it before the incident teaches you why it was necessary.


References

About Nimbus

Nimbus is a Collaborative AI Operating System built around four core pillars that bring human teams and autonomous AI together into a single, unified workspace.

Communication: Keep context tied to the job. Unify emails, meeting recordings, transcripts, and operational files directly within active projects—ending knowledge silos buried in private inboxes, scattered Slack threads, or unrecorded calls.

Collaboration: Work alongside AI in real time. Bring people and AI agents onto the exact same brief, visual canvas, or initiative. Query company-wide data, invite agents into live calls, and co-create in one shared space—eliminating the split between human group chats and isolated AI sidebars.

Automation: Put routine workflows on autopilot. Connect more than 2,000 enterprise tools and standardize repetitive operations. Background loops run on schedules or data triggers with full execution logs, ensuring operational knowledge is shared across the team rather than trapped in one person’s head.

Governance: Deploy AI with absolute control. Enforce strict role-based access controls across workspaces. AI agents can analyze, summarize, and draft—but no live system changes or external communications occur without explicit, verified human sign-off.

Short answers

A deterministic envelope around the agents

Why do frontier models stall on enterprise work?

They pass isolated coding benchmarks but fail on long-horizon jobs where the company — not a code repo — is the environment, and where a wrong write costs money or reputation.

What is the outer harness?

The deterministic envelope around autonomous agents: which systems they may read, which changes require approval, and what gets recorded so you can reconstruct the decision later.

What is the lifecycle graph for?

A bi-temporal record of work — what was true when the decision was made, and when the system learned about it — so auditors and operators can replay the exact state that informed an AI-proposed change.

See what governed AI looks like on your stack.

Connect your tools, run a workstream, and keep every decision on your ledger. Start on Free.