Explainer

What is an Enterprise Agent Harness

An enterprise agent harness is the outer runtime for operators: wiki, scoped connectors, agent teams, write gates, and a ledger — not a SWE-bench score and not a chat with every production login.

An enterprise agent harness is the outer runtime that lets a model work on company jobs: policy it actually loads, connectors with least privilege, a loop that can stop for a named signer, and a record you can query after the people change.

It is still Agent = Model + Harness. The workspace is not a git root. The sensor is not only pytest. The stop is not only max steps. Thoughtworks calls the missing piece an organisational harness: identity, ownership, economics, and learning around whatever builder harnesses (Claude Code, Cursor, LangChain graphs) teams already bought. An enterprise agent harness is that layer made operable — whether you assemble it or hire it.

Inner vs outer is the cut. This page is the outer object in full. An enterprise AI operating system is the product category that usually ships it: wiki, workstreams, teams, gates, ledger. You can have OS-class products (Nimbus, Palantir AIP, Salesforce Agentforce) and still fail the harness test if writes are a boolean on an API key. You can assemble an enterprise harness in LangGraph and pass the test. The noun is the runtime properties, not the logo.

Words you’ll hear

  • Outer harness. Company workspace. See inner vs outer.
  • Organizational harness. Thoughtworks’ fourth layer after model, builder harness, and user harness. Governance architecture, not another markdown file.
  • Workstream. Isolation domain: roster, connectors, budget, finish line. The job folder. Not a chat title.
  • Write quoting. The human sees the change in the language of the live system before sign-off. Write-back governance.
  • Fail-closed. Missing approval, detached grant, or down interceptor means nothing mutates. Fail-open is a faster incident.
  • Ledger / Lifecycle Graph. AI operations events: brief, agents, policy version, signer, payload. Distinct from the warehouse’s business events. See What is a lifecycle graph.
  • SWE-bench / Terminal-Bench. Inner evals. Useful for engineering vendors. Not a SOX control. Eval loops.
  • Forward-deployed programme. Vendor engineers for months. AIP at scale. Capability can be real. Time-to-value is staffing. Self-service vs forward-deployed.

Nimbus is one self-service enterprise harness: wiki, workstreams, agent teams, governance, graph, routing. Score it as an example of the shape, next to AIP and Agentforce, not as the definition of the category.

Why you should care

McKinsey’s 2025 State of AI keeps separating use from scale. Copilots and coding harnesses can produce the first. Enterprise harnesses are how writes to systems of record become the second without becoming shadow AI in the CRM.

It affects you if:

  • RevOps, Legal, and Finance must share a job, not a Slack channel of screenshots
  • Salesforce or NetSuite can change because a model proposed it
  • last quarter’s pricing chat is unrecoverable
  • security cannot list the AI actors that may write
  • the vendor demo is a SWE-bench plot and a “we have MCP”

In 2024 Air Canada was held to a chatbot’s invented policy (CBC). That is an outer-harness failure: a commitment left the building without a quote or a signer. GDPR constrains personal data in payloads. Sarbanes–Oxley constrains who may change revenue truth. EU AI Act Article 14 wants people who can interpret, interrupt, and leave a record. A coding-agent hook that formats Python does not satisfy those.

NIST AI RMF and ISO/IEC 42001 assume operational controls, not a slide titled governance. OECD AI Principles are a board checklist. They do not implement a gate. The harness does.

What “enterprise” adds to a harness

Start from what a harness contains — loop, tools, memory, permissions, feedback, orchestration — and raise the bar.

Policy that loads. Inner harnesses inject AGENTS.md. Enterprise harnesses inject asserted company policy for this job, versioned. A Drive dump is not policy. A wiki that agents cite, with the revision on the run, is. If Legal’s discount cap lives only in a PDF nobody attached, the model will invent a number. That is not hallucination as a personality. That is a missing guide.

Connectors as grants, not a toolbox. Default read. Write is a separate plane. Least privilege is a workstream property. Connector architecture. MCP may be the plug; it must inherit the grant. MCP for enterprise integrations. A Finance team assigned to a GTM-only stream still must not reach ERP “because it is Finance.” Agent team architecture.

A hiring object for operators. Not a folder of personal GPTs. A mandate, required systems, approval triggers — agent teams as a roster. How to evaluate agent teams vs single agents. Nimbus ships functional teams on that roster; AIP and Agentforce have their own packaging. The test is: can an operator inspect the mandate and the required systems before assign.

Human wait as a step. HITL is not a kill switch in a dashboard. It is quoted payload, named role, fail-closed adapter. HITL approval architecture. Soft / Hard / Critical matched to blast radius. A six-month zero-reject rate on CRM writes is a finding.

A ledger of AI operations. Who briefed, which team, which wiki revision, who signed, what executed. Exportable without the vendor in the room. The warehouse is not this ledger. How to evaluate AI audit and observability.

Evals that match the job. Did the executed write match the signed quote. Can you replay. Inner leaderboards are a vendor quality signal for coding. They are not the enterprise eval. See eval loops.

Economics of the loop. Routing compact extract vs frontier judgement. Spend quotes. Seat pricing that includes unlimited flagship is an unengineered cost harness. Model routing. Nimbus meters NTUs; copilots meter seats. Different jobs.

Self-service vs programme. If every new connector is a six-month SOW, you have bought a deployment, not a harness operators can tighten. That can still be the right buy for Ontology-scale complexity. It is the wrong buy for a standard Salesforce write this quarter.

What it is not

A coding harness with SSO. Inner vs outer.

A copilot with an admin console. Copilot vs work OS.

A framework. LangGraph can host an enterprise harness if you build grants, quotes, and a ledger. Out of the box it hosts a graph. Harness vs framework.

“We integrate with Salesforce.” Integration is a slide. A scoped connector plus a blocked unsigned write is a harness.

SWE-bench-first marketing. Anthropic and LangChain are writing about coding and general agents. Steal the discipline (stops, artifacts, sensors). Do not steal the benchmark as your control framework.

Thoughtworks’ organisational harness, in operator language

The Thoughtworks OS essay (10 July 2026) argues that most AI programmes fail because the organisation never built the operating system around the model: accountability, ownership, measurement, learning. They name four layers. An enterprise agent harness, as this article uses the term, is layers 3–4 made runnable for company jobs — not only for coding-agent users.

Delegation failures are the tell. The model was fine. The platform ran. Practitioner guides existed. The agent did what it was allowed to do. The company still took harm. Layer 4 questions: who approved that autonomy, who owns the policy, what was the escalation, how do we prevent the same miss on another team. Layers 1–3 cannot answer those. A chat product cannot either.

Thoughtworks’ control matrix is worth stealing even if you never hire them. Use deterministic controls where the boundary is knowable: allowed actions, residency, spend ceilings, blast-radius limits. Use probabilistic controls only where judgement is required. Pair every guide with a sensor. Temporal constraints — consistency across a multi-step workflow, not a single dropdown — are the ones they say teams miss most. A scheduling agent that is locally plausible on each step and globally inconsistent is not a “hallucination.” It is a missing temporal sensor.

Their public examples (Parloa’s repo-resident rules/skills/commands; Morgan Stanley’s tiered autonomy on CVE triage) are coding-adjacent. Translate them: discount policy as a versioned wiki skill; “what delegation tier does this CRM write require?” instead of “do we trust the agent.” Nimbus’s Soft / Hard / Critical is that tiering in product form. AIP will have a different packaging. The architectural claim is the same.

Databricks and Wikipedia describe the runtime. Thoughtworks describe why a runtime without ownership still fails at scale. You need both descriptions when you buy.

What “good” looks like on a live job

A renewal write: workstream isolation; Salesforce attached read-only until write is enabled; wiki revision with the cap cited on the run; team cannot start if Legal’s connector requirement is missing; model proposes a quote; Hard gate; reject leaves Stage unchanged; export shows signer without a vendor screen-share. That is an enterprise harness. A demo that only answers “what should we do about Acme” is a copilot with a logo.

Spend an hour asking where each Thoughtworks layer lives in the vendor’s product. If layer 4 is “our professional services team,” you are buying a programme. That can be the right buy. Name it. Self-service vs FDE.

Operators already know the human version of this harness. Maker-checker on journals. Segregation of duties on payments. Change-advisory on production. The enterprise agent harness is those instincts encoded so a model cannot talk through them. Sarbanes–Oxley did not wait for LLMs; it waited for a named signer. GDPR did not wait for MCP; it waits for purpose limitation on the payload. If your AI programme cannot point to the interceptor that enforces those, you have a chatbot with a risk register.

What failure looks like in the first ninety days: every department clones a GPT with the same Salesforce key; Legal’s cap lives in a slide; the only eval is “the demo was impressive”; coding-agent MCP is pointed at production “just for a spike”; the ledger is Slack. What success looks like: one roster of teams, workstream isolation, default read, a Hard refuse on the first PoV, a wiki revision on the graph, inner harnesses still compiling in repos. Nimbus is built to make the success path a product week rather than a services year. Verify that claim with the refuse. AIP may be the right path when Ontology-scale complexity is real — then the harness is a programme, and you should staff it as one.

How this shows up in Nimbus

Nimbus’s outer loop is: brief a workstream → assign a team whose connector contract is satisfied → retrieve under scope → draft on the canvas (Conflux) → quote writes → governance pause → execute the signed payload → commit to the Lifecycle Graph. Perception orients; it does not silently write. Routing picks model class per step.

That mapping is how we productised harness engineering for operators. It is not a claim that AIP or Agentforce are “not harnesses.” They are different time and scope. How to evaluate an enterprise AI OS and how to evaluate an agent harness are the two sheets; use both.

Questions people actually ask

Do we need this if we already have Claude Code?

You need it for jobs whose workspace is the company. Keep Claude Code for repos. Do not share production SoR write tokens into the inner harness.

Is Palantir AIP an enterprise harness?

It can be, as a programme-shaped outer runtime. Ask deployment time, who sets a gate without vendor engineers, and whether the ledger is yours. Category yes; evaluation still required.

Is Agentforce enough?

If the job is CRM-anchored and stays there, maybe. Cross-system jobs with Legal on the canvas usually need a harness that is not only Salesforce. Clear scopes; avoid two writers.

Can we build this on LangChain?

Yes, with time. You will rebuild grants, quoting, roster, and replay. Build vs buy. Frameworks assemble loops; operators still need a loop they can hire.

What’s the first proof?

A real cross-department write: operator attaches OAuth; unsigned payload blocked; reject leaves SoR unchanged; export shows signer. Proof of value. A chat demo is not this.

How to evaluate an agent harness. Agent harness architecture. What is harness engineering.

What is write-back governance and RFP questions for enterprise AI agents.

Sources

See what governed AI looks like on your stack.

Connect your tools, run a workstream, and keep every decision on your ledger - free for 7 days.