How to Choose Between a Coding Harness and an Enterprise Harness
A coding harness runs a repository — Claude Code, Cursor, Codex. An enterprise harness runs company jobs with connectors and signers. Most organisations need both; they are not substitutes.
Choosing between a coding harness and an enterprise harness is choosing the workspace. A coding harness (Claude Code, Cursor, Codex, open shells) wraps a model for a developer and a repository. An enterprise harness wraps a model for operators and systems of record. Same equation — Agent = Model + Harness — different loop.
This is the buying companion to inner vs outer agent harness. It sits beside how to choose between a copilot and a work OS: copilots are personal assistants; coding harnesses are agentic inner loops with tools and tests; enterprise harnesses are outer loops with grants and signers. Do not collapse all three into “we need ChatGPT.”
Böckeler documents how coding-agent users add guides and sensors. Osmani tells engineers to own verify-and-release. Thoughtworks argues the organisational layer is still the gap. The purchase mistake is using one budget line for all three layers.
McKinsey’s 2025 State of AI is the organisational backdrop: usage is easy; scale is redesign. A Cursor rollout can scale pull requests. It will not, by itself, scale governed CRM writes. An OS-class rollout can scale those writes. It will annoy engineers if you force “rewrite this function” through a Critical gate.
Words you’ll hear
- Coding / inner harness. Repo workspace, sandbox,
AGENTS.md/CLAUDE.md, hooks, CI. Eval: SWE-bench, Terminal-Bench, your tests. - Enterprise / outer harness. Job workspace, connectors, roster, write quotes, ledger. Eval: signed payload vs SoR. What is an enterprise agent harness.
- Copilot. Personal completion surface. Often no repo loop. ChatGPT Enterprise, Microsoft 365 Copilot, Claude for Work. Keep for mail. Do not hand it the NetSuite token.
- Framework. How you assemble a loop in code. Not a purchase of a company workspace. Harness vs framework.
- MCP. Plug into either. Dangerous when both share a production write server. MCP for enterprise.
Nimbus is an enterprise / outer option: workstreams, teams, governance. Claude Code is a coding / inner option. The rational stack is both, with a hard rule: no unsigned SoR writes from the inner harness. How to solve unapproved CRM writes.
Why the choice is usually “both”
The tools look similar in a first meeting. Both stream tokens. Both call tools. Both have “agents” on the website. The evaluation is what happens after the answer.
Buy a coding harness when the artefact is code in a repo you already trust with CI: features, refactors, tests, developer docs, infra-as-code that merges through the same gates humans use. Anthropic’s long-running harness is this world: git, progress files, end-to-end checks.
Buy an enterprise harness when the artefact is a change to Salesforce, NetSuite, a policy commitment, or a cross-department decision that must be replayed. Write-back. HITL. NIST RMF context of use is operations, not a checkout.
Keep a copilot when the job is a paragraph in a mailbox. Do not scale it into an approval architecture.
Build on a framework when engineers own a unique loop and will maintain grants. That is a programme, not a seat.
Stanford HAI’s 2025 AI Index charts the explosion of coding-agent tooling. Procurement that only reads that chart will under-buy the outer layer. Procurement that only reads ISO 42001 will over-process inner loops and lose developers.
Decision tests
1. What is the system of record for the outcome? Git: inner. CRM/ERP/customer commitment: outer. Both: two harnesses, one write plane (the outer quotes).
2. Who is the signer? The author of the PR (inner, plus CODEOWNERS). A named RevOps/Finance/Legal role (outer). If you cannot name the role, you are not ready to buy the outer write path — buy read-only first.
3. What is the independent sensor? Pytest / tsc / CI (inner). Payload schema + SoR read-back (outer). “The model said it was fine” is neither. Eval loops.
4. What identity should the tools use? Developer sandbox and repo token (inner). Workstream-scoped OAuth (outer). A shared MCP god account fails both OWASP and SoD.
5. How will you ratchet failures? Inner: AGENTS.md + hooks + tests (harness engineering). Outer: wiki revision + gate tier + graph. If your plan is “we’ll prompt better,” you have not chosen a harness. You have chosen hope.
6. Time-to-value and staffing. Cursor can be a week for a team that already has CI. AIP can be a programme. Nimbus-style self-service claims a product week for a standard write — verify with a PoV. Self-service vs FDE.
Anti-patterns
Cursor for Salesforce. MCP connected to production. Tests on fixtures. Amount changes. No signer in the ledger. Inner loop on an outer record.
Work OS for a one-line refactor. Critical gate, three departments. Engineers route around. Outer loop on an inner job.
One mesh to rule them. IDE, chatbot, and OS all write through the same server. Two writers. Multi-agent architecture.
Benchmark shopping. Buying Agentforce because of a coding leaderboard, or buying Claude Code because of a governance white paper. Wrong evidence. How to evaluate an agent harness.
Banning inner harnesses until the OS ships. Usually slows software and does not stop paste-into-CRM. Ban the write path; allow the compile path.
Nimbus should lose the inner job on purpose. If a vendor tries to replace Claude Code for application engineering, ask for sandbox, hooks, and merge sensors — evaluate the harness — and expect to keep a coding tool anyway. If a coding-tool vendor tries to replace the OS for NetSuite journals, ask for quoted GL lines and a Finance signer.
A simple portfolio
| Job | Buy |
|---|---|
| Mail, slides, one-off Q&A | Copilot |
| Application and infra repos | Coding harness |
| Cross-department SoR writes | Enterprise harness |
| Unique simulation / exotic tools | Framework + your grants |
Most enterprises tick all four rows. Budget them separately. Share policy intent (discount cap) via wiki and via AGENTS.md where relevant; share enforcement only on the plane that can execute the write.
See Overview for how Nimbus maps to the third row, models for routing, integrations for connectors. See Claude Code / Cursor docs for the second. Do not let a single SOW blur the rows.
Procurement sequence that does not waste a quarter
Week 1 — inventory loops, not vendors. List jobs that already have a finish line. Tag each: git artefact, SoR artefact, mailbox artefact, unique research. You now have four shopping lists. McKinsey programmes that skip this step buy one platform and force every row into it.
Week 2 — freeze the write rule. Unsigned SoR writes are impossible from copilots, coding agents, frameworks, and the OS. That rule is cheaper than any bake-off. It also tells Security what to revoke this month (god MCP servers). Unapproved CRM writes.
Week 3 — inner bake-off only if you lack a coding harness. Hooks, sandbox, CI independence, model swap on the same tools. Terminal-Bench and SWE-bench as vendor quality, not as Legal’s control. Anthropic hooks vs Cursor rules vs Codex — pick for your repos.
Week 4 — outer bake-off only for SoR jobs. Run the refuse/replay script from how to evaluate an agent harness. Include Nimbus, AIP, Agentforce, or a LangGraph programme as fits the staffing model. Self-service vs FDE.
Do not hold week 3 until week 4 ships. Engineers will adopt inner tools anyway; you will only lose the chance to standardise hooks. Do not skip week 4 because week 3’s coding agent “can also call Salesforce.” That is the anti-pattern.
Budget: copilot seats (predictable, personal); coding harness seats or usage (developer count); enterprise harness by work, not by mailbox count if you care about routing. Mixing all three into one “AI budget” is how flagship models burn on classify and how CRM writes go unquoted to save a line item.
Thoughtworks’ organisational harness is the steering cadence after purchase: incidents become controls across both inner and outer. Buy tools that allow that ratchet. A coding harness that forbids custom hooks, or an OS that forbids adding a gate without FDE, will stall week 5.
Expect political arguments that are actually workspace arguments. Engineering will say the OS is slow. They are right for a one-line refactor. RevOps will say Cursor is unsafe. They are right for a production Opportunity. The CISO will say “one approved agent.” Translate: one write rule, many loops. EU AI Act oversight can be satisfied per system of use, not per brand. NIST RMF Map is the same advice.
If budget forces a single purchase this half, buy the loop that matches the highest-harm unfinished job. Ungoverned CRM writes usually outrank “we could use a better coding agent” — paste already exists; unsigned APIs are new blast radius. If the highest-harm job is shipping software and SoR writes are still human, buy the coding harness and freeze the write rule until the outer product lands. Either way, write the rule down before the PO.
Nimbus should win the outer row on self-service quoting and graph export, and should lose the inner row on purpose. If a bake-off ranks us against Claude Code on SWE-bench, the scorecard is wrong. If it ranks us against a copilot on mail quality, also wrong. Rank us against AIP and Agentforce on the refuse/replay script, and against “we’ll build LangGraph” on time-to-first-governed-write.
The copilot row still matters. People will keep ChatGPT Enterprise for drafts. That is healthy if the write path is the easy official one. Banning unofficial drafts usually fails; making unofficial writes fail-closed usually works. Shadow AI is often a write-path problem wearing a chat-policy costume.
Questions people actually ask
We already paid for GitHub Copilot.
That is often a completion copilot, not a full coding harness. You may still want Claude Code or Cursor for agentic repo work. Evaluate hooks and tests, not the seat.
Can the enterprise harness include a coding specialist?
Yes, as a bounded tool that opens a draft PR. The SoR write still quotes in the outer harness. Specialists are hands. Agent teams.
What if Legal wants one vendor?
One vendor for identity and logging is reasonable. One vendor for repo loop and CRM loop is how you get a mediocre both. Prefer two harnesses and one interceptor rule: unsigned SoR writes are impossible everywhere.
How do we score Nimbus vs Claude Code in a bake-off?
Different jobs. Run inner tests on a repo. Run outer tests on a quoted CRM write. A combined “winner” is a category error unless you only have one job.
What should I read next?
Inner vs outer for architecture. How to evaluate an agent harness for the live tests. What is an enterprise agent harness for the outer object.
Related reading
How to choose between a copilot and a work OS and Build vs buy an enterprise AI OS.
Sources
- LangChain, Agents
- Böckeler, Harness engineering for coding agent users
- Addy Osmani, Own the outer loop
- Thoughtworks, The operating system for enterprise AI
- Anthropic, Effective harnesses for long-running agents
- Anthropic, Claude for Work
- OpenAI, ChatGPT Enterprise
- Microsoft 365 Copilot
- SWE-bench
- Terminal-Bench (arXiv:2601.11868)
- McKinsey, The state of AI in 2025
- Stanford HAI, 2025 AI Index
- NIST AI RMF
- ISO/IEC 42001
- OWASP Top 10 for LLM applications
- Model Context Protocol specification
See what governed AI looks like on your stack.
Connect your tools, run a workstream, and keep every decision on your ledger - free for 7 days.