How to Evaluate an Agent Harness
Evaluating an agent harness means checking whether it can stop a write, replay who signed, swap the model without rewriting tools, and fail a real sensor — not whether the demo answered a question.
Evaluating an agent harness is checking whether the runtime around the model can finish a job under a stop you trust — not whether a demo answered a question.
LangChain defines the object: Agent = Model + Harness. The scoring sheet is therefore about the harness. If your RFP starts with context-window size and SWE-bench, you are scoring a model (and maybe an inner coding loop). You will miss whether an unsigned Salesforce PATCH is possible. How to evaluate an enterprise AI OS is the cousin sheet for wiki, workstreams, and routing as a product category. This page is the runtime tests that apply to Claude Code, a LangGraph deployment, AIP, Agentforce, and Nimbus alike — then specialised by inner vs outer.
McKinsey’s 2025 State of AI already measured the trap: widespread use, limited scale. A fluent demo produces the first. A harness that can refuse, replay, and ratchet produces the second. NIST’s AI RMF Playbook is the measurement language. ISO/IEC 42001 is the management-system language. Neither is “the model seemed careful.”
Words you’ll hear
- Harness vs framework. Library versus running loop. Harness vs framework. “We use LangChain” is not a passed test.
- Sensor. Independent check. Böckeler; Thoughtworks. Inner: tests. Outer: quote vs SoR.
- Hook / interceptor. Always runs. Claude Code hooks. Outer: fail-closed adapter.
- Quote. Structured payload, not a paragraph. Write-back.
- Replay. Can you reconstruct signer, policy version, tool grants. Lifecycle graph.
- Model portability. Swap weights without rewriting tools. Not a logo on a slide. Model routing.
When you evaluate Nimbus, run these tests on workstreams and governance, not on a homepage video. When you evaluate Claude Code, run them on a repo hook and CI, not on a blog SWE-bench screenshot. Same sheet, different workspace.
Why evaluation usually fails
People score agents like they score chat: quality of the paragraph, latency, brand of the model. That produces three false passes:
- The copilot pass. SSO, a usage dashboard, a good answer. No loop ownership. Copilot vs work OS.
- The benchmark pass. SWE-bench or Terminal-Bench for an outer job. Inner eval, outer purchase. Eval loops.
- The framework pass. A graph in a notebook with every production tool attached. OWASP excessive agency with extra nodes.
Anthropic is blunt: encode the job, bound the tools, define done. Your proof of value should force those three. Written answers without a failed action are still a slide. How to run an enterprise AI proof of value.
Red flags: chat as the entire proof; “we integrate” with no scoped grant; governance as PDF; memory as a long window; “model-agnostic” with a flagship default and seat pricing; MCP write tools that inherit a god service account; vendor database offered as the new system of record.
Checklist
1. Can it stop an action the model wants? Inner: PreToolUse denies a matched command; tests fail the merge. Outer: unsigned write does not execute; reject leaves SoR unchanged. If the only stop is max tokens, you have a fuse, not a control plane. HITL architecture.
Why this matters: Air Canada and the sanctioned ChatGPT brief are ungated generation reaching a record. Your demo must show a failed write.
2. Can you replay who signed and which harness version ran? Signer identity, wiki or AGENTS.md revision, tool grants, payload hash, model class. If the answer is Slack search or “the transcript,” you do not have a ledger. How to evaluate AI audit and observability. Nimbus’s Lifecycle Graph is one implementation; demand the export without a vendor engineer.
3. Can you swap the model without rewriting tools? Change compact vs frontier on extract vs judgement. If tools are bound to one vendor’s function-calling dialect in application code with no adapter, portability is a hope. LangChain’s model interface exists for this; product harnesses must expose it as policy, not as a rewrite.
4. Are tools grants or a belt? Least privilege per job. Missing Salesforce is a configuration error, not a hallucination. Connector architecture. MCP servers inherit the same grant. MCP for enterprise.
5. Is verification outside the generator? Inner: CI the agent cannot mark skip without a hook. Outer: schema of the quote; SoR row matches. Anthropic’s long-running harness refuses “premature victory” by forcing artefacts and tests. Steal that instinct.
6. Can an operator add a sensor without a six-month SOW? Harness engineering is a ratchet. If only vendor FDE can add a gate, you bought a programme. Fine for AIP-scale. Wrong for a standard CRM field this quarter. Self-service vs FDE.
7. Is the workspace the job you are buying? Repo vs company. How to choose coding vs enterprise. A single scoring sheet with no workspace column will buy the wrong loop.
8. Economics of the loop. Max steps, spend cap, routing. Seat “unlimited” is often always-flagship. AI cost control architecture. Ask for a per-step model breakdown on a live run.
RFP questions
- Show an action the model attempted that the harness refused. What fired?
- After a successful write (or merge), show the signer, policy version, and payload (or diff) without Slack.
- Change the model on extract this week. Which tools broke?
- Attach a connector (or repo permission) as an operator, not as SE. Time?
- Detach the grant mid-job. Does the write fail closed?
- What is the independent sensor for “done”? Who can mark skip?
- Two departments, different scopes, one job — or one god toolbox?
- Price: seats, tokens, NTUs, or a services quote? What stops flagship on classify?
Put these in the RFP, then run them in a PoV. RFP questions for enterprise AI agents overlaps; keep both. Agents without a harness test are a persona list.
Proof of value (short)
Inner job: real repo, required hook, red test the agent must fix, no production SoR token.
Outer job: real cross-department write, quoted payload, reject path, export. Nimbus should pass the same live sequence as anyone else: OAuth attach, blocked unsigned write, graph export. Overview is not the proof.
Skip any refuse/replay/swap and you evaluated a chat product, a benchmark, or a framework notebook.
Score inner and outer without mixing oracles
Run two short scripts. Do not average them into one “AI score.”
Inner script (repo). Fresh checkout of a service you own. Required hook: deny a dangerous bash pattern. Agent must add a failing test then make it pass. CI is the merge sensor. No production CRM token in the environment. Record: did the hook fire, did CI stay independent, can you show the AGENTS.md revision. SWE-bench plots from the vendor are background, not this script.
Outer script (SoR). Sandbox Salesforce or equivalent. Operator (not SE) attaches OAuth. Model proposes a write. Unsigned path must fail. Reject path must leave records unchanged. Approve path: read-back matches hash. Export signer and wiki revision. Detach the connector and retry the write — must fail closed. Proof of value is this script with two departments on the canvas.
If a vendor refuses to run the outer script because “we are a coding tool,” believe them and buy them for inner only. If a vendor refuses the inner script because “we are an OS,” believe them and do not replace Cursor. If a vendor claims both and fails one script, you have a category error in their marketing. Nimbus should pass the outer script on workstreams and governance. Claude Code should pass the inner script. How to choose.
Thoughtworks’ layer check. After the scripts, ask where layer 4 lives: who owns the policy when the agent did what it was allowed to do and harm still happened. If the answer is a steering committee with no interceptor, you evaluated theatre. ISO 42001 will not save a missing refuse.
Economics check. Pull one live run’s step list: model class per step, tokens or NTUs, which sensor fired. Always-flagship with no cap is a failed harness eval even if the paragraph was good. Cost control.
MCP check. One write-capable server. Which workspaces may use it. If the answer is “any host that can see the URL,” fail. MCP for enterprise.
Weight the eight checklist items; do not add a ninth called “brand.” Stanford AI Index is useful context for how fast coding tools moved. It is not a substitute for the outer script.
Score vendors as systems, not as essays. A beautiful anatomy post does not pass the refuse test. A messy UI that blocks the unsigned PATCH does. Watch for “evaluation theatre”: the SE runs the happy path, the fail path is “we’ll configure that in phase two,” the ledger is a screenshot of LangSmith. Phase two is where McKinsey pilots go to die.
Bring your own oracle. For inner: a test the agent did not write. For outer: a sandbox row you control. If the vendor must supply the only success criterion, you are scoring their demo fixtures. Terminal-Bench’s strength is that the environment is the grader. Copy that.
People on the bake-off: an operator who will live in the product, someone who owns the SoR, someone who can say no for Legal, an engineer who will keep the inner harness. If only the vendor and an innovation lead attend, you will buy a narrative. Nimbus, AIP, Cursor, and a LangGraph SOW should all survive that room or be narrowed to the job they actually do.
Write the pass/fail before the demo so the SE cannot redefine success live. “Blocked unsigned write” is a boolean. “Felt enterprise-ready” is not. Record the session. If they cannot fail on camera, assume they cannot fail in production. NIST Playbook language helps here: you are Measuring a control, not a vibe.
How this shows up in Nimbus
Nimbus is an enterprise / outer harness: wiki as guides, connectors as grants, teams as the hiring object, governance as the interceptor, graph as replay, models as routing. Score those surfaces against the eight tests. Do not accept “we are a harness” as a substitute for a failed write. AIP and Agentforce deserve the same eight.
Questions people actually ask
Can we score Claude Code and Nimbus on one spreadsheet?
Yes, with a workspace column. Shared rows: refuse, replay, swap, sensors, operator change, economics. Inner-only rows: tests, sandbox, PR. Outer-only rows: SoR quote, roster signer, workstream isolation.
The vendor sent a SWE-bench plot.
File it under inner quality. If you are buying CRM writes, it is not sufficient. Eval loops.
We already completed a copilot RFP.
Keep it for personal tools. This sheet is for loops that act. Different job.
Is ISO 42001 certification the eval?
It is a management-system signal. Still watch a write fail. Certification without an interceptor is paperwork.
What should I read next?
Agent harness architecture to know the parts. How to evaluate write-back governance for the outer stop in detail. What is harness engineering for the ratchet after you buy.
Related reading
How to evaluate multi-agent platforms and How to evaluate AI governance platforms.
Sources
- LangChain, Agents
- LangChain, The anatomy of an agent harness
- Böckeler, Harness engineering for coding agent users
- Thoughtworks, Harness engineering and agent feedback
- Anthropic, Building effective agents
- Anthropic, Effective harnesses for long-running agents
- Claude Code, Hooks
- McKinsey, The state of AI in 2025
- NIST AI RMF Playbook
- ISO/IEC 42001
- OWASP Top 10 for LLM applications
- CBC, Air Canada chatbot lawsuit
- Reuters, ChatGPT legal brief sanctions
- Stanford HAI, 2025 AI Index
- Model Context Protocol specification
See what governed AI looks like on your stack.
Connect your tools, run a workstream, and keep every decision on your ledger - free for 7 days.