Explainer

Agents Should Be Disposable

The project survives the agent. Swap model, vendor, or harness — the job still holds the files, the decision, and the signer. A guide to continuity of the work, not a buying test.

Agents should be disposable because the project survives the agent. Microsoft and LinkedIn’s 2024 Work Trend Index, from a survey of 31,000 knowledge workers across 31 markets, found that 78% of people who use AI at work bring their own tools. Personal threads are convenient. They are also how reasoning leaves the company when a person changes seats, a vendor changes models, or a licence expires.

That finding is not an argument against assistants. It is an argument against treating the assistant as the place the work lives. When the model version rolls, the vendor changes, or the person who ran last week’s session is on leave, the named job should still hold the brief, the files, the stored rejection, and the name on the signed change. The agent is staff on a shift. The job is the unit that persists.

McKinsey’s State of AI (2025) measured the same gap from the other side: 88% of organisations now use AI in at least one function, yet most remain in experiment or pilot, and only about a third have begun to scale. Scale is not “more seats on the same chatbot.” Scale is whether a second shift, a second department, or an auditor can reopen the same artefact. If the artefact is a thread that expired with a seat, you did not scale. You multiplied private windows.

This page is a design claim about continuity of the work. It is not the same question as “can you swap the model without rewriting tools?” in how to evaluate an agent harness. That test is a buying sheet for the runtime. Disposable agents are an operator claim: finance can reopen the order hold on Monday even if nobody has the same Claude thread open.

Multiplayer AI vs multi-agent AI puts the social fact plainly: everyone leaves the room eventually. Disposable agents mean the room does not empty when they do. What is collaborative AI is the shared-job definition. This page is why the job must not be buried inside a persona.

What “disposable agent” means

It means you can remove the current model, harness, or vendor without losing:

  • The named job and its finish line.
  • The files the teams already attached — not copies in five inboxes.
  • The decision chain: who proposed, who rejected, who signed.
  • The authority model: the agent never exceeded the person who started the run.

It does not mean agents are worthless, or that you should churn vendors weekly. Stable vendors and stable models are fine. Disposable means the company’s memory of the job is not identical to the agent’s session. You should be glad the agent can leave. You should be worried if the project notices.

Anthropic’s guidance on building effective agents keeps returning to the same instinct: encode the job, bound the tools, define done. Encoding the job inside a model window is the opposite of that instinct. The window is labour. The job is the object.

The everyday version of burying the job in a thread is the habit search is not memory names. The company-scale version is institutional memory in enterprise AI. Disposable agents sit between those two: the job object is where decision memory lives while the work is open, so the next operator is not rebuilding from Slack.

Why operators should care

The failure mode is familiar. A fluent run finishes. Someone screenshots the answer. Three weeks later nobody can say what was approved.

That is not a model-quality problem. It is a container problem. The screenshot is a souvenir. The job needed a brief, a payload, and a name.

Operators feel this in four places:

Handover. Night shift asks day shift why the hold is still there. Day shift asks the person who “ran it in ChatGPT.” That person is off. The hold ages.

Audit. A later reader asks who authorised the CRM change. The trail is a personal thread, a forwarded email, and a channel that has since been renamed. NIST’s AI Risk Management Framework asks for measurable governance — Map, Measure, Manage. You cannot measure a window that no longer exists.

Vendor change. Prices move. A model is deprecated. A security review kills a consumer tool. The work that lived in that tool does not migrate. Teams reconstruct from memory and call it a migration.

Role change. The analyst who “knew the customer” leaves. Their assistant’s memory was their seat. Conway’s 1968 paper noted that organisations design systems that copy their communication structure. If the structure is one person and one chat, the AI copies that — including the bus factor of one.

Stanford HAI’s AI Index tracks rising capability and falling inference cost. Capability is not continuity. Cheaper tokens make it easier to re-ask the same question. They do not create a signer. Disposable design is how you stop paying to reconstruct last week.

How this differs from “swap the model” in evaluation

How to evaluate an agent harness asks vendors to change weights mid-proof and show which tools broke. That is portability of the runtime — adapters, grants, routing policy.

Disposable agents ask an operator question: if I swap the runtime tomorrow, does the order hold still have finance’s rejection on file?

Both matter. Conflating them is how teams pass an RFP and still lose Monday’s job. A harness can be portable and still store every decision in a session the next model cannot read. A job can be durable and still sit on a brittle, single-vendor runtime. Score both.

QuestionHarness evaluationDisposable agent
Unit of concernTools, sensors, stopsFiles, signer, finish line
ProofFailed write, replay exportReopen job without original thread
FailureCRM PATCH with no gate“Search Slack for the hold reason”
OwnerPlatform / engineeringOps, finance, RevOps on the roster

How to evaluate collaborative AI is the buying sheet for the shared-job shape this design depends on. Run it on any vendor. Do not treat a fluent multi-user chat as a pass.

What belongs on the job, not in the agent

Treat four artefacts as properties of the work object, not of the chat:

1. The brief. What we are trying to finish, in one sentence a stranger can read. “Release credit hold on SO-88412 if margin clears the floor” is a brief. “Look at this” is not.

2. The sources. CRM excerpt, WMS snapshot, partner email — attached once, visible to the roster. Re-forwarding is how versions fork. What an AI workstream is is the container that holds those attachments with the job.

3. The proposal and the payload. What the model wanted to change, field by field, before anything executed. Write-back governance is the control on execution. The payload still belongs on the job if the write never ran. A rejected write is data.

4. The signer and the stored rejection. A name, a timestamp, a “no” that stays visible. What is human-in-the-loop AI is not a label on a diagram. It is a person on this roster who can halt a live change while others watch.

The agent may draft all four. It should not be the only place they exist.

Governance as a multiplayer primitive is why the roster and the stop live in the room, not in a PDF emailed after the fact. Inherited authority — the agent cannot exceed the acting person — survives a model swap only if it is attached to the job, not to last week’s service account.

What disposable looks like on a real handover

Picture a credit hold on a strategic account. Sales opened a collaborative AI job last Thursday: the AR ageing extract, the customer’s PO, and a draft note to release the hold. Finance rejected — margin still below floor. The agent that drafted the release note was a frontier model on vendor A.

Monday, vendor B’s model is routed for cost. Ops reopens the same job. They see the rejection, the margin sheet, and the named finance delegate. They do not ask “what did ChatGPT say?” They continue the work.

If Monday’s operator had to paste the ageing extract again, re-explain the rejection, or trust a summary with no signer, the agent was the system of record. That is the opposite of disposable.

The same pattern shows up wherever a signed artefact crosses teams. Collaborative AI for revenue operations is the pipeline version — a stage move finance never saw. Collaborative AI for finance and planning is the books version — a journal draft that lived in a personal window through close. Collaborative AI for operations is the warehouse-and-hold version. The agent is interchangeable in all three. The job is not.

ISO/IEC 42001 wants named actors and operational controls. A job with a roster and a stored rejection is closer to that instinct than a shared login with a long context window. The standard does not mention agents. It mentions accountability. Accountability needs an object that outlives the labour.

When a personal assistant is enough

When one person owns the draft and nothing writes to a live system, keep the personal tool. Collaborative AI and personal assistants is the decision guide. Disposable design is not a moral campaign against copilots. It is a boundary.

Disposable design starts when:

  • A second department must stand on the result.
  • A handover will happen before the job finishes.
  • A customer-facing or system-of-record change can leave the room.
  • An auditor or partner will ask what was approved, not what was said.

Multiplayer AI vs multi-agent AI adds a cast of models only after the shared job exists. A swarm without a work object is several disposable agents and zero continuity. More models do not create a roster. They create more sessions to lose.

If the work is already a known Monday checklist — the same extract, the same rules, the same notify list — do not staff it with an improvising agent every week. Loop engineering is the discipline for compiling that path so it runs without a chat ritual. A loop is not an agent is the distinction. Disposable agents and compiled loops are complementary: the loop owns the known path; the agent, when you need one, is labour you can replace.

How to test disposable design without a programme

Run one recurring exception for four cycles. Do not start with a platform bake-off. Start with a job you already fight in chat.

  1. Open one named job with the two files teams already email. Name the finish line in a sentence a stranger can read.
  2. Record one rejection with a name — not “the channel said no.” The rejection stays on the job.
  3. Swap the model or vendor mid-cycle. Reopen the job. Can a stranger continue without the original thread?
  4. Remove the original prompter from the roster. Can finance still see the payload they rejected?

If step three fails, you bought a chatbot with integrations. If step four fails, RBAC for enterprise AI is missing on the job, not only on the login.

Measure three numbers, even if you measure them by hand:

  • Reopen time. Minutes for a stranger to continue. If it requires Slack archaeology, fail.
  • Signer completeness. Every refusal and approval has a name and a timestamp on the job.
  • Payload replay. The exact proposed change is visible without the original model.

NIST’s framework asks for measurable governance. What is AI governance is the programme. Disposable jobs are one enforceable habit inside it: you can reopen, you can name the signer, you can quote the payload.

Thoughtworks’ operating-system framing for enterprise AI stresses organisational ownership — not only a builder harness around a model. Ownership needs an object. The agent is not that object.

What people get wrong

“We export the transcript.” A transcript is not a payload. It is not a signer. It is not scoped to the CRM row. Transcripts are useful for debugging. They are a poor system of record.

“The agent remembers the customer.” Memory in a personal thread is the user’s seat. Company memory is the job, the asserted policy, and the systems of record — not the model’s window. Do not retell that as a retrieval project. Cite the memory pages and keep decisions on the job while it is live.

“Model-agnostic slide, single-vendor default.” Portability on a deck is not disposable jobs in production. Ask to reopen last month’s exception on this month’s model.

“Institutional memory is a vector store.” Indexing everything without owners produces sludge. Disposable agents keep the small set of assertions and decisions on the job while it is open. The index is not the signer.

“Multi-agent replaces multiplayer.” More models do not create a roster. Governance as a multiplayer primitive does. A cast of specialists that cannot show finance the same rejection is a demo.

“Disposable means cheap models only.” You can run a flagship model on a disposable design. You can run a cheap model on a design that still buries the job in a thread. Cost and continuity are different axes. What is AI token economics is the spend guide; this page is the continuity guide.

How to start this quarter

Pick one cross-team exception — hold, discount, stage conflict — already fought in chat every week.

Name the job. Attach the files once. Name the closer. Store the first rejection. Swap the model on week three on purpose. Invite a reader who was not in the original thread to reconstruct signer and payload without Slack.

Four pillars of an enterprise AI platform is the wider stack — wiki, workstreams, routing, governance as layers, not a single chat window. How to evaluate collaborative AI is the sheet for the shared-job shape. If the room is missing, start from what is collaborative AI and what is an AI workstream. If the stop is missing, start from write-back governance. If the agent is doing Monday’s known job, compile it — see loop engineering — instead of restaffing it with a persona.

The agent should be glad to leave. The project should not notice.

How this shows up in Nimbus

In Nimbus, the durable object is the workstream: brief, attachments, roster, proposed payload, stored rejection, and signer. Agents assigned to that workstream are labour. Swapping a model or a vendor does not empty the job.

Governance keeps inherited authority and fail-closed writes on the same object, so the next agent arrives with the same bounds, not a fresh superuser token. Standing orders that should not need an agent at all are compiled as loops on that workstream — see what is a Nimbus Loop.

Score the design with how to evaluate collaborative AI and the harness sheet, not a homepage video. If a proof of value cannot reopen last week’s rejection after a model swap, the agent is still the system of record.

Common questions

The job outlives the model

Is disposable the same as “swap the model” in an RFP?

No. A buying test for an agent harness asks whether adapters and tools still work when you change model weights mid-proof. That is a useful test, and it belongs on a different sheet. Disposable design asks whether the job still has its files, its stored rejection, and a named signer after the agent is gone. If Monday’s operator has to reconstruct the hold from Slack, the agent was the filing cabinet. Portability of the runtime and continuity of the work are easy to conflate in a demo. Score them separately, or you will pass an RFP and still lose the exception.

Do we lose history if we replace the agent?

You should not, if history lives on the job rather than in the session. Attachments, proposals, approvals, and rejections belong on the named work object, visible to the roster. A vendor thread the next model cannot read is not a record. Exporting a transcript is not the same as keeping a payload and a signer. If replacing the agent forces a re-upload parade or a Slack archaeology session, the history was never institutional. Design for the swap before you need it.

When should we worry about agent continuity?

Worry when a handover, an audit, or a Monday reopen depends on “the same chat.” If the job dies when someone clears the window, the agent was the system. Worry also when a second department must stand on the result, or when a customer-facing or system-of-record change can leave the room. Personal drafting can stay in a personal tool. Cross-team exceptions cannot. The test is simple: remove the original prompter and the original model, then ask a stranger to continue.

Does disposable mean we should churn vendors every quarter?

No. Disposable is a design claim, not a procurement hobby. Stable vendors and stable models are fine when the job object is the source of continuity. Churning for its own sake wastes integration time and teaches nobody. The point is that a forced change — a price hike, a deprecation, a seat that lapses, a person on leave — must not erase the work. Treat the agent as staff on a shift. Treat the job as the unit that persists.

See what governed AI looks like on your stack.

Connect your tools, run a workstream, and keep every decision on your ledger. Start on Free.