Thought Leadership

The Financial Mechanics of AI: Accounting for Compute over Seats

How agentic AI shifts enterprise IT budgeting from per-seat SaaS headcount pricing to variable compute and tool-execution costs.

Procurement knows how to buy a login. A person, a month, a price that moves when headcount moves. The comparison set is email, documents, a CRM seat. Finance can forecast it. The board can be told the cost is predictable because the workforce is predictable.

Agent work does not behave like a login. One analyst might ask three questions. A workstream might read a contract set, call a ledger, draft, get refused, draft again, and stop before anyone posts a journal — or it might not stop, and the meter continues. Two customers with the same headcount will consume wildly different inference once one of them lets runs touch real volume. The seat was the easy line. The workload is the bill that arrives as a variance, an overage, or a second product bought because the first could not be capped.

McKinsey’s 2025 survey is why this matters now rather than later. Eighty-eight percent of organisations use AI in at least one function. Nearly two-thirds have not scaled it. The invoice does not wait for the operating model. Companies are paying seat prices for a pattern of use that is still a pilot, and they cannot say which jobs consumed the money. A unit of work is how the bill becomes something a finance committee can approve, cap, and stop.

Why the seat breaks

A seat assumes the costly thing is the person being entitled to open the application. For a chat window used a few times a day, that assumption roughly holds, which is why the first wave of enterprise assistants was sold that way. It stops holding when the entitled actor is a run that can loop.

Loops multiply cost without multiplying headcount. A simulation that tries several treatments of a revenue line, a swarm that reviews a queue overnight, a retry that calls a tool again because the first response was malformed: none of these appear as a new employee. They appear as inference, retrieval, and orchestration. If the contract only knows seats, those units are either bundled until they are not, or they spill into a cloud invoice nobody mapped to a job. Finance then meets the programme as a surprise. Surprises are how standing orders survive reviews that should have killed them.

There is a second distortion. Seats are bought for people. A large share of actual use is not on the enterprise seat at all. Seventy-eight percent of AI users bring their own tools. Those subscriptions are small. The work done inside them is off the ledger the board approved, and off the log. The official seat bill is a lower bound on both spend and risk. IBM’s 2025 figures put a cost on the shadow path: one in five studied breaches involved shadow AI, and high shadow use was associated with about $670,000 in higher average breach cost. A seat true-up will not show that number. A workload view at least tells you which sanctioned jobs exist, so the unsanctioned ones are the residual rather than the whole practice.

SeatWorkload
Forecast from headcountForecast from jobs and their ceilings
One person, one login, roughly similar useA quiet user and an overnight review are different orders of cost
Overage is a surprise or a true-upQuoted work pauses. More units are a pack or a commitment you choose
“Adoption” is the success metricA stopped run is a success metric
Shadow subscriptions sit outside the true-upUnsanctioned use is visible as work that never hit a job identifier
The business case is a vendor calculatorThe business case names a job, a baseline, and a cancelled cost

The unit is the run, not the chair

Nimbus meters use in Nimbus Token Units. An NTU is a prepaid unit of completed work: analysis that retrieves or calls tools, governed runs, reports, writes, index jobs. It is a work credit. It is not a raw count of the tokens a model vendor bills wholesale, and it is not a fraction of a seat. Lightweight questions that do not retrieve or call tools can sit outside that meter; the moment the work becomes a run against company data, it draws the pool. Current plan figures, and the line between the two, belong on the pricing page. They will move. The mechanic is what finance should understand when the price changes.

The pool is organisational. Operators share one allowance for the period. A quiet login does not trap a private bundle the close cannot use, and a busy login does not receive a larger bundle merely because a seat was added. On the published plans, extra seats entitle more people to start work. They do not enlarge the NTU pool. Capacity is a separate decision: a higher allowance, a commitment, or a pack bought with eyes open. That split is the point. Headcount and workload are different budgets. Binding them back together is how the seat model hides the bill.

High-commitment work shows a quote and a ceiling before anyone confirms it. The operator sees the units this brief is expected to consume, and the point at which the run will halt. A failed run does not book units. A successful retry books the quoted work once. Near the top of the pool, quoted work pauses until someone adds capacity or the period resets. The pause is the control. A meter that continues and invoices is a different product, and it should be bought as that product, on purpose.

The unit also gives procurement a question the seat never did. What does this job cost in units, and what happens at the ceiling? A product that continues and bills is a different risk from a product that stops and records the stop. Put the stop in the order. A warning is a suggestion. The same discipline applies to any meter sold as intelligence: if you cannot join the unit to a job identifier you already use for cost centres, you have a recollection, not a cost system.

QuestionSeat answerNTU answer
What are we buying?The right to log inUnits consumed when runs execute, pooled across operators
Who can starve whom?A quiet seat hoards its bundleThe pool follows the job, inside a ceiling
What happens when use spikes?A surprise, or a true-up after the factQuoted work pauses until you add capacity on purpose
Can we stop a runaway job?Not if the seat is simply “on”A workspace or run cap halts execution and leaves a record
What do we show the board?Licences versus last yearUnits, stops, and the job they belonged to
What does another seat buy?Another login, and often a vague sense of more capacityAnother person who may start work. The pool stays the pool until you change it

What the unit is counting

A finance partner should be able to say, without a vendor engineer in the room, which activities draw the meter. Four do, in practice.

Model work that is actually a job. A short question with no company context is cheap and often unmetered. A pass that reads a contract set, drafts against a policy, and revises after a refusal is the workload. The unit collapses that activity into something a cost centre can hold. It does not pretend that every prompt costs the same.

Tool calls. Each reach into a ledger, a file store, or a ticket system is work, and it is also a place the run can loop. A malformed response that is retried is a second call. The ceiling is what stops “retry until it looks right” from becoming an unbounded invoice. If the connector can write, the retry is no longer only a cost. It is a second side effect, which is why the quote and the release belong together.

Retrieval and indexing. Grounding a run in the current playbook and the current records has a cost the first time the corpus is prepared, and a smaller cost when the same job runs again against material the workspace already holds. That is the economic content of “context that compounds.” The first forecast pack pays to assemble the sources. A later pack, on the same workstream, should be able to start from the chain and the playbook version rather than from a blank paste. Treat any specific decline in the quote as a design aim you can watch on your own usage, not as a saving a brochure has already booked.

Orchestration. Splitting a brief across specialist passes costs more than a single window, and it is often the correct spend. You are buying a finance pass that cannot send mail and a policy pass that cannot invent a clause. The unit makes that choice visible. A single generalist prompt can look thrifty on the meter and expensive in concessions. Attribute both the units and the outcome, or the cheap run will win every review and lose on the customer sentence.

ActivityWhy it belongs on the meterWhat a ceiling is for
A governed run against company dataThis is the workload that replaces a seat’s fiction of “similar use”Halt a brief that would otherwise loop
Tool callsExternal systems are where cost and side effects accumulateStop a retry from becoming a second write
Indexing and retrievalThe first assembly of sources is real workKeep corpus jobs inside a period allowance
Specialist orchestrationSeparation of mandates is worth paying forPrevent a swarm from spending past the job’s value

The same questions apply to any meter sold as intelligence. If the unit cannot be joined to a job identifier you already use for cost centres, you have a recollection. If failed work is billed the same as finished work, you will hesitate to stop a bad run. If another seat silently buys more inference, you are back to forecasting cost from headcount.

Caps before the swarm starts

Financial control that arrives in the monthly invoice is archaeology. The run has finished. The money is spent. The review explains. Caps move the decision to before the long job starts, which is the only moment a finance partner can still say no.

Three levels are enough. A workspace pool is the budget for a team’s runs this period. A workstream estimate is what this brief is expected to consume before specialists start, shown to the operator in units rather than discovered afterwards. A hard stop is the point at which the run halts even if the model would like another pass. Soft warnings exist for people who are watching. They are not the control. The control is the halt, recorded, so “we explored” cannot be confused with “we overran.”

Approval gates belong next to the cap when the run is allowed to cause a side effect or to spend past a threshold a manager cares about. A simulation that only drafts can be cheap to retry. A simulation that is about to post, send, or pay is no longer a metering question. It is a release question. The person who releases it should see the cost and the payload together. Splitting them — finance sees the bill next month, operations sees the sentence now — is how both miss the decision.

This is ordinary commitment control applied to a new meter. You would not let a plant order materials with no purchase-order limit because the vendor’s interface was conversational. Do not let a swarm do it because the interface is a brief. NIST’s govern-and-measure language is the public version. The internal version is a ceiling that fires.

GateWhen it firesWhat “fired” looks like in the pack
Estimate before startThe operator sees expected units for this briefA number, not “it depends”
Workspace poolThe team’s period allowance is exhaustedNew runs wait, or overage is explicitly on
Hard stop on a runThe brief hits its own cap mid-flightThe run halts. The halt is stored
Release on a write or a high spendA person must see payload and costA name, or the action did not happen

Attribute the unit to a signed outcome

Units are not a benefit. They are a cost. The benefit is a job that finished: a commentary a controller will sign, a claim answered against the current rule, a queue cleared without a rise in reopen rate. Roi that stops at “hours saved,” especially hours a vendor modelled, is the personal clock. The operating clock is elapsed time, quality, and a step you actually deleted, set against a baseline you froze before the run was allowed in.

Join them. For each workstream worth money, report units consumed, whether it stopped at a ceiling, whether a person released or refused the output, and what happened to the external measure you already trust — concessions, rework, a contractor line. A month of units and no refusals is a month you bought fluency without a control. A month of units and no movement in the job is a popular meter. Most organisations are still in that pattern. The ones that leave it will be the ones whose finance team can point at five jobs and say what the units bought.

Do not book a saving from the meter. Book a saving when the job got shorter or cheaper against the baseline, the error rate did not worsen, and something left the calendar or the contractor list. Hold headcount claims until that cycle has happened. Taking the seat-count “efficiency” in the announcement is how the cost returns later as complaints and contractors under another name.

Public cases about describing a capability you cannot show are a useful constraint on the business case itself. “AI-driven savings” is a claim. If the evidence is a unit total and a survey of people who feel faster, say that, and do not say more. The FTC’s instruction on AI claims is aimed at advertisers. It is also a sound rule for an internal pack.

Show thisDo not show this as the result
Units by workstream, and stopsA single platform total labelled “adoption”
Release or refusal beside the unitsA month of spend with no refusal
Baseline versus outcome on the jobA vendor calculator of hours saved
The contractor or meeting that disappearedA headcount reduction taken before the quality cycle
Overage, if you chose it, as a conscious rateA cloud variance nobody can name

One recurring pack, attributed

Consider a commentary workstream finance already trusts enough to put in front of a controller. Before the first run, the operator sees a quote in units and a ceiling. The brief names the finish: a pack that cites the ledger and the playbook, and does not post. The run either completes inside the ceiling or halts and records the halt. Units book when the run succeeds. The person who will sign the pack is the person who releases it. Next month the same workstream runs again. The quote is visible again. If the workspace already holds the sources and the prior release, the quote can come down. You will see that on the usage record if it happens. You should not put a brochure’s illustration into the board pack as if it were your result.

The attribution line is short. Workstream name. Units booked. Whether it halted. Whether a person released or refused. The external measure you already track for that job: cycle time to a signed pack, number of restatements, a contractor day you did not buy. A quarter of those lines is a cost system. A quarter of platform-wide totals labelled “AI spend” is a variance with a new name.

Procurement’s schedule can require the mechanic without freezing a price that will move. Define the unit as completed work, not as a wholesale model token and not as a seat. State that the allowance is pooled at the organisation, and that seats do not increase it. State what happens near exhaustion: an alert, then a pause of quoted work, until someone adds capacity on purpose. State that a failed run does not book. State that usage can be exported by workstream, beside the release. A vendor that can only invoice seats, or can only invoice after the overrun, has not met the schedule. Pay them for a chat window if that is what you want. Do not pay them as if they had sold a workload you can stop.

Put this in the orderThe answer that counts
What is the unit?Completed governed work, pooled, separate from the seat
What does a seat add?A person who may start work. Not a larger inference budget
What happens before a long run?A quote and a ceiling the operator can see
What happens if the run fails?It does not book. A successful retry books once
What happens when the pool is exhausted?Quoted work pauses. More capacity is a decision
What do we receive each month?Units, halts, and releases by workstream

A call to chief financial officers

Keep the seat if you need to know how many people may start work. Do not let the seat be the cost system. Ask for the unit, the pool, the ceiling, and the stop. Ask which jobs consumed last month’s units and which of those jobs changed a number you already manage. Freeze the lines you cannot map.

The argument worth having is no longer whether the company will “do AI.” People already are. The argument is which jobs are worth the next increment of units, and which runs should halt. That is a finance conversation. The meter has to be built so the conversation can happen before the invoice, not after it.


References

About Nimbus

Nimbus is a Collaborative AI Operating System built around four core pillars that bring human teams and autonomous AI together into a single, unified workspace.

Communication: Keep context tied to the job. Unify emails, meeting recordings, transcripts, and operational files directly within active projects—ending knowledge silos buried in private inboxes, scattered Slack threads, or unrecorded calls.

Collaboration: Work alongside AI in real time. Bring people and AI agents onto the exact same brief, visual canvas, or initiative. Query company-wide data, invite agents into live calls, and co-create in one shared space—eliminating the split between human group chats and isolated AI sidebars.

Automation: Put routine workflows on autopilot. Connect more than 2,000 enterprise tools and standardize repetitive operations. Background loops run on schedules or data triggers with full execution logs, ensuring operational knowledge is shared across the team rather than trapped in one person’s head.

Governance: Deploy AI with absolute control. Enforce strict role-based access controls across workspaces. AI agents can analyze, summarize, and draft—but no live system changes or external communications occur without explicit, verified human sign-off.

Short answers

A unit of work, with a ceiling

Why does a seat price misdescribe an agent?

A seat is a person with a login. An agent run is compute, tool calls, and a job that may never map to a headcount line.

How should the bill be shaped?

As a unit of work, with a ceiling, before it is a subscription.

What should finance see before the run?

A quote for the job and a cap that stops it. Unlimited AI is not a control.

See what governed AI looks like on your stack.

Connect your tools, run a workstream, and keep every decision on your ledger. Start on Free.