The Boardroom AI Dilemma: Spending Millions Without an Outcome
Why boards struggle to tie soaring API invoices to specific business outcomes, and how to implement job-level AI cost attribution.
Somewhere between the innovation slide and the cash-flow slide, the AI programme disappears. The board has been told that adoption is broad. The finance committee has been told that the bill is manageable. Neither statement is false. Together they are useless. A company can have assistants in every function and still be unable to say, in one sentence, which jobs consumed the money, which of those jobs changed a number that matters, and which of them should be stopped.
That gap is now a governance problem, not a tooling curiosity. McKinsey’s State of AI survey found that 88 percent of respondents say their organisations use AI in at least one business function, up from 78 percent a year earlier, while nearly two-thirds remain in experimentation or piloting and only about a third have begun to scale. Use has become ordinary. Enterprise-level benefit has not. The invoice, meanwhile, does not wait for the operating model to catch up.
| What the board has been told | What is true | What is still missing |
|---|---|---|
| Adoption is broad | 88 percent use AI in at least one function, up from 78 percent | Which jobs consumed the money |
| The bill is manageable | A seat price can be forecast from headcount | Which of those jobs changed a number that matters |
| The programme is progressing | Nearly two-thirds are still experimenting or piloting. About a third have begun to scale | Which jobs should be stopped |
A seat licence is not a cost system
The first distortion is the way vendors price the easy part. A per-user subscription looks like software the company already knows how to buy. It sits next to email and document tools. Procurement can compare it with last year’s collaboration suite. The board can be told that the cost is predictable because the headcount is predictable.
Production work does not behave like a seat. A quiet analyst and a team running overnight document review do not consume the same inference, the same tool calls, or the same human review. Seat maths hides that difference until a finance partner asks why a “flat” subscription produced a variable cloud bill, a professional-services overrun, and a second product bought because the first one could not write back to the system of record. By then the original business case has been replaced by a stack.
Microsoft and LinkedIn’s Work Trend Index makes the organisational version of the same point: 75 percent of knowledge workers use generative AI at work and 78 percent of those users bring their own tools. Seventy-nine percent of leaders say their company needs AI to stay competitive, while 60 percent worry that leadership lacks a plan to implement it. Employees are already spending — time, and often a personal subscription — outside the ledger the board approved. The official bill is therefore a lower bound.
What the board pack actually contains
Open a typical pack and the AI section is a narrative: a pilot count, a logo slide, a sentence about productivity. Sometimes a savings claim inherited from a vendor’s case study. What is missing is the thing a capital committee already requires of a plant, a warehouse system, or a sales-compensation change: a unit of work, an owner, and a stop.
Without those, three failures follow.
Double counting. Marketing books the hours an assistant saved drafting campaigns. Operations books the same hours because the campaign brief came out of a shared workspace. Finance adds both to an “AI ROI” appendix. The company congratulates itself on a saving that was never cash.
Invisible rework. A fluent draft that a manager rewrites is not free. It is labour moved from creation to inspection, often done by a more expensive person, and almost never coded as AI cost. If the inspection fails — a wrong price, a wrong clause, a wrong customer commitment — the cost leaves the AI budget entirely and lands in refunds, credits, or legal.
The shadow bill. Personal accounts, departmental credit cards, and “temporary” API keys do not appear in the enterprise agreement. IBM’s Cost of a Data Breach research found that one in five organisations studied reported a breach involving shadow AI, and that high levels of shadow AI were associated with about $670,000 in higher average breach costs. Only 37 percent had policies to manage AI or detect unsanctioned use. A board that has never seen the shadow bill is not conservative. It is uninformed.
| Failure | How it gets into the pack | Where the money actually sits |
|---|---|---|
| Double counting | Marketing and operations both book the same hours. Finance adds them | A saving that was never cash |
| Invisible rework | The draft looks free. A more expensive person rewrites it | Inspection, then refunds, credits, or legal if the inspection fails |
| The shadow bill | Personal accounts, departmental cards, temporary API keys | Outside the enterprise agreement. About $670,000 on the breach cost where shadow use is high. Only 37 percent can detect it |
Questions for a serious finance review
A useful review does not start with the model catalogue. It starts with jobs.
Which named processes were allowed to use an assistant this quarter? Not “marketing,” but the weekly price-exception memo, the vendor-onboarding pack, the customer-email queue. Who owns each process, and who is allowed to let an output leave the building? What was the ceiling, in money or in steps, and how many runs stopped because the ceiling was hit? A stop is a fact. “We are exploring” is not.
Then the attribution question: can finance join the inference, the tool calls, and the human approval to the same job identifier that already exists for cost centres? If the answer is a spreadsheet assembled the night before the committee, the company does not have an AI cost system. It has a recollection.
Then the substitution question: did the spend replace a contractor, a bureau, an outsourced review — or did it sit on top of all three? Leaders are right to be suspicious of savings that do not show up as a cancelled purchase order. McKinsey’s survey is blunt about the distance between use and scale: most organisations have not embedded AI deeply enough into workflows to realise material enterprise-level benefits. A bill that arrives before the workflow changes is not an investment. It is a pilot with a standing order.
| Ask this | A serious answer | Not an answer |
|---|---|---|
| Which named processes used an assistant? | The weekly price-exception memo, the vendor-onboarding pack, the customer-email queue | “Marketing” |
| Who owns it, and who may let an output leave? | Two names | A platform team |
| What was the ceiling, and how many runs stopped? | A number of stops | “We are exploring” |
| Can finance join inference, tool calls, and approval to a cost-centre job? | The same identifier the ledger already uses | A spreadsheet assembled the night before |
| Did the spend replace a contractor, a bureau, or an outsourced review? | A cancelled purchase order | A vendor case study sitting on top of all three |
What to put in front of the committee next quarter
Pick five jobs, not fifty tools. For each, name the owner, the system of record that may be read, the class of change that may be proposed, and the person who must sign before anything is written. Put a ceiling on the run. Record the refusals. If a job cannot name those four things, it is not ready for production spend, however impressive the demo.
Separate the sanctioned path from the personal one. Samsung’s decision to restrict generative tools on company machines, after staff uploaded sensitive material, is the blunt version. The durable version is a path employees will actually use, because a ban that people route around on their phones is a policy and not a control. Microsoft’s finding that most AI users already bring their own tools is the evidence. The official system has to be good enough that the personal subscription is no longer the path of least resistance.
Report the bill in the language of work. Not “AI spend versus last year,” but spend and stops by job, with the human owner beside the number. A director can challenge a job. A director cannot challenge a platform abstraction.
Finally, retire claims you cannot reconstruct. If the business case says a function will run with fewer contractors, show the contractor line. If it says customer commitments are faster, show the refusal rate and the complaint rate beside the cycle time. Speed without those companions is how a fluent assistant becomes a liability the finance committee meets later, under a different heading.
| Job | Owner | System that may be read | Change that may be proposed | Who signs before anything is written | Ceiling | Refusals this month |
|---|---|---|---|---|---|---|
| One sentence a customer would recognise | A person | Named | A class, not “whatever the model suggests” | A person, before it leaves | Money or steps. A hit stops the run | A number. Zero means you bought fluency without a control |
| Four more, then stop |
A blank cell means the job is not ready for production spend.
One invoice, read properly
Pull last month’s invoice for the largest assistant the company pays for. Do not start with the total. Start with the lines. If the lines are only seats, or only tokens, you do not yet have a bill you can explain. You have a consumption total. Ask the owner — if there is one — to map the largest lines to jobs: which team, which work, which system of record was read, whether anything was sent or posted, what ceiling would have stopped it. The lines that cannot be mapped are the lines to freeze at renewal. A freeze is a management act. A variance explanation that says “adoption was strong” is not.
Then put the personal spend beside it. Expense reports, corporate cards, and the tools people buy because the official one was slow to arrive. Most people who use AI at work bring their own. Those subscriptions are small. The work done inside them is not on your invoice and not in your log. The board does not need a lecture on shadow IT. It needs two numbers: sanctioned spend you can tie to a job, and a pulse on unsanctioned use that touches customers, code, or unpublished figures. The breach research is why the second number matters: incidents involving shadow AI were common in the studied set, and heavy shadow use lined up with higher cost.
Finally, match the invoice to the claim. If a business case said contractors would fall, show the contractor line. If it said the assistant does not invent policy, show a refusal from the month you are paying for. A month with spend and no refusals is a month you bought fluency without a control. Public companies and regulated advisers have already been sanctioned for describing capabilities they could not demonstrate. A private board should want the same demonstration before it approves the renewal. The demonstration fits on a page: five jobs, five owners, spend, stops, and one result against a baseline. Widespread use is not that page. The page is how use becomes something a director can govern.
| Start here | Freeze at renewal if |
|---|---|
| The largest lines, not the total | They are only seats or tokens, and nobody can name the job |
| Team, work, system of record read, whether anything was sent or posted, the ceiling | The owner cannot map the line |
| Sanctioned spend tied to a job | The only figure is a platform total |
| A pulse on unsanctioned use that touches customers, code, or unpublished figures | You have the enterprise invoice and not the personal one. The official bill is a lower bound |
| The contractor line, if the case said contractors would fall | The case is still in the pack and the line did not move |
| One refusal from the month you are paying for | Spend with no refusals |
A call to finance chairs and chief financial officers
The board does not need another vision deck. It needs an AI bill it can explain to itself: which jobs, which owners, which ceilings, which stops, which claims have been withdrawn because the evidence was not there. NIST’s AI Risk Management Framework puts the same idea in public language — govern, map, measure, manage — but the framework does not close the books. Finance does.
Companies that keep treating AI as a seat product will keep approving spend they cannot attribute and discovering risk they did not budget. Companies that attach the spend to named work will find that the interesting argument is no longer whether to “do AI.” It is which jobs are worth the next increment of money, and which should be turned off. That is a conversation boards already know how to have. The work now is to make the AI programme fit it.
References
About Nimbus
Nimbus is a Collaborative AI Operating System built around four core pillars that bring human teams and autonomous AI together into a single, unified workspace.
Communication: Keep context tied to the job. Unify emails, meeting recordings, transcripts, and operational files directly within active projects—ending knowledge silos buried in private inboxes, scattered Slack threads, or unrecorded calls.
Collaboration: Work alongside AI in real time. Bring people and AI agents onto the exact same brief, visual canvas, or initiative. Query company-wide data, invite agents into live calls, and co-create in one shared space—eliminating the split between human group chats and isolated AI sidebars.
Automation: Put routine workflows on autopilot. Connect more than 2,000 enterprise tools and standardize repetitive operations. Background loops run on schedules or data triggers with full execution logs, ensuring operational knowledge is shared across the team rather than trapped in one person’s head.
Governance: Deploy AI with absolute control. Enforce strict role-based access controls across workspaces. AI agents can analyze, summarize, and draft—but no live system changes or external communications occur without explicit, verified human sign-off.
Spend the board can attribute
Why can boards not explain the AI bill?
Usage is everywhere and the invoice is not tied to a job, a team, or an outcome.
What should finance be able to say?
Which work consumed the spend, what the ceiling was, and what would have been refused if the ceiling had been hit.
Is a seat price that explanation?
No. A seat describes a login. It does not describe the run.
See what governed AI looks like on your stack.
Connect your tools, run a workstream, and keep every decision on your ledger. Start on Free.