Thought Leadership

Why AI Pilots Fail the Quarterly Review Every Time

The top reasons AI proofs of concept stall at the quarterly review, and how establishing baseline metrics and clear owners saves them.

The pilot survives the demo. It rarely survives the quarter. A team shows a fluent summary, a faster draft, a chart that used to take an analyst a day. The room is impressed. Someone says “let’s scale this.” Three months later the same slide is still a pilot, the owner has moved teams, and the data it depended on has changed. The tool is not the thing that died. The absence of an operating owner did.

McKinsey’s 2025 global survey describes the pattern at industrial scale. Eighty-eight percent of respondents say their organisations regularly use AI in at least one function. Nearly two-thirds have not begun scaling it across the enterprise. About a third say they have started. Use is no longer the news. The news is that use and scale have come apart, and that the quarterly business review is where the gap becomes visible to people who control headcount and capital.

The following examines why pilots stall in that room, what leaders mistake for progress, and how to design an experiment that can either become a job or be shut without theatre.

What the survey is willing to sayWhat the quarterly review is then forced to ask
88 percent use AI in at least one functionWhat changed in the P&L, the cycle time, the error rate, or the risk register?
Nearly two-thirds have not begun scalingWhy is this still a pilot if people already depend on it?
About a third say they have started to scaleWhich job, with which owner, survived a baseline?

A pilot is a permission slip, not a product

In most companies a pilot is how a function gets to try something without promising a result. That was reasonable when the something was a sandbox. It is a poor container for a system that employees are already using every day. Microsoft and LinkedIn’s Work Trend Index found that three-quarters of knowledge workers use generative AI, and that 78 percent of AI users bring their own. The “pilot” on the roadmap and the “assistant” on the laptop are not the same programme. One is governed by a slide. The other is governed by habit.

When the quarterly review arrives, the sponsor is asked a question the pilot was never built to answer: what changed in the P&L, the cycle time, the error rate, or the risk register? A demo cannot answer it. A usage chart cannot answer it. The team reaches for anecdotes. The finance partner reaches for a baseline that was never frozen. The review ends with “continue to learn,” which is how organisations store work they are not willing to fund or kill.

What actually kills the pilot

The model is seldom the cause. Four management facts are.

No frozen baseline. If you did not record how long the job took, how often it was wrong, and what it cost before the assistant arrived, you cannot claim an improvement that a sceptical controller will accept. Teams discover this in the meeting. They should have discovered it in week one.

No system of record. The pilot lives in a chat beside the work. The work lives in the CRM, the ledger, the contract store, the ticket queue. When the quarter ends, nobody can reopen the pilot and see the decision. They can reopen a transcript, if it was retained, which it often was not. A quarterly review is a memory test. Chat is a bad memory.

No named signer. Someone has to be willing to say that this output may leave the team. If that person does not exist, the pilot is entertainment. The moment a customer, an auditor, or a regulator might rely on it, the sponsor retreats. Air Canada’s chatbot case is the external version of the same retreat: the airline argued, unsuccessfully, that the bot was a separate legal entity. The tribunal decision held that a company is responsible for information on its own website, whether the words come from a static page or a chatbot. Internally, leaders already know this. They just prefer not to attach their name until the pilot feels finished. It never feels finished.

A success metric borrowed from the vendor. “Time to first draft” is not a business metric. “Messages handled” is not a business metric if the messages are worse. Klarna and others have published striking service-automation figures; the leadership question is not whether a vendor can produce a large number. It is whether your complaint rate, your refund rate, and your reopen rate moved with it. If those were not in the pilot charter, the quarterly review will invent a verdict anyway, and it will be no.

What is missingHow it shows up in the roomThe verdict a controller will reach
A frozen baseline: time, error rate, cost, before the assistantThe team discovers the gap in the meetingNo improvement you can accept
A system of recordThe work is in the CRM or the ledger. The pilot is in a chat that may not have been keptYou cannot reopen the decision. Chat is a bad memory
A named signerThe sponsor retreats the moment a customer or an auditor might rely on the outputThe pilot was entertainment. A company is responsible for words on its own site
A metric that is not the vendor’s“Time to first draft” or “messages handled,” with complaint, refund, and reopen rates absent from the charterThe review invents a verdict. It will be no

The agentic version of the same stall

The newer disappointment has a new noun. Agentic systems — assistants that plan steps and call tools — are being trialled widely and scaled narrowly. McKinsey’s researchers have been plain that experimentation with agents is common and that scaling is usually confined to one or two functions, often IT or knowledge management. That is not a paradox. An agent that can act is an operating change. An operating change needs a boundary: what it may read, what it may propose, who stops it.

A pilot that grants an agent a broad connector “so we can see what it does” will not survive contact with security, and it should not. IBM’s 2025 breach research found that among organisations reporting an AI-related breach, 97 percent lacked proper AI access controls, and that 31 percent of those incidents produced operational disruption. A quarterly review that celebrates an agent demo and has not asked about access is celebrating the precondition of an incident.

The review that kills a pilot cleanly

A pilot dies usefully or it dies in disguise. The useful death happens in the quarterly review, in front of the sponsor, with the baseline on the table. The disguised death is “continue to explore,” which keeps the licence, the Slack channel, and the manager’s rework, and removes only the obligation to conclude.

Set the review up so disguise is hard. One job, described in a sentence a customer would recognise. The baseline, frozen before the pilot, for cycle time, error or reopen rate, and one external cost. The result on the same definitions. The step that was deleted, or the honest note that nothing was deleted. The refusal count. If any of those is missing, the outcome is stop, not a request for a follow-up deck. Follow-up decks are how a quarter becomes a year. The survey picture — widespread use, a minority scaling — is what those decks add up to.

Invite the manager who has to sign the output, and let her speak before the sponsor. If she says the assistant’s answer cannot be sent, booked, or paid, that is not resistance. That is the result. Leaders worry their organisations lack a plan even while staff use the tools every day. The plan is this meeting. It fails when the sponsor narrates adoption and the operator is given two minutes at the end.

Allow three outcomes only.

Promote. The job moves into the ordinary budget, the old step is removed, an owner is named.

Narrow. The pilot keeps only the slice that met the bar, and the rest is turned off, including the credentials.

Stop. The channel is closed, the story is written down, and the team is thanked for a negative result.

Negative results are how the next pilot gets a better question. A culture that punishes them will produce pilots that cannot fail and therefore cannot teach.

Write the stop in the same pack that once announced the launch. People remember the launch. If the stop lives only in a footnote, the next team will rebuild it. If you told the market or the board that the pilot was already a capability, correct that description with the same specificity. Firms have settled charges for describing AI they were not running. An internal myth is how an external sentence starts. Say what you can show.

There is a security reason to stop cleanly as well. A pilot that “continues” often continues with the broad service account that made the demo work. AI-related compromises are overwhelmingly associated with missing access controls. Ending the pilot includes revoking the token. A token without an owner is not exploration. It is an account.

On the tablePresentMissing
One job, in a sentence a customer would recogniseThe review can judge a jobYou are judging a demo. Stop
Baseline frozen before the pilot: cycle time, error or reopen rate, one external costThe result uses the same definitions“Continue to learn.” That is how a quarter becomes a year
The step that was deleted, or an honest note that nothing wasA saving you can see on a calendarNothing was deleted. Do not claim one
Refusal countThe control was in the pathA token with no owner is still an account
The operator speaks before the sponsor“Cannot be sent, booked, or paid” is a resultAdoption narrative, two minutes for the person who signs
OutcomeWhat changes the same weekWhat does not count
PromoteOrdinary budget, the old step removed, a named ownerA pilot that keeps its side channel
NarrowOnly the slice that met the bar stays on. Credentials for the rest are revokedA smaller demo of the same unbounded tool
StopThe channel is closed, the story is in the same pack that announced the launch, the token is revoked“Continue to explore” without new money

Write the stop where you wrote the launch

The review allows promote, narrow, or stop. If the baseline is missing, the outcome is stop. Record the stop in the same pack that announced the pilot, and revoke the credentials the same week. A pilot that “continues to explore” without new money is a system with no owner. That is how use stays widespread and scale stays rare. Negative results are a product of a functioning review. Protect them.

A call to operating executives

Stop collecting pilots. Collect jobs that survived a quarterly test. The companies that pull ahead will not be the ones with the longest innovation backlog. They will be the ones whose reviews can point to a handful of processes that are faster, whose errors are visible, and whose owners would defend the result without a vendor in the room.

If a pilot cannot be defended that way, end it in the meeting. The kindness is not another quarter. The kindness is giving the organisation permission to spend its attention on work that is real.


References

About Nimbus

Nimbus is a Collaborative AI Operating System built around four core pillars that bring human teams and autonomous AI together into a single, unified workspace.

Communication: Keep context tied to the job. Unify emails, meeting recordings, transcripts, and operational files directly within active projects—ending knowledge silos buried in private inboxes, scattered Slack threads, or unrecorded calls.

Collaboration: Work alongside AI in real time. Bring people and AI agents onto the exact same brief, visual canvas, or initiative. Query company-wide data, invite agents into live calls, and co-create in one shared space—eliminating the split between human group chats and isolated AI sidebars.

Automation: Put routine workflows on autopilot. Connect more than 2,000 enterprise tools and standardize repetitive operations. Background loops run on schedules or data triggers with full execution logs, ensuring operational knowledge is shared across the team rather than trapped in one person’s head.

Governance: Deploy AI with absolute control. Enforce strict role-based access controls across workspaces. AI agents can analyze, summarize, and draft—but no live system changes or external communications occur without explicit, verified human sign-off.

Short answers

Praised, then cancelled

Why do AI pilots die in the quarterly review?

They are praised as experiments and then cancelled because nobody changed the job they were supposed to replace.

Is that a model failure?

Usually it is a management failure. The pilot never had an owner, a baseline, or a decision it was allowed to change.

What should the review ask?

What job changed, what it cost, and whether it will run next quarter without a special team.

See what governed AI looks like on your stack.

Connect your tools, run a workstream, and keep every decision on your ledger. Start on Free.