Thought Leadership

Customer Trust and the Cost of Your First Public AI Mistake

Why customers equate chatbot promises with core brand failure, and how to build incident response plans for public AI errors.

Customers do not experience your architecture. They experience a sentence with your name on it. When that sentence is wrong, the argument that “the AI said it” lands the way “the intern said it” would have landed, except the intern did not speak at the scale of a website, and the intern could be asked what they meant. The model cannot. Trust, after the first public mistake, is rebuilt by showing a change in how sentences are allowed to leave, not by a post that says you take the issue seriously.

The case law and the market have already supplied the examples. Air Canada told a grieving passenger, through its website chatbot, that a bereavement fare could be claimed after travel. The policy page said otherwise. The tribunal refused the idea that the bot was a separate entity. In February 2023 Google’s Bard demo included a factual error about a satellite; Reuters reported that Alphabet shed on the order of $100 billion in market value in the session. Different stakes, same mechanism: a public sentence, an identifiable company, an audience that does not grant a hallucination discount.

Air Canada, 2024Alphabet, February 2023
The sentenceA bereavement fare could be claimed after travelA factual error about a satellite, in a promotional demo
What the company wished were trueThe policy page, which said otherwiseThe model was not yet the product customers should rely on
Who decidedA tribunal. The bot was not a separate entity. The decision held the airline to the answer. CBC’s account is the operating versionThe market, in one session, on the order of $100 billion in value
What a customer or investor could checkA screenshot against the policyThe claim against a known fact

Why the apology fails

The standard crisis statement promises a review, a model update, and a commitment to accuracy. Customers have learned what that sequence means. The review will be internal. The model update will not be visible. The next mistake will sound just as confident. Trust is not a tone. It is a prediction that the next interaction will be governed. If nothing in the customer’s experience changes — no clearer policy, no easier path to a person, no visible correction of the specific claim — the prediction does not change.

There is also a fairness problem the apology skips. The customer who relied on the wrong sentence has often already paid, travelled, or declined another offer. Air Canada’s passenger had flown. A correction that arrives as a lesson for the company, without a remedy for the person, teaches the market that reliance is the customer’s mistake. The tribunal did the opposite. It treated reliance as reasonable. Your recovery design should assume a future decision-maker will do the same.

Line in the statementWhat customers have learned it meansWhat would change the prediction
“We are reviewing”The review will be internalThe sentence, preserved, and the name of who can now stop the next one
“We are updating the model”The update will not be visibleThe rule, published where the assistant lives, identical to the assistant
“We take accuracy seriously”The next mistake will sound just as confidentA person sees the class of promise that costs money, before it is sent
A lesson for the company, and no remedyReliance was the customer’s mistakeThe remedy the correct policy would have given, quickly

What to do in the first week, and how to prevent a second

In the first week, find the sentence. Not the topic. The sentence. Preserve it. Identify who could have stopped it and why they did not. If the answer is that nobody was assigned, say that internally without euphemism. Offer the customer the remedy the correct policy would have given, or a better one, quickly. Speed of remedy is the only part of the apology people can verify.

Then change the gate. Customer-facing assistants should be bound to the current policy, not to a pile of pages in which the current policy is merely the most popular. When the assistant and the policy disagree, the assistant loses, and the disagreement is ticketed. A human sees the exact outbound words for the classes of promise that cost money: fares, credits, coverage, delivery, exceptions. The EU AI Act’s oversight language — people able to interpret outputs and intervene — is a good design test even for uses the Act does not classify as high-risk. A footer that says “AI-generated” is not intervention.

Publish, where you can, the rule you corrected. Customers trust companies that show the rule more than companies that show the model. You do not need to publish weights. You need to publish the bereavement rule, the return window, the coverage trigger, in the same place the assistant lives, and keep them identical.

By whenActionDone when
Day 1Preserve the exact sentence. Name who could have stopped itThe sentence is in the incident file, not paraphrased
Day 1 to 3Remedy the customer who reliedThey have what the correct policy owed them, or better, and they can see that they do
Day 7Bind the assistant to the current rule. Retire the page that contradicted itA disagreement between bot and policy opens a ticket, and the assistant loses
Day 7A person sees outbound words for fares, credits, coverage, delivery, and exceptionsThe approver sees the sentence, not a summary of the model’s intention
Same weekPublish the corrected rule where the assistant livesA customer can compare the next answer with the rule without asking which page is real

The internal audience is also the public

Staff read the incident. If leadership’s lesson is “don’t get caught,” staff will route around the sanctioned bot and draft answers in personal tools. Seventy-eight percent of AI users already bring their own. A public mistake followed by a ban, with no usable alternative, increases that share. The next error will be harder to find because it will not be in your log. Samsung’s restriction after a leak made sense as a pause. As a permanent strategy it hides the work. Recovery includes a path staff prefer to their phones, because the public brand is whatever path they actually use.

Marketing will want to move on. Let them move on only after the gate exists. The FTC’s warning on AI claims is a constraint on the comeback campaign. Do not announce that the assistant is now “trusted” or “accurate” because you retrained something. Announce the control a customer can understand: a person checks promises over a threshold; the policy page and the bot cannot diverge without an alert; here is how to reach a human. Claims you can demonstrate are the only claims that repair anything.

The second mistake ends the argument

Customers will forgive a first error if the second interaction proves the system changed. They will not forgive a second error of the same kind. The second error says the apology was a holding statement. It is also the moment journalists, regulators, and plaintiff firms stop treating the incident as a glitch and start treating it as a practice. A public demo error moved a market in a day. A repeated customer-facing error moves a reputation more slowly and more permanently, because each screenshot confirms the last.

Design the fortnight after the incident as a control sprint, not a communications sprint. Freeze the class of promise that failed. Route it to humans. Sample every answer in adjacent classes. Publish the corrected rule where the assistant lives. Tell staff, in writing, what they may not paste into personal tools while the freeze holds, and give them the sanctioned draft that already contains the corrected rule. If you only freeze the website bot, the phone channel will recreate the sentence by lunchtime.

Then look for siblings. The bereavement rule was wrong because a stale answer and a current page were both allowed to speak. Search for other pairs: a return window in the bot and a different window in the terms; a coverage phrase in sales macros and a different phrase in the policy. Each pair is a future screenshot. Retire one side this fortnight. The tribunal’s point was that customers are not obliged to know which of your pages is the real one. Neither are your new employees.

Report to the board in sentences, not themes. What was said. Who relied. What remedy was given. What gate now stops a repeat. What remains unfixed. A board that receives “we have reinforced our commitment to accuracy” has not been briefed. A board that receives the sentence and the control can govern. Claims you cannot demonstrate should stay out of the comeback, including the claim that the problem is solved.

Trust is a lagging indicator. The leading indicators are contradictions found, promises routed to a person, and time-to-remedy for the customer who already relied. Put those three in the weekly customer review until they are boring. Boring means the operating change took. A campaign means it did not.

IndicatorWhat you countWhat “boring” looks like
Contradictions foundPairs of answers for the same noun: bot versus terms, macro versus policyThe count is non-zero while you are looking, then falls because one side was retired
Promises routed to a personFares, credits, coverage, delivery, exceptions that waited for a humanThe class that failed no longer leaves unattended
Time-to-remedyHours from the screenshot to the customer receiving what the correct rule owedThe customer can verify the speed. The model update is not a substitute

What “sorry” must contain

Name the sentence that was wrong. Name the remedy for people who relied on it. Name the rule that replaces it, and the hour the old source was withdrawn. Name the person who now sees that class of promise before it is sent. Leave out the sentence about your commitment to innovation. Customers cannot test a commitment. They can test the next answer. Your repair will be judged the same way the Bard error was: by the next transcript, not by the statement.

Search the siblings in the same week. Two numbers for one noun, anywhere a customer or a seller can see them. Each pair gets an owner and a retirement date. Report the open pairs to the board until there are none. Do not claim the channel is now safe in the abstract. Claim the pairs you closed. Trust is the customer’s ability to predict you. Prediction is one version of the rule, in every tool your staff actually use.

Put this in the noteLeave this out
The sentence that was wrong“We take this seriously”
The remedy for people who reliedA commitment to innovation
The rule that replaces it, and the hour the old source was withdrawn“The model has been updated” with nothing a customer can compare
The person who now sees that class of promise before it is sent“The channel is now safe”

A call to chief marketing and customer officers

Assume the screenshot. Design the assistant as if the worst answer will be attached to a complaint. That is not cynicism. It is how Air Canada’s case was proved. Then rehearse the recovery before you need it: who freezes the bot, who authorises the remedy, who changes the rule, who tells the board the sentence rather than the vibe.

Trust after a public AI mistake is not a communications project. It is an operating change the customer can feel on the next visit. If they cannot feel it, they are right not to come back.


References

About Nimbus

Nimbus is a Collaborative AI Operating System built around four core pillars that bring human teams and autonomous AI together into a single, unified workspace.

Communication: Keep context tied to the job. Unify emails, meeting recordings, transcripts, and operational files directly within active projects—ending knowledge silos buried in private inboxes, scattered Slack threads, or unrecorded calls.

Collaboration: Work alongside AI in real time. Bring people and AI agents onto the exact same brief, visual canvas, or initiative. Query company-wide data, invite agents into live calls, and co-create in one shared space—eliminating the split between human group chats and isolated AI sidebars.

Automation: Put routine workflows on autopilot. Connect more than 2,000 enterprise tools and standardize repetitive operations. Background loops run on schedules or data triggers with full execution logs, ensuring operational knowledge is shared across the team rather than trapped in one person’s head.

Governance: Deploy AI with absolute control. Enforce strict role-based access controls across workspaces. AI agents can analyze, summarize, and draft—but no live system changes or external communications occur without explicit, verified human sign-off.

Short answers

After the wrong sentence is public

Why is a wrong public answer a trust event?

Customers will not separate your chatbot from your brand. The sentence is yours the moment it is on your site or in your name.

Is a better apology the recovery?

No. The recovery is operational: what was said, who could have stopped it, what changed so the same sentence cannot go out again.

What should we keep before the incident?

The exact output, the source it used, and the person or rule that was supposed to approve it. Without that, the apology is all you have.

See what governed AI looks like on your stack.

Connect your tools, run a workstream, and keep every decision on your ledger. Start on Free.