Why Buying AI Demos Leaves Enterprise Procurement Exposed
Why standard RFPs fail to catch systemic AI risk, unowned connectors, and data retention traps, and how to fix your procurement scorecard.
Enterprise buying was built for products that sit still. A database has a capacity. A suite has a seat. A security review has a questionnaire that mostly stays true after signature. Generative systems do not sit still. The model behind the demo changes. The vendor’s subprocessors change. The customer’s own staff discover uses the statement of work never named. Procurement, which is supposed to be the adult in the buying process, is still scoring the demo.
The result is a contract that buys a capability and inherits an operating model by accident. The operating model — who may paste what, what is retained, whether a human must approve a customer-facing sentence, what happens when the vendor’s claim exceeds the product — is where the cost and the liability are. It is rarely the section with the weighting in the scorecard.
What the scorecard rewards
A typical AI procurement scores accuracy on a canned set, integration logos, price per seat, and a security packet. Those are not worthless. They are incomplete.
Accuracy on a canned set does not tell you what the system does with your stale policy. Air Canada’s chatbot was not failing a benchmark. It was contradicting the company’s own bereavement rule, and the tribunal held the company to the contradiction. A buyer who never ran the demo against their own live rule bought a spokesperson.
Price per seat hides the production meter. Two customers with the same headcount will consume wildly different inference once one of them lets the tool touch real volume. If the contract converts overages into a surprise, finance will meet the tool again as a variance, not as a project. Ask for the unit of work and the ceiling, and ask what the product does when the ceiling is hit. A product that continues and bills is a different risk from a product that stops and records the stop.
The security packet is a snapshot. IBM’s 2025 findings — 97 percent of organisations with an AI-related breach lacking access controls, shadow AI in one in five breaches — describe customers more than they describe vendors. A vendor can be certified and still be deployed, by you, without access control. Procurement’s job is to make the deployment conditions contractual: identity, least privilege, logging, a ban on training on your prompts unless you opt in in writing. NIST’s framework is a reasonable exhibit to attach. A logo on a slide is not.
| What the scorecard weights | What it does not tell you | The test that belongs in the RFP |
|---|---|---|
| Accuracy on a canned set | What the system does with your stale policy | Run it against the live rule. If it contradicts you, you are buying a spokesperson |
| Integration logos | Whether a connector can write, and who can cut it off at 9 p.m. | Read versus write, an owner, and a kill switch |
| Price per seat | The production meter once the tool touches real volume | Unit of work, ceiling, and whether hitting the ceiling stops or bills |
| A security packet | How you will deploy it. Certified vendors are still deployed without access control | Identity, least privilege, logging, and no training on your prompts unless you opt in in writing |
Claims are a procurement issue
Vendors are under the same pressure as their customers to sound advanced. Some will overshoot. The SEC’s cases against Delphia and Global Predictions were about advisers, not software firms, but the pattern — “AI-driven” as an untrue statement of fact — is one a buyer should assume exists in the market. The FTC has told advertisers to keep AI claims in check. Translate that into the RFP. Require the bidder to describe, for each headline claim, the feature that implements it today, not on the roadmap. Reserve a right to test that feature on your data before acceptance. If the claim is “human in the loop,” make them show the screen the human sees and the state in which the system is blocked without that human.
ISO/IEC 42001 certification, where a vendor has it, is a sign they run a management system. It is not a sign their sales deck is precise. Read the deck against the admin guide.
| Headline claim | What you require before it scores | Fail |
|---|---|---|
| “AI-driven,” “ensures,” “the first” | The feature that implements it today, not on the roadmap, tested on your data | The deck describes a roadmap. The admin guide does not |
| “Human in the loop” | The screen the human sees, and the state in which the system is blocked without them | The human sees a digest the next morning |
| “Will not invent policy” | A refusal, in your environment, on your stale document, on purpose | They cannot refuse. The claim is advertising |
| A management-system certificate | Hygiene. Still read the deck against the admin guide | The certificate is asked to stand in for the afternoon you did not spend breaking the product |
What a good pilot contract looks like
Procurement’s leverage is highest before signature and almost gone after. Use it to buy a pilot that can fail, not a subscription that can only expand.
The pilot has a named job inside your company, not the vendor’s script. One team, one class of work, a frozen baseline from the month before: cycle time, error or reopen rate, and the concessions or rework you already measure. The vendor may help instrument the run. The vendor does not get to define success as “users liked it.” Liking is not a control. Most organisations are still piloting. A pilot that cannot fail is how that sentence stays true for another year.
The pilot also has a ceiling. Spend, messages, and — if a write exists — the value of the write. Hitting the ceiling stops the run and pages a person. A ceiling that only sends a warning is a suggestion. Put the stop in the order form. Put the owner, a company employee, in the order form too. If the sponsor cannot name the owner, you do not have a deployment. You have a demo with a purchase order.
Demand the artefacts the demo skipped. Where an answer came from, shown to a non-specialist. What is retained, for how long, in which region. Whether prompts are used to improve a model, and how the enterprise agreement changes that default. How you export prompts, configurations, and logs. How fast a connector that can write can be cut off, and who can cut it off at 9 p.m. These are buying questions you already ask of systems that move money or personal data. The fact that the interface is a conversation does not retire them. A customer-facing answer has already been treated as the company’s answer. Buy the log before you buy the fluency.
Score claims against the documentation, not against the deck. If the proposal says the system will not invent policy, have the vendor show the refusal in your environment, on your stale document, on purpose. If they cannot refuse, the claim is advertising. Advertising claims about AI are already a regulatory subject. You do not need a novel clause. You need the refusal recorded in the pilot notes, and a warranty that matches what you saw. A management-system certificate can support the vendor’s hygiene. It cannot replace the afternoon you spent trying to make the product break.
Finally, price the labour you inherit. Someone must keep the cited policy current, review the class of output that can become a promise, and answer the screenshot. If that labour is not in the business case, send the business case back. Procurement is allowed to be the function that says the saving is not real yet. That is the job. A seat price times a headcount, with no unit of work and no reviewer, is how the company acquires a tool it cannot explain and a bill it cannot attribute. The SEC cases were about public descriptions. An internal description you cannot open is how you get there. The EU AI Act will make some of this documentary duty statutory for higher-risk uses. Buyers who wait for the statute to teach them the questions will pay for the lesson in a tender they cannot score.
| In the order form | If it is missing |
|---|---|
| One named job, one team, a baseline frozen last month | The vendor’s script will become the success metric |
| Cycle time, error or reopen rate, concessions or rework | “Users liked it” will be the result. Liking is not a control |
| A ceiling on spend, messages, and the value of any write. Hitting it stops the run | A warning is a suggestion. A product that continues and bills is a different contract |
| A company employee as owner | A demo with a purchase order |
| Retention, region, training default, export, and who can cut a write connector at 9 p.m. | You bought fluency and not the log |
| Labour: who updates the cited rule, who reads the sample, who answers the screenshot | The saving is not real yet. Send the business case back |
After signature, the scorecard is yours
The contract does not operate the tool. In the first month, run the job you named, against the baseline you froze, and keep the refusal when the product invents a policy you do not have. If the vendor’s success manager wants to substitute a different job because it demos better, that is a no. You bought a job. A customer-facing answer remains yours when the job is live, which is why the first month is a legal event and not only an onboarding plan.
Hold the ceiling. If spend or writes blow through it, the run stops. A vendor that treats the ceiling as a notification has sold you a meter, not a control. Escalate that as a contract performance issue, not as a training issue. At the same time, name the internal labour: who updates the cited rule, who reads the sample, who answers the screenshot. If those names are vacant, pause rollout. A vacant name is how a purchase becomes the pile of tools surveys keep finding inside companies that have not scaled.
When the quarter ends, decide with the scorecard you wrote before you were charmed. Promote, narrow, or stop. Do not let a renewal auto-execute because the purchase order is easier to repeat than to revisit. If you stop, export the configuration and revoke the credentials in the same week. A stopped pilot with a live token is still a system. Claims in the proposal that the documentation does not support should already have been struck. If they survived into the contract, surviving them into the website is how a buying mistake becomes a public one.
The first month is the contract
Run the named job against the frozen baseline. Keep every refusal. If the vendor swaps in a prettier job, decline. Hold the ceiling as a stop, not a warning. If the internal owner of the cited rule is still unnamed, pause. At the quarter, promote, narrow, or stop, and if you stop, export the configuration and revoke the token the same week. A live token is not a cancelled pilot. Auto-renewal is how a demo becomes a permanent system nobody chose twice. Choose twice. A second signature is the point of having a scorecard. Without it, the renewal is inertia.
| First month | Hold | Pause or stop |
|---|---|---|
| The job | The one named in the order form, against the frozen baseline | A prettier job the success manager prefers |
| The refusal | Kept, especially when the product invents a policy you do not have | A month of fluency and no refusal. You cannot see the control |
| The ceiling | A stop that pages a person | A warning, or a bill. That is a meter |
| The names | Who updates the rule, who reads the sample, who answers the screenshot | Any vacancy. Rollout waits |
| Quarter end | Promote, narrow, or stop. A stop exports the configuration and revokes the token the same week | Auto-renewal. A live token is not a cancelled pilot |
A call to chief procurement officers
You are not late to a technology trend. You are on time for a buying discipline you already know, applied to a product that talks. Make the vendor show the control, not the fluent sample. Make your own sponsor show the owner and the baseline. Refuse seat maths that do not state the unit of work.
The companies that look sophisticated in this market will not be the ones with the most logos. They will be the ones whose contracts describe the workflow that actually runs on the Tuesday after go-live.
References
About Nimbus
Nimbus is a Collaborative AI Operating System built around four core pillars that bring human teams and autonomous AI together into a single, unified workspace.
Communication: Keep context tied to the job. Unify emails, meeting recordings, transcripts, and operational files directly within active projects—ending knowledge silos buried in private inboxes, scattered Slack threads, or unrecorded calls.
Collaboration: Work alongside AI in real time. Bring people and AI agents onto the exact same brief, visual canvas, or initiative. Query company-wide data, invite agents into live calls, and co-create in one shared space—eliminating the split between human group chats and isolated AI sidebars.
Automation: Put routine workflows on autopilot. Connect more than 2,000 enterprise tools and standardize repetitive operations. Background loops run on schedules or data triggers with full execution logs, ensuring operational knowledge is shared across the team rather than trapped in one person’s head.
Governance: Deploy AI with absolute control. Enforce strict role-based access controls across workspaces. AI agents can analyze, summarize, and draft—but no live system changes or external communications occur without explicit, verified human sign-off.
The risk arrives after the RFP
What does a typical AI RFP miss?
It asks for features. The risk arrives later as a connector, a retention clause, and a workflow procurement has not seen run.
What should procurement ask to see?
A real job: which system changes, who signs, what is retained, and what happens when the vendor is removed.
Why is a demo a poor proxy?
A demo shows the happy path. Operating risk is the path where the assistant is allowed to act.
See what governed AI looks like on your stack.
Connect your tools, run a workstream, and keep every decision on your ledger. Start on Free.