Evaluation

How to Evaluate Collaborative AI

An RFP sheet for shared jobs — not copilots. Questions on roster, stored rejection, second department, files in one place, and what survives after the session. Plain language for operators.

Evaluating collaborative AI means scoring whether several people can share one job — same files, same history, a named person who can stop a change — and whether that job still exists after the session ends. McKinsey’s State of AI (2025) found that 88% of organisations use AI in at least one function while most remain in experiment or pilot. A common pilot is one person and one assistant. The next purchase mistake is rebranding that pilot “collaborative” because five people share a login. This checklist is written to catch that false pass early, on any vendor.

It is not the same as evaluating a copilot, an agent swarm, or a shared ChatGPT account. What is collaborative AI is the definition. Collaborative AI and personal assistants is when to use a copilot versus a shared room — read that first if the purchase question is still “do we need both?” This page is the RFP sheet for the shared job: questions to put in procurement, demos, and proof-of-value scripts.

How to evaluate an agent harness is the runtime sheet — refused writes, replay, model swap. This page is the job sheet — roster, rejection, handover. Run both when the work crosses departments and may write to a live system. Do not run only the one the vendor prefers.

NIST’s AI RMF Playbook is the measurement language. Translate it to demos: Map the job, Measure reopen time and signer completeness, Manage with stored rejections before write tokens. “The model seemed careful” is not a measure.

What you are actually buying

You are buying a container for shared work, not a better paragraph generator.

Minimum properties:

  • A named job, not “the Slack channel.”
  • Shared context — files and history visible to the roster.
  • Tools with recorded steps — read, propose, sometimes write with a payload.
  • A named stop — someone other than the prompter can refuse a change.

Multiplayer AI vs multi-agent AI separates people in the room from models in a loop. You can fail collaborative eval while passing multi-agent demos: agents pass tickets to each other; finance still cannot see the brief. What is multi-agent AI is the cast. This sheet scores the stage.

Four pillars of an enterprise AI platform situates workstreams and governance in a full stack. This sheet scores whether the product you are viewing implements the collaborative pillar for one real job — Salesforce sidebar, Microsoft copilot, Slack bot, specialist agent platform, or anything else on the shortlist.

Anthropic’s guidance on building effective agents says encode the job, bound the tools, define done. Your proof of value should force those three in a room with two departments, not in a solo sandbox.

Yang and colleagues, in Nature Human Behaviour (2022), showed that firm-wide remote work made collaboration networks more static and siloed, with fewer bridges between groups. Shared jobs already fight that pull. A vendor that adds a fluent assistant to each silo will make the silo more confident, not more shared. Score the bridge.

RFP questions: shared job vs shared login

Put these verbatim in the RFP. Require live answers, not slides. If professional services will “configure that later,” note the time-to-value and price the configuration as part of the buy.

1. Is this a shared job or a shared login?

Pass: One work object with its own roster, files, and audit — independent of which user opened the UI today.

Fail: Five people in one chatbot account, or five parallel threads that cannot see each other’s attachments.

Proof: Show two users on the same job ID. Remove one user’s access. The job remains for the roster.

Why vendors fail this: shared seats are cheap to demo and expensive to govern. Collaborative AI and personal assistants explains why a shared login is still a personal-assistant shape. Ask for the job ID in the URL or export. If the vendor cannot point at an object, they pointed at a session.

2. Can a second department join mid-run?

Pass: Finance joins Thursday’s sales exception without a re-upload parade. They see the same CRM excerpt, the same draft payload, the same history.

Fail: “Export and email the transcript.” “Start a new session and paste context.”

Proof: Add a finance delegate mid-proof. They reject a proposal while sales watches. No side channel required.

This is the core multiplayer AI test dressed for procurement. Do not accept a pre-seeded “war room” that was built overnight by the vendor’s solutions team unless you can repeat the join on a job your people created.

Microsoft and LinkedIn’s 2024 Work Trend Index found that 78% of AI users bring their own tools. Mid-run join is how you find out whether the product can absorb that habit or whether finance will open a second private window.

3. Is there a named stop — not “human review” in the abstract?

Pass: A person on this roster can halt this class of write while others on the job see the payload.

Fail: A generic approval workflow outside the job, or a prompt that says “ask manager.”

Proof: Attempt a live-system change. Show the signer field tied to a human identity. Show a stored rejection with name and timestamp on the job.

What is human-in-the-loop AI and write-back governance define the stop. This question tests whether they are on the job. A ServiceNow ticket opened after the write is not a stop. A Slack reaction is not a signer.

OWASP’s LLM Top 10 lists excessive agency and insecure output handling. Collaborative eval adds organisational agency: who could have stopped this change on this job. If the only stop is max tokens, you have a fuse, not a control plane.

4. Are files in one place, not five inboxes?

Pass: Attachments live on the job — WMS snapshot, ageing extract, partner PO — visible to the roster without re-forwarding.

Fail: “Paste into the chat window.” “The model will fetch from SharePoint if you paste the link.”

Proof: List attachments on the job object. Remove the original uploader from the roster. Files remain.

Search is not memory is why “we’ll find it in Slack later” fails eval. Do not let the vendor substitute a retrieval demo for an attachment that survives user removal. Institutional memory in enterprise AI is the company-scale layering; this question only asks whether this job still has its two files on Monday.

5. What happens after the session?

Pass: Monday reopen shows brief, files, last proposal, last rejection or signature — even if the original prompter is out and the model vendor changed.

Fail: Session expiry deletes context. “Memory” is the user’s personal thread.

Proof: Close the browser. Reopen with a different user. Continue the job without reconstruction.

Agents should be disposable is the design claim this question tests. What an AI workstream is is the container name. If continuity requires the original model or the original person, you scored a session, not a job.

ISO/IEC 42001 wants records and named actors. A session that evaporates is not a record. Ask for an export a later reader can use without vendor professional services.

6. Does governance live in the room?

Pass: Roster, inherited authority, spend-cap pause, and notify-to-roster on this job — see governance as a multiplayer primitive.

Fail: Governance PDF emailed quarterly while writes succeed unsigned.

Proof: Show spend-cap alert to roster members only. Show agent grants bounded to the acting person’s authority. RBAC for enterprise AI is the access vocabulary; demand it scoped to the job.

A tenant-wide “AI policy acknowledged” checkbox is not this test. Neither is an SOC 2 report. Those are programme artefacts. This question is whether finance sees the same payload ops sees before execute.

Scorecard: false passes to reject

Demo looks likeLikely false passAsk instead
Fluent multi-user chatShared loginJob ID, roster, survive user removal
Agent orchestraMulti-agent without multiplayerSecond department join + named stop
Copilot in CRM sidebarPersonal assistantCross-team exception with finance on job
“We integrate Salesforce”Connector without payload quoteShow field-level payload before write
Long context windowMemoryReopen Monday without original thread
“Human review” nodeAbstract HITLNamed signer + stored rejection on job
Quarterly attestationPDF after the writeFail-closed unsigned attempt

NIST’s AI Risk Management Framework Measure function assumes you can observe outcomes. A demo that cannot produce a stored rejection has nothing to measure except fluency.

Proof-of-value script (two weeks)

Do not let the vendor script a happy-path email draft. Use one exception you already run.

Week one — read-only, multiplayer:

  1. Name one real exception: credit hold, discount outside grid, order hold in WMS.
  2. Put ops and finance on the roster. Attach two files they already email.
  3. Mid-week, add a second-department guest with a scoped view.
  4. Model drafts release; finance rejects; rejection must stay on the job.

Week two — continuity and optional write:

  1. Swap the prompter. Reopen. A stranger continues without Slack archaeology.
  2. Swap model tier or vendor if the product claims portability.
  3. If write is in scope: enable fail-closed write with payload quote; show unsigned attempt failed.
  4. Independent reader reconstructs signer and payload without authors in the room.

If steps 4 or 5 fail, you do not have collaborative AI — you have a group chat with AI autocomplete. Stop the proof. Do not “save write for phase two” as a way to skip the rejection test. The rejection is the point.

Function-specific walkthroughs if you need a scenario library: operations, revenue operations, customer support, human resources, finance and planning. Use one. Do not run six proofs.

How this differs from harness evaluation

How to evaluate an agent harness asks:

  • Can the harness refuse a write?
  • Can you replay signer and policy version?
  • Can you swap the model without rewriting tools?

Collaborative eval asks:

  • Can two departments see the same refusal?
  • Does the job survive people and agents leaving?
  • Are files and stops on the job, not in personal threads?

You need both when the purchase is “AI for cross-team exceptions that may touch CRM, ERP, or WMS.” Harness without collaborative passes produces a gated write finance never saw coming. Collaborative without harness passes produces a shared room where unsigned writes still slip through.

Agent harness vs agent framework is the build-versus-compose warning: a graph in a notebook is not a passed test. Inner vs outer agent harness is why a coding-harness scorecard will mis-score an outer operations job. Use the right sibling sheet. Do not grade a warehouse hold with SWE-bench.

What is AI governance is the programme within which both sheets fit. Collaborative AI is not a feature checkbox. It is whether shared work survives the people and models that staffed it this week.

Vendor questions to copy into procurement

Number them. Require a live show, a recording, or a written fail.

  1. Show one job ID with two departments and different tool grants on the same roster.
  2. Show a rejection stored on the job with signer identity — not an email log.
  3. Remove the original uploader; attachments remain.
  4. Add a guest; guest cannot inherit write token.
  5. Pause on spend cap; notify roster members; show delegate approval on the job.
  6. Reopen after 72 hours with a different model; show continuity.
  7. Attempt unsigned write to a system of record; show fail-closed.
  8. Export replay — signer, payload, policy version — without vendor professional services.

If the vendor answers questions 1–6 with “our SI will configure that,” treat configuration time as part of the price. A product that needs six months of graph work before a second department can join is a framework purchase, not a collaborative-AI purchase. Say so in the scoring notes.

Ask incumbents the same questions you ask specialists. A CRM copilot that cannot put finance on the job fails this sheet even if it drafts beautiful emails. A multi-agent platform that cannot store a rejection fails this sheet even if the orchestra is elegant.

When collaborative eval is not the first buy

Stay with personal assistants when:

  • One owner drafts and nothing writes to a live system.
  • No second department must stand on the result this quarter.
  • The pain is blank-page speed, not lost outcomes.

Collaborative AI and personal assistants is the decision tree. Buy collaborative when the recurring meeting already exists and the pain is “we never keep the outcome.”

Do not force a roster onto solo research. Theatre rosters teach people that governance is ceremony. Save the sheet for jobs that already have a fight in chat.

How to run the bake-off

Score every shortlisted product on a cross-team job, not the AE’s private email draft. Use the same exception, the same two files, the same two departments. Keep a shared scorecard with pass/fail per question above — not a 1–5 “wow” rating on fluency.

Include at least one incumbent copilot and at least one agent platform if both are on the table. The point of the sheet is comparison, not a single-vendor script. What is an enterprise agent harness and what is an enterprise AI operating system are category pages if you need language for the stack around the job. This sheet still scores the job.

Put this sheet in the RFP, then run it in a proof of value. A fluent demo without a stored rejection is still a slide.

How this shows up in Nimbus

When you include Nimbus in a bake-off, run this same sheet on workstreams and governance. Do not substitute a homepage video or a pre-built demo room.

Ask for the stored rejection, the mid-run finance join, the reopen after a model swap, and the unsigned write that fails. Score incumbents on the same exception. Staffing and time-to-value for the bake-off itself are a separate frame — see self-service vs forward-deployed.

If Nimbus cannot show the “no” on the job Monday, it fails this sheet the way any other vendor would. Collaborative eval is not a product tour.

Common questions

Score the shared job

Is this the same as evaluating a copilot?

No. Copilot evals score solo drafting — latency, paragraph quality, seat SSO. This sheet scores whether two departments share one job, with a named stop and artefacts that survive the session. A product can pass every copilot test and still fail the moment finance joins Thursday’s exception. If your RFP only asks context-window size and brand of model, you will buy an assistant and call it collaboration. Read the personal-assistant comparison first if you are still deciding whether you need a shared room at all.

Do we need multiplayer and collaborative AI in the RFP?

Ask both. Multiplayer is who is in the room now — can a second department join, see the same brief, and reject a proposal while the work is happening. Collaborative is whether Monday still has the files, the rejection, and the signer after everyone hangs up. You can fail one and pass the other. Five people in a call with nothing written down is multiplayer for an hour. A job two teams open on different days can be collaborative without a live huddle. Cross-department work that may write to a live system needs both.

What is the fastest proof-of-value test?

Invite a second department mid-run, reject a proposal, swap the prompter, reopen Monday. If the job is empty, you evaluated a shared login. Do not let the vendor pick a solo email-draft scenario. Name a real exception you already fight in chat. Require a stored rejection with a name and a timestamp before anyone enables write. An independent reader who was not in the demo should reconstruct signer and payload without Slack. If they cannot, fail the proof even if the paragraph was fluent.

Can we reuse our agent-harness RFP instead?

Reuse it as a sibling, not a substitute. The harness sheet scores refused writes, replay, and model swap on the runtime. This sheet scores roster, handover, and whether files live on the job. You need both when the purchase is AI for cross-team exceptions that may touch CRM, ERP, or WMS. Harness without collaborative passes produces a gated write finance never saw. Collaborative without harness passes produces a shared room where unsigned writes still slip through. Copy both question lists into procurement.

See what governed AI looks like on your stack.

Connect your tools, run a workstream, and keep every decision on your ledger. Start on Free.