What separates an AI vendor who has shipped from one who has demoed?
A demo is a rehearsal with a babysitter present. Procurement is where you find out whether the system has ever run on an ordinary Tuesday, against your data, with nobody from the vendor in the room.
I have sat on both sides of these evaluations, as the person building the workflow and as the invited second opinion on someone else's shortlist. The vendors who survive diligence are rarely the ones with the loudest model claim. They are the ones who answer questions about their own failures without flinching, because they have seen those failures often enough to have a procedure for them.
Seven questions carry most of a first-round evaluation. Six are ordinary diligence that any competent buyer would ask anyway. The seventh decides the deal.
Ask a vendor to walk you through a real case their system got confidently wrong — what the operator saw on screen and what happened next — and you will learn more about whether they have shipped than from a full day of happy-path demos and three reference calls combined.
A vendor who has only demoed answers that request with a slide about human oversight. A vendor who has shipped answers it with a screenshot of the escalation queue and a case reference.
The seven questions, and what each answer actually tells you
| Question | What a demo-stage vendor says | What a shipped vendor says |
|---|---|---|
| 1. Show me a case you got confidently wrong on data like ours. What did the operator see? | A slide on "human in the loop" | A screen, a case, and who resolved it |
| 2. What is the exception rate on this process type, measured over a full month? | A benchmark figure with no denominator | A number, its denominator, and the period it covers |
| 3. What do the next 10,000 runs cost us at list price? | A seat price or a licence tier | A per-run figure including retries and failed attempts |
| 4. Which customer runs this in production with none of your engineers present? | Logos | A name who will take the call |
| 5. What has to change in our process before your software works? | "It is configurable" | A list of process changes, each with an owner |
| 6. What leaves our environment, to whom, and for how long? | A compliance badge | A subprocessor list, retention periods, and the region data lands in |
| 7. If we stop paying, what do we keep? | "Your data" | Export formats, the evaluation set, and who holds the configuration |
Questions two through seven are checklist items. Question one is a diagnostic, and the reason is structural: a happy-path demo is selected by the vendor, so it carries almost no information. A failure path is selected by reality.
Why is the failure path the decisive question?
One answer to that question exposes five things at once.
It shows whether the system can tell that it does not know — whether there is abstention, confidence routing, or a way to hand a case to a human before a write reaches a system of record. It shows whether the operator interface presents the evidence behind a decision or merely the decision, which is the difference between a reviewer checking the work in fifteen seconds and redoing it. It shows whether a per-case record exists, so a wrong outcome is reconstructible six months later when an auditor asks. It shows who carries the cost when a case falls out of the automation. And it produces the exception rate, the number that decides the economics more than the model choice does.
The practical version of this is work you do before the meeting, not during it. Pull sixty to one hundred recorded cases from last quarter, deliberately favouring the ones that took longest, escalated, or were later reversed. Send them to the vendor and ask for a written result per case: pass, fail, or route to a human, with the reasoning attached. Two things happen. You get an accuracy figure measured on your own traffic, and you find out how the vendor behaves when the inputs are not theirs. Some will decline. That decline is an answer, and it costs you a week to obtain.
No vendor can tell you your exception rate before running your cases. Anyone who quotes one to a decimal place from a generic benchmark is guessing, and the guess will be wrong in the direction that helps them. What you can require is the mechanism: which denominator they count against, over what period, and whether exceptions that a human corrects quietly in a downstream queue are counted at all. The quietly corrected ones are the ones that break the business case a year later, because nobody has a number for them and so nobody notices the trend.
What does an exception actually cost?
Here is the arithmetic for a process I scoped, with every input stated so you can substitute yours. Accounts payable, 4,200 invoices a month, a twelve per cent exception rate, and six minutes of human work per exception at a fully loaded rate of $62 an hour. That is 504 exceptions, roughly fifty hours, about $3,125 a month. The inference and orchestration line on the same volume, at two cents a run including retries, is about $88 a month.
The tail is roughly thirty-five times the model cost. The inputs are mine and the ratio is the point: if your exception rate is four per cent rather than twelve, the same shape holds at a third of the size. What matters for procurement is which side of that table the vendor is pricing.
| Cost line | Illustrative monthly, 4,200 runs | Who carries it | Present in the vendor quote? |
|---|---|---|---|
| Inference and orchestration | About $88 | Vendor platform | Yes, folded into a seat or platform fee |
| Exception handling | About $3,125 | Your operations team | Almost never |
| Monitoring and someone on call | Your staff time | Your team | No |
| Re-measurement after go-live | Analyst days | Your team | No |
| Exit and re-implementation | Unknown until asked | You | No |
This is why a per-seat price is not a comparable unit. Seats do not scale with volume; exceptions do. A quote that prices the $88 line and stays silent about the $3,125 line has not misled you deliberately, it has simply quoted the part it controls. Your job in the RFP is to make the other four lines someone's stated responsibility, even when the answer is "you".
What has to change in our process before the software works?
"Configurable" is not a process design. Ask for the list in writing: what the operator must now do that they did not do before, what the upstream team must now do, and which of your existing controls stops being exercised once the system writes directly to a system of record. Every workflow I have shipped changed the process around it, usually in a small way nobody predicted during the sales cycle — a validation step becoming advisory, a queue being checked once a day instead of continuously.
The useful test is to ask the vendor to name the customer whose process changed the most during implementation, and what changed. A vendor who cannot name one has either never implemented, or is not close enough to their own deployments to know.
What leaves our environment, and who can prove it?
Ask for the subprocessor list, the retention period per data type, and the region data is stored and processed in — in writing, as a contract appendix rather than a link to a trust page that changes without notice. Then ask whether prompts and outputs are used for training, whether that is a contractual commitment or a product setting, and what notice applies if the default changes. Contracts get skimmed by busy people; the subprocessor list is where the risk sits.
If we stop paying, what do we keep?
Ask for export formats, the evaluation set, the configuration and prompts, and any integration code that touches your ERP. If orchestration lives with the vendor, the decision you are making is not a tool purchase but the transfer of your workflow boundary, and the framework I use for a build-versus-buy AI workflow decision is worth applying before the RFP goes out rather than after the renewal.
A vendor who owns the orchestration has a legitimate reason to keep it, and you have a legitimate reason to want the boundary documented. Write down who holds each component and what happens on termination, because ambiguity here is not resolved in your favour later.
What belongs in the RFP itself?
Six mechanics, all of which cost the vendor hours rather than dollars, which is exactly why they filter.
Require each of the seven questions as a numbered response item with a named artifact attached, not narrative prose. Require the evaluation-set run before the commercial proposal is submitted, not as a proof of concept after shortlisting. Require unit pricing at one times, three times, and ten times your current volume, and ask which components are contractual and which are list-price estimates. Require the production reference to be a named person at a customer running the system without vendor staff on site. Put a date on every accuracy claim so it can be re-tested against the same evaluation set after any model change. And state in the document that a refusal to run your failure cases is scored as a failure to respond.
That last line is not a trick. It converts a question vendors can dodge in a meeting into a scoring rule a procurement team can apply without judgement calls, which is why the seven questions survive contact with a shortlist.
Where is the cheapest place to learn this?
The procurement cycle you are already running is the cheapest place to find out whether a system fails safely, because the alternative is discovering it after your team has reorganised its work around software that does not hold. A quote that cannot describe a failure path is not evidence of a working system, only evidence of a rehearsed demonstration, and the two are indistinguishable until you ask. Put the failure path at the centre of the evaluation, score the artifacts rather than the answers, and let the vendors who will not run your worst cases remove themselves. That is a shorter shortlist, and it is a better one.
Keep reading
- The Cost of Delaying AI Adoption Is a Run-Rate, Not a Purchase Price2026-03-247 minAI Economics
- How to Measure AI ROI So the Number Survives Finance Review2026-03-227 minAI Economics
- Your AI Project Does Not Need More Data, It Needs a Schema2026-02-267 minAI Economics
- AI Vendor Lock-In Is the Cheap Kind: Your Prompts and Evaluation Set Are Not2026-01-198 minAI Economics
- The Expensive AI Failure Is the System You Kept, Not the Pilot You Cancelled2026-02-148 minAI Economics