Why do AI ROI calculations fail in finance review?
I have watched the same proposal die in three different companies, and never once because the technology was wrong. It died because the benefit line said "time saved", the baseline was a manager's estimate, and the cost side stopped at the licence fee. A reviewer who owns a P&L does not need to be a sceptic to reject that. They only need to be competent.
An AI ROI number survives finance review when every benefit is either cash already on the ledger or explicitly labelled capacity, every figure is measured against a dated baseline, and the cost side states a per-run floor that still holds at ten times the volume.
That is the whole test. Everything below is the mechanics of passing it, and none of it requires a working system. All three gaps can be closed with two weeks of measurement and one spreadsheet, before the proposal is written.
| The gap | What the reviewer does with it | The artifact that closes it |
|---|---|---|
| Benefit stated as hours saved | Discounts the whole line to zero, or asks which role is being removed | A benefit ledger splitting cash from capacity, with a named account or role for every cash line |
| No measured baseline | Asks "compared to what?", defers a quarter, the sponsor loses momentum | A dated, signed measurement: monthly volume, touch time, exception rate, cost per transaction |
| Cost stated as a licence or build price | Asks what happens at five times the volume, gets no answer, rejects the case on risk | A per-run cost model with exception handling priced separately from fixed platform cost |
When does saved time become cash?
Apply three tests in order, and be strict about the order, because the third is where sponsors talk themselves into a figure they cannot defend later.
First: does the change remove a hire that is already approved and budgeted? Then the benefit is cash with a date, which is the start date of the role that no longer starts. Second: does it reduce spend already sitting on the ledger, such as overtime, agency temps, or outsourced processing? Then it is cash now, and the invoice trail proves it. Third: does it release capacity in a queue that is already sold, such as a billable backlog or a throughput limit with penalty clauses attached? Then it is revenue, which is a different and riskier line. Model it separately and discount it, because revenue depends on demand you do not control.
If none of the three applies, the benefit is capacity, not cash. Keep it in the model, label it "capacity released, non-cash", and value it only as an option: the hire you do not make if volumes rise thirty per cent next year. Reviewers trust that line. What they do not trust is a cash label stuck onto it, because it is the first thing they will test and the easiest thing to disprove.
Here is the arithmetic on a process I scoped last year, with every input stated so you can substitute yours. An accounts payable team processes 4,200 invoices a month at 3.5 minutes of touch time each. The clerk's fully loaded cost is $62 an hour, which comes from an $85,000 salary plus 28 per cent payroll overhead, divided by 1,760 productive hours after leave and public holidays — a controller can check that against payroll in five minutes. Gross hours: 4,200 × 3.5 ÷ 60 = 245 hours, or $15,190 a month.
That $15,190 is the gross benefit, and it is not the number that belongs in the proposal. The number that belongs in the proposal is $15,190 minus the cost floor, and the cost floor is where most cases quietly fall apart.
How do you build a baseline that survives a controller's questions?
Measure, do not ask. Self-reported time is inflated by the same instinct that inflates project estimates, and asking people to estimate their own efficiency is asking them to argue against their own job security. Pull timestamps from the system of record — ERP, ticketing, the CRM audit log — instead of running a stopwatch survey.
Ten business days, or 200 transactions, whichever comes first, in the same queue, same team, same period, with no process changes in flight. Report the median and the ninetieth percentile rather than the mean: the mean hides the tail, and the tail is where the cost lives. A process whose median touch time is two minutes and whose ninetieth percentile is eleven minutes has a very different automation economics from one with a flat distribution, and the average cannot tell you which you have.
Then freeze it. One page, dated, stating monthly volume, median touch time, exception rate, and cost per transaction, signed by the process owner. A baseline without a date and a signature is an opinion, and opinions get renegotiated the first month actuals diverge from plan. Re-measure ninety days after go-live. If you skip that, you end up comparing a quiet February against a December close and then explaining to finance why the benefit evaporated.
The one figure most teams omit from the baseline is the exception rate: the share of transactions a human has to touch even after automation. That number decides the economics more than the model choice does, and you cannot estimate it after the fact without looking like you are reverse-engineering the result toward the answer you wanted.
What is a cost floor, and why does its absence end the discussion?
A cost floor is the monthly cost of running the process with every scaling component priced per unit. Four components belong in it: per-run inference and orchestration; exception handling labour; fixed platform, storage, and monitoring; and amortised build cost. Only the first two scale with volume. The third steps up in chunks when you cross a threshold. The fourth is sunk, but finance will want it in the payback calculation regardless.
For the same process, at 4,200 runs a month:
| Cost line | Monthly | Basis |
|---|---|---|
| Inference and orchestration | $88 | $0.021 per run, including retries and failed attempts |
| Exception handling | $3,125 | 12% of 4,200 = 504 exceptions, 6 minutes each, $62 per hour |
| Platform, storage, monitoring | $1,200 | fixed below roughly 20,000 runs a month |
| Amortised build | $2,000 | $72,000 over 36 months |
| Total cost floor | $6,413 | marginal cost of the next transaction: about $0.76 |
Net monthly benefit is $15,190 minus $6,413, which is $8,777, and that pays back the build in 8.2 months. Those figures hold up because every input is either measured or stated outright.
Then show the reviewer what breaks, because that is the part that buys credibility. Exception handling is 49 per cent of the cost floor here, so this workflow's economics are decided by the exception rate, not by the model. If that rate doubles to 24 per cent, the floor rises to $9,538 and payback stretches to 12.7 months. Still fundable, and saying so before anyone asks is worth more than a better projection. If you want the per-run figure to hold under that kind of stress test, the mechanics are in a per-token cost model that prices caching, retries, and failures into a single number.
What belongs on the single page you bring to the review?
Six lines, and no more. The baseline, dated and signed. Gross benefit split into cash and capacity. The cost floor, with the marginal cost per transaction stated explicitly. Payback in months. Sensitivity on the two variables you do not control, which are exception rate and volume. And the conditions that would falsify the case.
Name those conditions yourself. If measured touch time falls below two minutes, the gross benefit cannot carry the build. If the exception rate exceeds 25 per cent at month three, stop and re-scope rather than defend a number that has already moved. A sponsor who publishes kill criteria in advance is a sponsor whose projections get believed, because they have demonstrated a willingness to be wrong in public. The opposite posture — protect the number at all costs — is what taught your CFO to discount the next AI proposal on sight.
Why the second review is the one that kills the case
Most AI cases survive the first approval and die at the second, when actuals arrive and nobody can explain the gap between the plan and the ledger. The three gaps are not analytical failures so much as sequencing failures: the baseline and the cost floor have to be measured before the number is committed to, and by the time the proposal is written the sponsor is usually attached to a figure. Two weeks of timestamp data and one honest column labelled capacity, non-cash will not make the case larger, but they will make it the case that gets funded and then keeps its funding. Build the smaller number you can defend, because the larger one you cannot defend ends up costing more in lost credibility than the delay ever would have.
Keep reading
- The Cost of Delaying AI Adoption Is a Run-Rate, Not a Purchase Price2026-03-247 minAI Economics
- Your AI Project Does Not Need More Data, It Needs a Schema2026-02-267 minAI Economics
- Seven AI Vendor Procurement Questions, and Why the Failure Path Is the Decisive One2026-02-168 minAI Economics
- LLM Token Cost per Workflow Needs Three Terms, and Retry Rate Is the One People Omit2026-03-279 minAI Economics
- AI Vendor Lock-In Is the Cheap Kind: Your Prompts and Evaluation Set Are Not2026-01-198 minAI Economics