Why does the speed show up in the tool and not in the numbers?
AI makes a task faster and leaves company throughput almost exactly where it was, and the reason is not model quality. Throughput is set by the handoffs around the work — the queue in front of it, the approval after it, the re-explanation between them — and none of those get shorter when one person types less.
I have watched this shape repeat often enough that I now expect it. A team adopts an assistant, individuals report saving five to eight hours a week, the people reporting it are not exaggerating, and six months later the operations dashboard shows the same cycle time, the same backlog and the same headcount. The gain was real. It was also local, and the only numbers leadership reads are not local.
AI shows up as individual speed long before it shows up as organisational throughput, because a task is bounded by one person's effort while throughput is bounded by the handoffs between people — so making the work faster moves nothing until the handoffs move.
What does the evidence actually show, and what does it not?
Two studies are worth knowing precisely, because both are usually cited loosely.
The first measured a role that owns its work end to end. Brynjolfsson, Li and Raymond, in NBER working paper 31161 (2023), found that giving customer support agents an AI assistant raised issues resolved per hour by about 14 percent on average, with the largest gains among less-experienced agents. That is a throughput result inside a single role, with no upstream dependency and no downstream approval: the agent opens the ticket and closes it.
The second measured people working inside a system they could not immediately verify. METR's July 2025 randomised trial put 16 experienced open-source developers on 246 real tasks in repositories they knew well. With early-2025 AI tools they were about 19 percent slower, while reporting that they had been roughly 20 percent faster. Small sample, unusual population, tools that are now a year old — and the direction of the gap between belief and measurement is the interesting part.
Together they bound the argument rather than settle it. Where one person owns a whole task and sees the result immediately, gains are real and measurable. Where the output disappears into someone else's queue, perceived speed and measured speed can point in opposite directions. Robert Solow's 1987 line — you can see the computer age everywhere but in the productivity statistics — was about an earlier technology with the same accounting.
Where does the saved hour actually go?
Three places, and none of them is throughput.
It becomes slack. The person finishes at 4pm instead of 5.30pm, which is a good outcome for them and invisible in any process metric.
It becomes inventory. More drafts are produced and they wait in the same queue they waited in before, so the queue grows rather than shrinks.
It becomes someone else's work. A draft still has to be read, checked, corrected and approved. If the drafting output doubles and the reviewing capacity does not, the reviewer is now the ceiling, and each item gets less attention than it did.
Which handoffs are eating the gain?
Four kinds, and they behave differently. The column that matters is the last one: each handoff has a signature that shows up in data you already own.
| Handoff | What it looks like in the process | Why a better model does not shorten it | The number that exposes it |
|---|---|---|---|
| Queue | Work sits in a shared inbox or ticket queue waiting for anyone to pick it up | Waiting is not labour, so no assistant reduces it | Hours in queue per case, from creation and pickup timestamps |
| Translation | The request arrives without the reasoning behind it, so a human re-explains context on a call or in a thread | The model can draft the explanation, but only from a rationale someone wrote down once | Clarification rounds per case, and the hours they consume |
| Approval | A named person with signing authority has to accept the output | Authority is a governance decision, not a text-generation problem | Hours waiting on that one person, and their queue depth at end of day |
| Rework | Work comes back with corrections and re-enters the flow | Faster drafting raises the number of items entering the return loop before it raises the number leaving | Return rate, plus calendar days spent inside a second revision cycle |
Most process maps show the four queue-shaped handoffs as arrows. An arrow is drawn as instantaneous, which is why the improvement case always looks better on the map than it does in the month.
How do you compute the ceiling in an afternoon?
With arithmetic on stated assumptions, not a benchmark. Take a process with eight steps across four roles, 1,000 cases a month and 14 calendar days end to end. Measure two things from system timestamps rather than interviews: hands-on work per case, and the wait between role transfers. Say hands-on work is 4.5 hours spread across the four roles, and 11 of the 14 calendar days are waiting.
Now apply AI to the largest single work step — two hours of drafting — and assume a generous 50 percent cut. One hour of labour is recovered per case. Nothing about the queues, approvals or the return loop has changed.
| Measure | Before | After halving the largest task | What it means |
|---|---|---|---|
| Hands-on labour per case | 4.5 hours | 3.5 hours | 22 percent less work |
| Calendar days per case | 14.0 | 13.4 | The saved hour sits in one station; the waiting is untouched |
| Cases per month | 1,000 | 1,000 | The constraint was never the drafting step |
| Cost per case at a loaded $65 an hour | $292 | $227 | Capacity, not cash, until the constraint role absorbs it |
The $65 comes from a standard payroll convention rather than anything I measured: base salary times roughly 1.3 for employer costs, divided by 2,080 paid hours, then divided again by utilisation, because close to a third of a paid day goes to meetings and context switching. Use your finance team's multiplier instead of mine. The shape is the point. A 50 percent cut to a step that holds 22 percent of the labour and none of the waiting produces a 22 percent labour reduction and a 4 percent improvement in the number the customer actually experiences. If you want the fuller mechanism, I have written separately about why the calendar days sit in the handoffs between roles rather than in the work itself, and it is the piece I hand to whoever owns the process map.
Why can a faster step make the whole process slower?
Because drafting and reviewing are two stations on one line, and the output of a line is set by the slower station plus the amount of work sitting between the stations. Little's law is the whole argument in one line: throughput equals work in progress divided by cycle time. There are only two ways to raise throughput — shorten the cycle time, or reduce the work in flight — and neither is what an assistant does by default. Speeding one station raises the arrival rate into the next, which raises work in progress, which lengthens cycle time.
The compounding version is worse. When a reviewer faces 40 drafts cleared at once, they batch: they read faster, skim, and approve on pattern rather than content. Returns go up, and a returned item pays the waiting cost a second time through the same queue. I have watched a content team roughly double its drafting output with no change in calendar time to publish, more revision cycles, and a heavier load on the single editor who could approve anything. The writers were faster. The process was not. The tool had made the constraint's job bigger.
So is the productivity paradox real?
The aggregate number is real and the conclusion usually drawn from it is wrong. The paradox is not evidence that AI does not work; it is evidence that firms measure the unit the tools act on — a task, a document, a ticket — while throughput is produced by the units the tools do not touch, which are transfers of work between people with different authority and different queues. Individual speed is a leading indicator, not an outcome, and a leading indicator that never converts looks identical to a failed investment from the finance seat while feeling like progress on the floor. That mismatch does the damage. The people closest to the tool become advocates, the people closest to the numbers become sceptics, and both are reading their own instrument correctly.
What I cannot give you is a decomposition of the aggregate gap: how much of it is handoffs, how much is absorption into slack, how much is measurement that never captured the gain in the first place. Firm-level data on that split is thin, and a precise percentage quoted for it is almost always a survey rather than a measurement. What I can give you is a decomposition of your own process, and the instrument for that is a spreadsheet and a week.
What to do before the next quarterly review
Pick one process, take twenty closed cases, and write the timestamp of every transfer between roles in a single spreadsheet, approvals included. That is the whole instrument, and it shows you the ceiling before you spend anything on a pilot. Then set the pilot's success metric as calendar days per case at a fixed volume, name the constraint role in writing, and pre-commit what happens to the freed hours: they go into the constraint role, or they do not become throughput at all. If you cannot move anyone into the constraint, do not promise a throughput gain — fund the work as a quality and morale investment, which is often the honest and still worthwhile case.
Keep reading
- Enterprise AI Adoption Fails on Process Ownership, Not Model Capability2026-03-287 minAI Adoption
- Why AI Pilots Fail in Production: The Demo Had a Babysitter, the Rollout Did Not2026-03-267 minAI Adoption
- Design AI to Propose, Not Commit, and Most Compliance Risk Disappears2026-03-048 minAI Adoption
- Eleven Gates Between an AI Pilot and Production, and Only Two of Them Are Technical2026-02-088 minAI Adoption
- Shadow AI Is Already in Your Company: Stop Asking Whether to Adopt It2026-02-248 minAI Adoption