A client described the moment precisely. Six weeks after rolling out an assistant across a forty-person team, the licence dashboard was healthy and nothing had been reported as broken. Then a manager said, almost in passing, that she had stopped using it for anything that mattered, because she had to read the output as carefully as if she had written it herself. She had not complained, because nothing had gone wrong. That is the failure mode. It leaves no trace.
The cost nobody put on the invoice
The bill for an AI tool has three lines. There is the licence, which finance can see. There is the training and the setup, which someone estimated once. And there is verification overhead: the time your people spend establishing whether what came back is true.
Only the third one varies with quality, and only the third one is invisible. It does not appear in a benchmark, because benchmarks measure whether the model was right, not how long it took a person to find out. It does not appear in usage analytics, because a person reading an output carefully looks exactly like a person reading an output.
The arithmetic is unforgiving. If a task takes twenty minutes unaided, and the assistant produces a draft in one minute that takes eighteen minutes to check properly, you have saved a minute and acquired a new risk. If it takes twenty-two minutes to check, which happens whenever the errors are subtle rather than obvious, the assistant is now costing you time and you are paying a licence fee for the privilege.
Why 95 per cent can be worse than 70 per cent
This is the part that surprises people, and it is the most useful idea in this article.
Suppose a tool is wrong three times in ten, and wrong in ways that are obvious. You glance at the output, you see the nonsense, you fix it or you redo it. Verification is fast, because the errors announce themselves. Your trust calibrates quickly: you know what it is bad at, and you route around it.
Now suppose a tool is wrong once in twenty, and the errors are plausible. A number that is the right shape but the wrong value. A citation to a regulation that exists, with an article number that does not. A summary that is accurate about everything except the one clause that mattered. You cannot tell which one is the bad one without checking all of them, so the verification cost of the whole set is driven by the undetectability of the errors, not by their frequency.
The first tool is honest about its limits and cheap to supervise. The second is more accurate and more expensive to use. Accuracy is not the variable that determines value. Accuracy plus the cost of finding the exceptions is.
This is why teams sometimes abandon a better model for a worse one, and why the abandonment looks irrational from the outside. It is not irrational. They are optimising for the total, and only they can see the total.
The selection rule that follows
If verification is the binding constraint, then the right question when choosing where to deploy AI is not "what is the model good at". It is: for which of our tasks is checking an answer much cheaper than producing one?
That asymmetry is where the value is, and it maps onto real work quite cleanly.
- Strongly asymmetric, deploy here first. Drafting from a template you will read anyway. Translating text a bilingual colleague will skim. Writing a query whose result you can sanity-check. Classifying items where a wrong label is visible at a glance. Finding candidate passages in a long document, where you then read the passage. Transcription you will correct while listening.
- Weakly asymmetric, deploy with care. Summarising a document nobody else will read, because the only way to check the summary is to read the document, which was the job.
- Inverted, do not deploy. Anything where confirming the answer requires the same expertise and roughly the same time as producing it: legal conclusions on unusual facts, a diagnosis, a figure that goes into a filing, a security judgement about a specific system. Here the assistant does not remove work, it relocates it and adds a plausible-sounding anchor that makes the reviewer lazier.
The uncomfortable corollary is that the most impressive demonstrations are usually in the third category, because that is where the work looks hardest and therefore the automation looks most valuable. The boring first category is where the money actually is.
Trust is a property of the team, not the person
Individual users calibrate fast. Teams do not, and this is where the real damage accumulates.
When a colleague is burned once by a confident wrong answer, they tell three people at lunch. The story is much more memorable than the eighty times the tool was fine, and it propagates faster. Within a month you have an informal reputation, and informal reputations are sticky: they persist after the tool improves, and they are almost never revised, because nobody re-tests a thing they have already decided about.
There is a second-order effect that matters more for anyone selling or building. Disengagement is silent. There is no error message that reads "the user gave up", no support ticket, no churn signal until renewal. The tool is still licensed, still installed, still occasionally opened for trivial things so that the usage graph looks alive, and it has been quietly removed from every workflow that carries consequence.
How to measure something invisible
Time one real task, twice. Pick three people and three genuine tasks. Do each once with the tool and once without, and include the checking in both. This takes an afternoon and produces a number your finance director can use. Most organisations have never done it, and a surprising share discover that their flagship use case is a wash.
Watch usage by the hard cases, not usage in aggregate. Total prompt volume rises for months after a rollout because people experiment. The signal that matters is whether the people with the most demanding problems are still using it in month three. If your senior people have drifted away and your juniors have not, you do not have an adoption success, you have a supervision problem.
Ask the question directly, and make it safe to answer. "Do you trust it?" gets you a polite answer. "When did you last decide not to use it, and why?" gets you the truth, because it presumes the behaviour rather than judging it.
Log the corrections. If output flows through any system you control, keep the before and after. The edit distance between what the model produced and what shipped is the closest thing to a direct measurement of verification overhead that most organisations can get cheaply.
What to do about it
Three things reduce verification overhead, and none of them is a better model.
Make the tool show its working, so that checking can be targeted rather than total. An answer with the source passage attached can be verified in fifteen seconds; the same answer without it takes five minutes. Narrow the task until the output is checkable at a glance, because a well-scoped assistant that does one thing verifiably beats a general one that does everything provisionally. And be explicit, in writing, about what it is not for, since a documented boundary is what stops the informal reputation from generalising from one bad experience to the whole tool.
Whether you are buying an assistant or building one, the question is the same, and it is not how capable it is. It is whether it saves time, or merely moves the time into checking it. Verification overhead is the real bill, and it arrives whether or not anyone is measuring it.
This article grew out of a Focus note from September 2026 on the hidden cost of a bad assistant. Related reading: Before You Adopt AI and The Private AI Assistant.