There is a specific failure mode worth naming, because a lot of businesses are living inside it right now: a team adopts an AI tool, everyone agrees it is impressive, and six months later nobody can point to an hour that was saved.
The tool works. The workflow around it does not.
The verification tax
Any system that produces output a human must check has a hidden cost: the time to check it. Call it the verification tax.
If a task takes 30 minutes to do by hand and 5 minutes to do with AI assistance, the naive saving is 25 minutes. But if verifying the AI's output properly takes 20 minutes, the real saving is 5 minutes — and that is before you count the times the output is wrong and has to be redone.
This is why the same technology can be excellent in one place and useless twenty feet away in the same office. It depends entirely on the ratio between doing the work and checking the work.
AI saves time when checking is much cheaper than doing. Summarising a long document: you can skim the summary against the headings in a minute. Drafting a first version of a routine email: you can read it in ten seconds. Extracting fields from an invoice: you can eyeball six numbers against the scan.
AI creates work when checking is nearly as expensive as doing. Producing a financial calculation you have to recompute to trust. Writing code in a system nobody on the team understands well enough to review. Researching facts you then have to independently confirm. In these cases you have not removed the work; you have added a review step to it.
Four patterns that reliably save time
Draft-then-edit. The AI produces version zero, a person edits to version one. This works because editing is genuinely faster than starting from a blank page, and because the human remains the author. It works best for text that follows a familiar shape: proposals, job descriptions, standard replies, meeting notes.
Extract-and-confirm. The AI pulls structured fields out of unstructured input — a PDF, an email, a voicemail transcript — and presents them for a one-glance confirmation. The check is visual and fast because the fields are next to the source.
Search across your own material. The AI finds the relevant paragraph in ten years of internal documents. Verification is trivial: the answer links to the source and you read the source. The value is in the finding, not the answering.
Triage and routing. The AI reads an incoming message and decides which queue it belongs in. Errors are cheap and self-correcting, because a misrouted item gets moved by the person who receives it.
Four patterns that reliably create work
Unreviewable volume. Generating a hundred pages of anything a person is nominally supposed to check. Nobody checks it. The output accumulates unverified, and the first time it matters it is wrong.
Confident numbers. Asking a language model to compute, reconcile or forecast. Numbers look identical whether they are right or wrong, so the verification cost is the full cost of doing the calculation yourself.
Answers with no source. Any workflow where the AI states a fact and there is no link back to where it came from. Every claim becomes a research task.
Automation without a stop. A process that runs end to end with no review point, in a domain where errors are expensive. This does not create work immediately. It creates work in a large lump, later, when someone discovers three months of wrong output.
Designing the workflow, not the prompt
Most disappointing results are workflow problems dressed as model problems. The productive moves are structural:
- Put the source next to the output. If the person reviewing can see what the answer was based on, review takes seconds instead of minutes.
- Constrain the output shape. Free-form text is hard to check. A filled-in form with six fields is easy.
- Choose the review point deliberately. One review at the end of a ten-step process is worse than one review after step three, where the error is cheap to fix.
- Make corrections do double duty. When a person fixes an output, capture the correction somewhere it improves the next run — a template, a rule, an example. Otherwise you correct the same thing forever.
- Measure the end-to-end time. Not the model's time. The time from work arriving to work being done and trusted.
A worked comparison
Two firms adopt the same assistant for customer email.
The first drops it into the inbox and asks staff to "use it where helpful." Staff draft with it, then rewrite most of it because the tone is wrong for their customers, then spend extra time second-guessing whether facts in the draft are real. Net effect: slower, with a vague sense that they should be getting more out of it.
The second spends a week first. They collect their forty most common enquiries, write down the actual answers, and set the tool up to draft only from that material with the source visible. Anything outside those forty goes straight to a person, unassisted. Staff approve most drafts in under fifteen seconds because the answer is drawn from text they already trust.
Same tool. The difference is that the second firm decided what the tool was allowed to not know.
Before adopting an AI step
- Name the task precisely — not "customer service", but "first reply to a delivery-status question"
- Estimate the time to check the output, not just the time to produce it
- Confirm a wrong answer would be visibly wrong to the reviewer
- Make sure the output can be traced to a source
- Decide where the review point sits and who owns it
- Agree what the tool should refuse to attempt
- Set a date to measure end-to-end time and compare it to today
Judge these tools on the total time from question to trusted answer. That number is unforgiving, and it is the only one that matters. Related reading: where AI actually fits and why human approval still belongs in important workflows.