The tasks AI handles badly don't feel any harder than the ones it handles well. That's the whole problem.
Run five questions against any task: can you check the answer, does it need current or proprietary information, how many steps does it take, what does a wrong answer cost, and are you the person who is supposed to be good at this. Score it, and you get one of four answers rather than yes or no: hand over the whole task, take a draft, use it as a checker only, or keep it by hand. And whatever you decide, time it for two weeks. The one study that measured both perceived and actual time found people overestimated AI’s effect on their own tasks by around 40 percentage points.
Say you’ve got a 60-page supplier agreement and a meeting at two. You paste it in and ask for a summary of the commercial terms and anything unusual. What comes back is clear, well-organised and correct about the pricing, the term, and the renewal mechanism. You take it into the meeting.
What it got wrong was the termination clause, which has a carve-out that changes who pays what if you exit in year one. Nothing in the summary looked uncertain. There was no hedge, no lower-confidence sentence, no visible seam.
Now compare that to a task that sounds much harder: draft a position paper on how the company should think about a regulatory change. Genuinely difficult, and AI does it usefully, because you’re going to read every line and argue with it.
So the hard task went fine and the easy one caused a problem. There’s nothing unlucky about that, it’s the shape of the thing, and once you see it you stop trying to judge fit by how hard the work feels.
This article is a test for that judgement. I’d rather give you something you can run on Monday than a philosophy of AI, so most of what follows is tables you can copy.
Somebody did test this properly, with BCG consultants, and the design is the interesting part.
They gave 758 knowledge workers 18 realistic tasks that sat inside AI’s capability, and one complex managerial task deliberately chosen to sit outside it. Inside, the AI groups completed 12.2% more tasks, 25.1% faster, at better quality; on the single task outside, they were 19% less likely to produce a correct solution than the people working without it[1].
The phrase the authors use for this is the jagged frontier, and their own description of it is the line I’d underline: the impact of AI is uneven “even within the same knowledge workflow and with a seemingly similar level of difficulty”[1].
Same workflow. Similar apparent difficulty. Opposite result.
Which means the instinct most people rely on, this feels simple enough to hand over, is measuring the wrong thing entirely. Difficulty is what the task feels like to you. Fit is about whether the answer sits in territory the model handles well, and those two are not related in a way you can sense.
The natural response is: fine, I’ll just notice when it goes wrong. That turns out to be harder than it sounds too.
METR ran a randomised trial with experienced developers and, unusually, asked them to estimate their own speed both before and after. The headline result from that 2025 study has since been superseded by METR’s own follow-up work, and they’ve said so publicly, so I’m not going to quote it. What survived, and what METR still states in its current writing, is the calibration finding: people “overestimated AI’s effect on their time spent on tasks by 40 percentage points on average”[2].
Forty points, in the flattering direction, by people who had just done the work.
Feeling faster isn’t evidence. If it matters, time it.
Since neither difficulty nor your sense of how it went is reliable, you need something external. These are the five questions I use. They take about ninety seconds and they’re deliberately blunt.
| Question | Score 2 | Score 1 | Score 0 |
|---|---|---|---|
| Could you check the answer? | Yes, in minutes, against something definite | Yes, but it takes real effort | Not really, you’d be trusting it |
| Does it need current, local or proprietary information? | No, it’s general knowledge or you supply the source | Some, and you can attach it | Yes, and it lives in people’s heads |
| How many steps, and does step three depend on step two? | One or two, independent | Three or four, loosely linked | A long chain where one wrong step poisons the rest |
| What does a confident wrong answer cost? | A minute of your time | An awkward correction | Money, a customer, a legal position, or someone’s trust |
| Is this the skill you’re paid to be good at? | No, it’s overhead | Adjacent to it | Yes, it’s the core of the job |
Score each row 0, 1 or 2. Interpretation bands: 8 to 10, hand the whole task over. 5 to 7, take a draft and edit. 3 to 4, use AI only to check your own work. 0 to 2, do it by hand. Any single zero on the cost row caps you at ‘draft’ regardless of total.
That last override is the one that saves you. A task can score well on everything else and still be a bad candidate because the downside is asymmetric. Fast and cheap on the upside, expensive and slow to unwind on the downside.
The second question, about proprietary information, is the one people underrate. METR offers a useful way to think about it: benchmark tasks are self-contained, whereas most real work “draws on prior context, such as previous conversations, tacit knowledge, or familiarity with an existing code base,” so a model’s performance is better understood as what someone with no prior context, “like a new hire or freelance contractor,” could manage[5]. That’s the right mental model. You wouldn’t hand a first-day contractor the question of why churn moved last quarter.
Here’s the test run against tasks that actually come up in a normal week, so you can see the shape of the answers rather than just the rubric.
| Task | Score | Verdict | Why it lands there |
|---|---|---|---|
| Turning meeting notes into actions | 9 | Hand over | You were in the meeting, so checking takes a minute |
| Rewriting a paragraph you already wrote | 10 | Hand over | You know what you meant; wrong is instantly visible |
| Summarising a contract you’ll act on | 4 | Check only | Checking properly means reading it anyway, and the cost row is a zero |
| Drafting a job description | 7 | Draft | Generic is the failure mode, not wrong; you edit in the specifics |
| Explaining last quarter’s numbers | 3 | Check only | The explanation lives in things only your team knows |
| First-pass research on an unfamiliar market | 7 | Draft | Fine as a starting map, not as a source. Verify anything you’ll repeat |
| Deciding who to promote | 0 | By hand | Judgement, consequence and your actual job, all at once |
| Writing a difficult message to a colleague | 4 | Check only | Useful as a second opinion on tone; the words should be yours |
The fit test applied to eight recurring professional tasks. The scores are a worked illustration of the rubric above, not measured data.
Most advice about this treats it as a yes or no question, and that’s what makes it feel unhelpful. Almost every real task has a version where AI helps and a version where it doesn’t, and the difference is how much of the job you actually hand over.
The four handover levels used by the fit test above. Level 3 is the most underused and often the safest place to start on a high-stakes task.
Level 3 deserves more attention than it gets. When you write the analysis yourself and then ask AI what’s missing, you’ve reversed the risk entirely. A wrong suggestion costs you ten seconds of consideration, because you already know the material well enough to dismiss it. The failure mode of level 1 on the same task is invisible; the failure mode of level 3 is just a bad suggestion you ignore.
If you want a fuller version of the checking move, our guide to fact-checking what ChatGPT tells you covers how to interrogate an answer rather than just re-reading it.
Everything so far has been about capability. There’s a second reason not to hand something over, and it has nothing to do with whether AI would do it well.
A study published in The Lancet Gastroenterology & Hepatology in August 2025 looked at over 1,400 colonoscopies across four centres in Poland. It compared how well 19 experienced endoscopists, each with more than 2,000 procedures behind them, detected precancerous growths without AI assistance, before and after AI was routinely introduced. Detection in the non-AI-assisted procedures fell from 28.4% to 22.4%, a 20% relative drop, in work the authors are careful to describe as observational rather than randomised, so hold it as a strong signal rather than proof[3].
A linked commentary called it the first real-world clinical evidence of deskilling, and warned about “the quiet erosion of fundamental skills.”
These were specialists with thousands of repetitions. Several months of good AI support measurably changed what they could do without it.
I don’t think that means avoid AI. I think it means choose, deliberately, a small number of things you keep doing by hand, and choose them by the same logic a musician keeps practising scales. Not because it’s efficient. Because the skill is the thing you’re actually selling.
Keep doing by hand the things you’re paid to be good at.
For most people that list is short. Mine has three things on it:
Everything else is fair game, and I’d rather spend the saved hours on those three than protect a general principle about doing things the hard way.
The obvious objection to all of this is that checking properly takes as long as doing it yourself. Sometimes true. Usually it’s true because people check the wrong thing, reading the output for whether it sounds right rather than testing the specific part most likely to be wrong.
Verification is also where the real-world failure sits. In a survey of 48,340 people across 47 countries by the University of Melbourne and KPMG, 56% of employees reported making mistakes in their work as a result of AI use, and on the underlying chart 44% said they do not verify AI output most of the time[4]. Worth flagging that a headline figure of 66% circulates from the same study, but that number counts people who have ever relied on unchecked output “including rarely,” so 44% is the more defensible one and I’d rather use it.
What actually keeps checking cheap is knowing in advance which single thing to check.
If you can’t check it, you can’t hand it over. Decide the check before you decide to use AI.
| Task type | The thing most likely to be wrong | The check | Time |
|---|---|---|---|
| Summary of a document | Conditions, exceptions and carve-outs | Search the original for ‘unless’, ‘except’, ‘provided that’ | 2 min |
| Anything with numbers | The arithmetic and the column it came from | Recalculate one figure; make it name its source | 2 min |
| Research or background | Sources that don’t say what it claims | Open two citations and read the sentence | 4 min |
| Anything about your company | Plausible details it filled in for you | Scan for specifics you never supplied | 2 min |
| A list of any kind | A missing item, not a wrong one | Ask what a reasonable person would add | 1 min |
A verification ladder by task type. The point is to check one targeted thing rather than re-reading for general correctness, which is what makes checking feel expensive.
Two habits make the whole thing cheaper. Ask it to mark what it’s least confident about before you read anything, which gives you a place to start even though the confidence estimate itself isn’t reliable. And ask it to name the specific source for any claim, because a claim that can’t name a source is the one to check first. Our anti-hallucination toolkit goes further into the techniques that reduce how often you need to.
The fit test gets you a sensible starting decision. It doesn’t tell you whether you were right, and given that people misjudge their own AI time savings by around 40 percentage points, your memory in a fortnight won’t tell you either.
So run something closer to an experiment. It’s less effort than it sounds and it takes about thirty seconds a day.
| Run | Expected | Actual, including checking | Rework needed? |
|---|---|---|---|
| Weekly performance summary, run 1 | 10 min | 18 min | Yes, two figures didn’t match the sheet |
| Run 2 | 10 min | 12 min | No |
| Run 3 | 10 min | 11 min | No |
| Run 4 | 10 min | 25 min | Yes, invented a channel name |
| Verdict: average 16.5 minutes against a felt 10. Still faster than the 40 it used to take, so keep it, but the checking step is not optional and the estimate was never right. | |||
A worked example of the two-week log. The illustrative point is the gap between the expected column and the actual one, which is invisible without writing both down.
Most of the time the answer is keep going. That’s fine, and now you know rather than assume. Occasionally the answer is that a task you’d handed over was quietly costing you more than it saved, and you’d never have found that by reflecting on it.
Pick the task you’re least sure about. Run the five questions on it this afternoon, write down the score, and start the log tomorrow.
Run five questions against it: could you check the answer, does it need current or proprietary information, how many dependent steps does it involve, what does a confident wrong answer cost, and is this the skill you are paid to be good at. Score each 0 to 2. Below 5 out of 10, don’t hand the whole thing over. And treat the cost question as an override: if a wrong answer costs money, a customer or a legal position, take a draft at most, no matter how well the task scores everywhere else.
Because difficulty for you and difficulty for a model are unrelated. Researchers call this the jagged frontier. In a Harvard Business School field experiment with 758 consultants, people using AI did better on 18 tasks inside its capability and 19% worse on one complex managerial task just outside it, and the authors note the unevenness shows up even within the same workflow at a seemingly similar level of difficulty. Summarising a contract feels easier than writing a strategy paper, but you will read every line of the strategy paper and probably not re-read the contract.
Only if you check the whole thing. Targeted checking is fast: for a document summary, search the original for ‘unless’, ‘except’ and ‘provided that’, because conditions and carve-outs are what gets dropped. For anything with numbers, recalculate one figure and make the tool name the column it came from. For research, open two citations and read the actual sentence. Each of those takes two to four minutes, which is a very different proposition from re-reading for general correctness.
There’s real evidence it can, for the specific skills you stop practising. A 2025 study in The Lancet Gastroenterology and Hepatology found that experienced endoscopists’ detection rate in procedures done without AI fell from 28.4% to 22.4% after AI was routinely introduced, which a linked commentary described as the first real-world clinical evidence of deskilling. The study is observational rather than randomised, so hold it as a strong signal rather than proof. The practical response is to deliberately keep a short list of tasks you do unassisted, chosen because the skill is what you are actually selling.
Measure it, because your impression is unreliable in a specific and consistent direction. In the one study that captured both perceived and actual time on the same tasks, people overestimated AI’s effect on their time by around 40 percentage points. Pick one recurring task, write down your expected time before each run, record the actual time including checking and rework, and note anything you had to redo. After two weeks compare the columns. Thirty seconds a day gives you an answer that reflection cannot.
One widely quoted finding was deliberately left out of this article. METR’s 2025 randomised trial reporting that AI tooling slowed experienced developers down has been explicitly superseded by METR’s own 2026 follow-up work, and METR now states on that page that the results no longer reflect current model impact. Quoting the headline would have been an error in an August 2026 piece, so only the calibration finding that METR still restates, the roughly 40-percentage-point gap between perceived and actual effect, is used here. Similarly, a 66% figure circulates from the KPMG and University of Melbourne study for employees relying on unchecked AI output; the underlying chart shows that includes people who do so rarely, so the stricter 44%-do-not-routinely-verify figure is used instead. The Harvard Business School working paper and the KPMG global report are both freely readable and were read directly. The Lancet study itself is paywalled, so the figures quoted come from The Lancet’s own press release, which carries them verbatim, and the study’s observational design is stated rather than glossed.