Explore our AI courses, practical training for non-technical teamsExplore courses Explore AI courses
Learning AIAI LiteracyJudgement

When Not to Use AI at Work: A Fit Test for Any Task

The tasks AI handles badly don't feel any harder than the ones it handles well. That's the whole problem.

TLDR: Deciding when not to use AI is a skill, and it isn’t intuition. In a Harvard Business School field experiment with 758 consultants, people using AI did measurably better on 18 tasks inside its capability and 19% worse on one task just outside it, chosen to look similarly difficult. So the useful test isn’t how hard a task feels. It’s whether you could check the answer, whether the task needs information only your organisation holds, and what happens if a confident wrong answer goes unnoticed.
19%Less likely to reach the correct answer with AI on a task just outside its capability (HBS/BCG, 758 consultants)
40ptsHow much people overestimated AI’s effect on their own task time, in the one study that measured both
44%Of employees say they do not verify AI output most of the time (KPMG and University of Melbourne)

Share this article

The Short Version

Run five questions against any task: can you check the answer, does it need current or proprietary information, how many steps does it take, what does a wrong answer cost, and are you the person who is supposed to be good at this. Score it, and you get one of four answers rather than yes or no: hand over the whole task, take a draft, use it as a checker only, or keep it by hand. And whatever you decide, time it for two weeks. The one study that measured both perceived and actual time found people overestimated AI’s effect on their own tasks by around 40 percentage points.

The task that looked like a perfect job for AI

Say you’ve got a 60-page supplier agreement and a meeting at two. You paste it in and ask for a summary of the commercial terms and anything unusual. What comes back is clear, well-organised and correct about the pricing, the term, and the renewal mechanism. You take it into the meeting.

What it got wrong was the termination clause, which has a carve-out that changes who pays what if you exit in year one. Nothing in the summary looked uncertain. There was no hedge, no lower-confidence sentence, no visible seam.

Now compare that to a task that sounds much harder: draft a position paper on how the company should think about a regulatory change. Genuinely difficult, and AI does it usefully, because you’re going to read every line and argue with it.

So the hard task went fine and the easy one caused a problem. There’s nothing unlucky about that, it’s the shape of the thing, and once you see it you stop trying to judge fit by how hard the work feels.

This article is a test for that judgement. I’d rather give you something you can run on Monday than a philosophy of AI, so most of what follows is tables you can copy.

Why you can't feel where the edge is

Somebody did test this properly, with BCG consultants, and the design is the interesting part.

They gave 758 knowledge workers 18 realistic tasks that sat inside AI’s capability, and one complex managerial task deliberately chosen to sit outside it. Inside, the AI groups completed 12.2% more tasks, 25.1% faster, at better quality; on the single task outside, they were 19% less likely to produce a correct solution than the people working without it[1].

The phrase the authors use for this is the jagged frontier, and their own description of it is the line I’d underline: the impact of AI is uneven “even within the same knowledge workflow and with a seemingly similar level of difficulty”[1].

Same workflow. Similar apparent difficulty. Opposite result.

Which means the instinct most people rely on, this feels simple enough to hand over, is measuring the wrong thing entirely. Difficulty is what the task feels like to you. Fit is about whether the answer sits in territory the model handles well, and those two are not related in a way you can sense.

You also can’t tell afterwards

The natural response is: fine, I’ll just notice when it goes wrong. That turns out to be harder than it sounds too.

METR ran a randomised trial with experienced developers and, unusually, asked them to estimate their own speed both before and after. The headline result from that 2025 study has since been superseded by METR’s own follow-up work, and they’ve said so publicly, so I’m not going to quote it. What survived, and what METR still states in its current writing, is the calibration finding: people “overestimated AI’s effect on their time spent on tasks by 40 percentage points on average”[2].

Forty points, in the flattering direction, by people who had just done the work.

The Fit Rule

Feeling faster isn’t evidence. If it matters, time it.

Five questions that tell you whether to hand a task over

Since neither difficulty nor your sense of how it went is reliable, you need something external. These are the five questions I use. They take about ninety seconds and they’re deliberately blunt.

The five-question fit test, scored

QuestionScore 2Score 1Score 0
Could you check the answer?Yes, in minutes, against something definiteYes, but it takes real effortNot really, you’d be trusting it
Does it need current, local or proprietary information?No, it’s general knowledge or you supply the sourceSome, and you can attach itYes, and it lives in people’s heads
How many steps, and does step three depend on step two?One or two, independentThree or four, loosely linkedA long chain where one wrong step poisons the rest
What does a confident wrong answer cost?A minute of your timeAn awkward correctionMoney, a customer, a legal position, or someone’s trust
Is this the skill you’re paid to be good at?No, it’s overheadAdjacent to itYes, it’s the core of the job

Score each row 0, 1 or 2. Interpretation bands: 8 to 10, hand the whole task over. 5 to 7, take a draft and edit. 3 to 4, use AI only to check your own work. 0 to 2, do it by hand. Any single zero on the cost row caps you at ‘draft’ regardless of total.

That last override is the one that saves you. A task can score well on everything else and still be a bad candidate because the downside is asymmetric. Fast and cheap on the upside, expensive and slow to unwind on the downside.

The second question, about proprietary information, is the one people underrate. METR offers a useful way to think about it: benchmark tasks are self-contained, whereas most real work “draws on prior context, such as previous conversations, tacit knowledge, or familiarity with an existing code base,” so a model’s performance is better understood as what someone with no prior context, “like a new hire or freelance contractor,” could manage[5]. That’s the right mental model. You wouldn’t hand a first-day contractor the question of why churn moved last quarter.

Here’s the test run against tasks that actually come up in a normal week, so you can see the shape of the answers rather than just the rubric.

The test run against eight real tasks

TaskScoreVerdictWhy it lands there
Turning meeting notes into actions9Hand overYou were in the meeting, so checking takes a minute
Rewriting a paragraph you already wrote10Hand overYou know what you meant; wrong is instantly visible
Summarising a contract you’ll act on4Check onlyChecking properly means reading it anyway, and the cost row is a zero
Drafting a job description7DraftGeneric is the failure mode, not wrong; you edit in the specifics
Explaining last quarter’s numbers3Check onlyThe explanation lives in things only your team knows
First-pass research on an unfamiliar market7DraftFine as a starting map, not as a source. Verify anything you’ll repeat
Deciding who to promote0By handJudgement, consequence and your actual job, all at once
Writing a difficult message to a colleague4Check onlyUseful as a second opinion on tone; the words should be yours

The fit test applied to eight recurring professional tasks. The scores are a worked illustration of the rubric above, not measured data.

Full, draft, check, keep: four ways to use AI on one task

Most advice about this treats it as a yes or no question, and that’s what makes it feel unhelpful. Almost every real task has a version where AI helps and a version where it doesn’t, and the difference is how much of the job you actually hand over.

Four levels of handover, with what each looks like in practice

1FullIt does the task, you glance at the result. For work where wrong is obvious and cheap. Example: reformatting a list.
2DraftIt produces version one, you rewrite meaningfully. Example: a job description, where the failure mode is blandness.
3CheckYou do the work, it looks for what you missed. Example: ‘what would a sceptical CFO ask about this?’
4KeepYou do it, unassisted, on purpose. Example: the calls that define your professional judgement.

The four handover levels used by the fit test above. Level 3 is the most underused and often the safest place to start on a high-stakes task.

Level 3 deserves more attention than it gets. When you write the analysis yourself and then ask AI what’s missing, you’ve reversed the risk entirely. A wrong suggestion costs you ten seconds of consideration, because you already know the material well enough to dismiss it. The failure mode of level 1 on the same task is invisible; the failure mode of level 3 is just a bad suggestion you ignore.

If you want a fuller version of the checking move, our guide to fact-checking what ChatGPT tells you covers how to interrogate an answer rather than just re-reading it.

The tasks worth keeping by hand even when AI could do them

Everything so far has been about capability. There’s a second reason not to hand something over, and it has nothing to do with whether AI would do it well.

A study published in The Lancet Gastroenterology & Hepatology in August 2025 looked at over 1,400 colonoscopies across four centres in Poland. It compared how well 19 experienced endoscopists, each with more than 2,000 procedures behind them, detected precancerous growths without AI assistance, before and after AI was routinely introduced. Detection in the non-AI-assisted procedures fell from 28.4% to 22.4%, a 20% relative drop, in work the authors are careful to describe as observational rather than randomised, so hold it as a strong signal rather than proof[3].

A linked commentary called it the first real-world clinical evidence of deskilling, and warned about “the quiet erosion of fundamental skills.”

These were specialists with thousands of repetitions. Several months of good AI support measurably changed what they could do without it.

I don’t think that means avoid AI. I think it means choose, deliberately, a small number of things you keep doing by hand, and choose them by the same logic a musician keeps practising scales. Not because it’s efficient. Because the skill is the thing you’re actually selling.

The Keep Rule

Keep doing by hand the things you’re paid to be good at.

For most people that list is short. Mine has three things on it:

  • The analysis my judgement is built on. If I stop doing the thinking, I lose the ability to tell when someone else’s thinking is off.
  • The first draft of anything I’ll be quoted on. Editing an AI draft toward what I meant is a different skill from working out what I mean.
  • Conversations that need me to have actually thought about the person. A well-worded message about someone you haven’t considered properly is worse than a clumsy one about someone you have.

Everything else is fair game, and I’d rather spend the saved hours on those three than protect a general principle about doing things the hard way.

Verification that doesn't cost more than the task saved

The obvious objection to all of this is that checking properly takes as long as doing it yourself. Sometimes true. Usually it’s true because people check the wrong thing, reading the output for whether it sounds right rather than testing the specific part most likely to be wrong.

Verification is also where the real-world failure sits. In a survey of 48,340 people across 47 countries by the University of Melbourne and KPMG, 56% of employees reported making mistakes in their work as a result of AI use, and on the underlying chart 44% said they do not verify AI output most of the time[4]. Worth flagging that a headline figure of 66% circulates from the same study, but that number counts people who have ever relied on unchecked output “including rarely,” so 44% is the more defensible one and I’d rather use it.

What actually keeps checking cheap is knowing in advance which single thing to check.

The Verification Rule

If you can’t check it, you can’t hand it over. Decide the check before you decide to use AI.

What to check, by task type, and how long it should take

Task typeThe thing most likely to be wrongThe checkTime
Summary of a documentConditions, exceptions and carve-outsSearch the original for ‘unless’, ‘except’, ‘provided that’2 min
Anything with numbersThe arithmetic and the column it came fromRecalculate one figure; make it name its source2 min
Research or backgroundSources that don’t say what it claimsOpen two citations and read the sentence4 min
Anything about your companyPlausible details it filled in for youScan for specifics you never supplied2 min
A list of any kindA missing item, not a wrong oneAsk what a reasonable person would add1 min

A verification ladder by task type. The point is to check one targeted thing rather than re-reading for general correctness, which is what makes checking feel expensive.

Two habits make the whole thing cheaper. Ask it to mark what it’s least confident about before you read anything, which gives you a place to start even though the confidence estimate itself isn’t reliable. And ask it to name the specific source for any claim, because a claim that can’t name a source is the one to check first. Our anti-hallucination toolkit goes further into the techniques that reduce how often you need to.

How to find out for real, in two weeks

The fit test gets you a sensible starting decision. It doesn’t tell you whether you were right, and given that people misjudge their own AI time savings by around 40 percentage points, your memory in a fortnight won’t tell you either.

So run something closer to an experiment. It’s less effort than it sounds and it takes about thirty seconds a day.

  1. Pick one recurring task you currently hand to AI. One, not five.
  2. Before you start each time, write down what you expect it to take. Before, not after.
  3. Note the actual finish time, including the checking and any rework.
  4. Note separately whether you had to redo anything, and what specifically was wrong.
  5. After two weeks, compare the columns. Then decide.

The two-week log, filled in

RunExpectedActual, including checkingRework needed?
Weekly performance summary, run 110 min18 minYes, two figures didn’t match the sheet
Run 210 min12 minNo
Run 310 min11 minNo
Run 410 min25 minYes, invented a channel name
Verdict: average 16.5 minutes against a felt 10. Still faster than the 40 it used to take, so keep it, but the checking step is not optional and the estimate was never right.

A worked example of the two-week log. The illustrative point is the gap between the expected column and the actual one, which is invisible without writing both down.

Most of the time the answer is keep going. That’s fine, and now you know rather than assume. Occasionally the answer is that a task you’d handed over was quietly costing you more than it saved, and you’d never have found that by reflecting on it.

Pick the task you’re least sure about. Run the five questions on it this afternoon, write down the score, and start the log tomorrow.

Sana Mian
Sana Mian, Co-Founder of Future Factors AI

Sana is an AI educator and learning designer specialising in making complex ideas stick for non-technical professionals. She has trained 2,000+ learners across corporate teams, bootcamps, and keynote stages. Future Factors offers AI Bootcamps, Corporate Workshops, and Speaking & Consulting for businesses ready to adopt AI without the overwhelm.

More about Sana →

Frequently Asked Questions

How do I know when a task is a bad fit for AI?

Run five questions against it: could you check the answer, does it need current or proprietary information, how many dependent steps does it involve, what does a confident wrong answer cost, and is this the skill you are paid to be good at. Score each 0 to 2. Below 5 out of 10, don’t hand the whole thing over. And treat the cost question as an override: if a wrong answer costs money, a customer or a legal position, take a draft at most, no matter how well the task scores everywhere else.

Why does AI get easy-sounding tasks wrong and hard ones right?

Because difficulty for you and difficulty for a model are unrelated. Researchers call this the jagged frontier. In a Harvard Business School field experiment with 758 consultants, people using AI did better on 18 tasks inside its capability and 19% worse on one complex managerial task just outside it, and the authors note the unevenness shows up even within the same workflow at a seemingly similar level of difficulty. Summarising a contract feels easier than writing a strategy paper, but you will read every line of the strategy paper and probably not re-read the contract.

Isn't checking AI output as slow as just doing the task myself?

Only if you check the whole thing. Targeted checking is fast: for a document summary, search the original for ‘unless’, ‘except’ and ‘provided that’, because conditions and carve-outs are what gets dropped. For anything with numbers, recalculate one figure and make the tool name the column it came from. For research, open two citations and read the actual sentence. Each of those takes two to four minutes, which is a very different proposition from re-reading for general correctness.

Does using AI make you worse at your job over time?

There’s real evidence it can, for the specific skills you stop practising. A 2025 study in The Lancet Gastroenterology and Hepatology found that experienced endoscopists’ detection rate in procedures done without AI fell from 28.4% to 22.4% after AI was routinely introduced, which a linked commentary described as the first real-world clinical evidence of deskilling. The study is observational rather than randomised, so hold it as a strong signal rather than proof. The practical response is to deliberately keep a short list of tasks you do unassisted, chosen because the skill is what you are actually selling.

How can I tell whether AI is genuinely saving me time?

Measure it, because your impression is unreliable in a specific and consistent direction. In the one study that captured both perceived and actual time on the same tasks, people overestimated AI’s effect on their time by around 40 percentage points. Pick one recurring task, write down your expected time before each run, record the actual time including checking and rework, and note anything you had to redo. After two weeks compare the columns. Thirty seconds a day gives you an answer that reflection cannot.

About This Article

One widely quoted finding was deliberately left out of this article. METR’s 2025 randomised trial reporting that AI tooling slowed experienced developers down has been explicitly superseded by METR’s own 2026 follow-up work, and METR now states on that page that the results no longer reflect current model impact. Quoting the headline would have been an error in an August 2026 piece, so only the calibration finding that METR still restates, the roughly 40-percentage-point gap between perceived and actual effect, is used here. Similarly, a 66% figure circulates from the KPMG and University of Melbourne study for employees relying on unchecked AI output; the underlying chart shows that includes people who do so rarely, so the stricter 44%-do-not-routinely-verify figure is used instead. The Harvard Business School working paper and the KPMG global report are both freely readable and were read directly. The Lancet study itself is paywalled, so the figures quoted come from The Lancet’s own press release, which carries them verbatim, and the study’s observational design is stated rather than glossed.

Sources

  1. Dell’Acqua, Fabrizio et al. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality. Harvard Business School Working Paper 24-013, with Boston Consulting Group. (Free full text via SSRN. 758 knowledge workers, preregistered.) https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4573321
  2. METR. How Much Are Researchers and Engineers Using AI? 11 May 2026, restating the calibration finding from Becker et al. (2025). (Freely readable. METR’s superseding note on the 2025 productivity result is at metr.org/blog/2026-02-24-uplift-update/.) https://metr.org/blog/2026-05-11-ai-usage-survey/
  3. Budzyń, K., Romańczyk, M., Mori, Y. et al. Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy. The Lancet Gastroenterology & Hepatology, 12 August 2025. (Article paywalled; figures quoted here come from The Lancet press office release, which carries them verbatim.) https://www.thelancet.com/journals/langas/article/PIIS2468-1253(25)00133-5/fulltext
  4. University of Melbourne and KPMG International. Trust, attitudes and use of artificial intelligence: A global study 2025. (48,340 respondents across 47 countries. Full report PDF freely downloadable; figures here are from pages 75 and the underlying Figures 44 and 45.) https://kpmg.com/xx/en/our-insights/ai-and-technology/trust-attitudes-and-use-of-ai.html
  5. METR. Measuring AI Ability to Complete Long Tasks. Page updated 8 May 2026. (Freely readable. Source for the reliability-by-task-length illustration and the ‘no prior context’ framing referenced in the fit test.) https://metr.org/time-horizons/

Psst, Hey You!

(Yeah, You!)

Want helpful AI tips flying Into your inbox?

Weekly tips. Real examples. Practical help for busy professionals.

We care about your data, check out our privacy policy.