Heavy AI use and sharp judgment aren't opposites. They just don't happen to coexist by accident.
MIT’s EEG study on AI-assisted essay writing found weaker brain connectivity and lower self-reported ownership among ChatGPT users, but it tracked 54 people, mostly students, on one task, for four months, not a career. A separate Microsoft and Carnegie Mellon survey of 319 knowledge workers found that confidence, not AI use itself, predicts whether people actually think critically: more confidence in the tool means less scrutiny, more confidence in yourself means more. Neither study proves AI causes lasting cognitive decline. What actually protects your judgment is specific: form your own view before you ask, interrogate the answer instead of accepting it, and keep at least one task manual on purpose so the underlying skill doesn’t go quiet.
You ask AI something you’re only half sure about yourself. The answer comes back fast, well organized, confident sounding. You skim it, it looks right, you move on. Nothing about that moment tells you whether you just got a genuinely good answer or a wrong one dressed up the same way a right one would be. That’s the real worry behind “is AI making me worse at thinking,” and it’s worth being precise about what’s actually been measured here, twice now, because the popular version of this story has run well ahead of the evidence, in both directions.
The most-cited piece of evidence comes out of MIT Media Lab, led by Nataliya Kosmyna, first posted in mid-2025 and revised since.[1] Fifty-four people were split into three groups writing essays over four months: one group used ChatGPT, one used a search engine, one used neither, and everyone wore an EEG cap. The finding that made headlines was clear: the ChatGPT group showed the weakest brain connectivity of the three while writing, the search-engine group sat in the middle, and the no-tools group showed the strongest, most distributed activity. People in the ChatGPT group also said they felt less ownership over what they’d written, and several struggled to accurately quote their own essays back afterward.
Here’s what that work didn’t show. Fifty-four people is a real but small sample, and they were mostly university students and postdocs in their twenties, recruited from five Boston-area schools, writing timed essays under lab conditions on one narrow task. That’s a genuinely useful window into what happens to engagement during unfamiliar, effortful writing. It isn’t evidence that a financial analyst who’s used AI daily for a decade has a measurably worse brain than one who hasn’t. The team behind it frames its own results as raising a question worth investigating further, not as settling one.
The second piece of work asked a more directly useful question: not what’s happening inside someone’s skull, but what actually predicts whether people think critically when AI is involved. A team from Microsoft and Carnegie Mellon, presenting at CHI 2025, asked 319 knowledge workers about their own use of generative AI and collected 936 specific, real examples of how they’d actually used it at work.[2] Two things stood out. People who felt more confident in the AI’s answer said they put in less critical-thinking effort. People who felt more confident in their own knowledge of the subject said they put in more.
That’s a self-assessed finding, not a measured one. Nobody’s actual output was graded for accuracy against a fixed answer key; the team asked people to describe their own effort and confidence, then looked for the pattern across hundreds of examples. That kind of self-assessment is a real signal. People generally have some sense of when they’re skimming versus actually checking something. But it’s a different kind of evidence from an EEG cap or a scored test, and it’s worth holding it as exactly what it is: what people say they did, not a direct measurement of what they got right.
Neither paper proves AI makes anyone permanently worse at thinking. What they add up to is narrower, and honestly more useful: thinking gets quieter when a tool takes over a task, and how much you personally scrutinize an AI answer has less to do with the tool than with how confident you already feel, in the tool and in yourself. That’s the thing worth building habits around. Not the tool.
Here’s the mechanism, stated plainly. The better an AI answer looks, the less most people check it. Not because they’re lazy. Because a fluent, confident-sounding answer and a genuinely correct one feel identical from the outside, and checking takes effort your brain would rather not spend if it doesn’t think it’s needed.
Say a supply chain planner is prepping for a Tuesday call with a vendor about switching materials suppliers. She asks AI to pull together a comparison across four candidates: cost, lead time, minimum order quantities. What comes back is a clean table, definite-sounding numbers, a confident recommendation at the bottom. It looks like something a specialist put together. She skims it, likes the shape of it, walks into the call with it open on her second monitor. Halfway through, the vendor rep corrects one of the lead-time figures on the spot, off by three weeks, sourced from a page updated more recently than whatever the model had learned from. Nothing about the table had signaled which numbers were solid and which were closer to a guess dressed up the same way.
That’s the pattern Lee and her Microsoft and Carnegie Mellon colleagues found in their work: confidence in the AI’s output and confidence in your own knowledge pull in opposite directions on how much scrutiny you apply.[2] The less sure you are of the subject yourself, the more the tool’s fluency quietly substitutes for your own checking, which is exactly the situation where you’re least equipped to catch a mistake. If you want the deeper mechanics of why a fluent wrong answer can look identical to a fluent right one, we’ve covered that separately in why does AI hallucinate.
Before you act on an AI answer that actually matters, score yourself honestly, one to five, on two separate questions.
| Question | Score 1 | Score 5 |
|---|---|---|
| My own knowledge of this subject | I’d have to guess | I could explain it to someone else |
| How finished this answer sounds | Rough, hedged, clearly a first pass | Polished, definite, ready to send as-is |
If your score on the second question is higher than the first, that gap is the exact condition the research ties to skipped scrutiny. Slow down and verify before you use it. [2]
A polished answer and a correct one feel the same from the outside. The variable the research ties to catching the difference is how much you already know, not how confident the AI sounds.
None of this means treat every AI answer with the same suspicion. It means notice the specific moment when your own uncertainty and the tool’s confidence are furthest apart, because that’s the moment your guard is lowest exactly when it should be highest.
Not every task deserves this much caution, and treating all of them the same is its own mistake, the exhausting kind where you triple-check a meeting summary as carefully as you’d check a number that’s about to go in front of a client.
The honest split isn’t about how impressive the task sounds. It’s about three narrower questions: is it reversible if wrong, is it cheap to check, and is your job specifically to have the judgment, not just produce the output. A task can look small and still fail that third question. A comp analyst recommending someone’s raise isn’t dealing with a dramatic decision on paper, but getting it slightly wrong doesn’t announce itself the way a broken spreadsheet formula does, so it belongs in a more careful tier even though the tool being used looks identical to the one drafting a meeting recap.
| Tier | What belongs here | Example |
|---|---|---|
| Hand it over | Reversible, cheap to check, no one’s judgment is really on the line | A financial analyst’s first-pass variance report: the numbers either tie out or they don’t, and checking is fast. |
| Draft, then verify | Useful as a starting point, but something you’d be accountable for if wrong | A marketer’s first draft of a competitor summary going out under her name to leadership. |
| Form your view first | The judgment itself is the job, or the mistake wouldn’t announce itself | An HR business partner’s recommendation on a contested promotion, or a pricing call that’s hard to reverse. |
The task in your hands, not the tool you’re using, decides which tier applies. The same AI chat window can sit in any of the three.
Watch for the tasks that quietly slide out of tier one without you noticing:
If two or more of those are true, the task has moved into tier three, whatever it looked like at the start. We built a fuller version of this fit test in when not to use AI at work, which is worth reading alongside this if you want the task-level version rather than the habit-level one this piece focuses on.
Ask AI first and its answer becomes the frame everything else gets measured against, including your own instincts. Psychologists have a name for this: anchoring. Whoever answers first sets the reference point, and it’s genuinely hard to move away from a plausible-sounding number once you’ve read it, even when you’re consciously trying to.
Say a marketing director is putting together a pricing recommendation for a leadership meeting next week: raise the mid-tier plan by a set amount, or hold. Before she opens any AI tool, she writes two sentences by hand: her own gut call, and the one number that would change her mind. Then she asks AI the same question, with the same context she’d give a sharp colleague. The tool comes back leaning the other way, citing a churn risk she hadn’t weighted heavily enough. Now she has something real to reconcile, her own reasoning against a specific counter-argument, instead of a confident-sounding answer she has no independent read on.
If she’d asked AI first, that answer would have become the frame, and anything she thought afterward would have been reacting to it, not arriving there independently. The order matters more than people expect.
Write your own answer in two sentences before you open the AI. If the two answers match, you’ve confirmed something real. If they don’t, you’ve found the one place actually worth your attention.
This is deliberately a small ask, not “think harder about everything you send to AI.” Just thirty seconds, on the specific decisions where your name is attached to the outcome, tier two and tier three from the table above. For a routine meeting recap, skip it. The habit only earns its keep where the stakes do.
Once you’ve got an AI answer in front of you, and it doesn’t just match what you already thought, the next move matters more than the first one. Most people re-read the same answer more carefully, which mostly just confirms they understood what it said, not whether it’s right.
Paste this after any AI answer you’re about to actually act on:
“Which single claim in that answer are you least certain about, and why? What specific piece of information, if it turned out to be wrong, would change your conclusion? Now argue the strongest case against your own answer.”
This does two things a re-read can’t: it forces the model to rank its own claims by confidence instead of presenting them all with the same tone, and it hands you the counter-case you’d otherwise have to build yourself.
The goal isn’t verifying every sentence evenly. That’s exhausting and it isn’t actually how good judgment works, even without AI in the picture. It’s finding the one claim that would flip the decision if it were wrong, and spending your real attention there.
Spend your checking time on the one claim that would change the decision if it were wrong, not evenly across every sentence the model wrote.
For claims you can independently verify, source names, dates, numbers you can look up yourself, our fact-check workflow for ChatGPT walks through the actual steps. Interrogating the answer and fact-checking it are two different moves: one surfaces where the weak point probably is, the other confirms whether it’s actually wrong. The anti-hallucination toolkit is worth keeping open for the second step.
| What AI does | What you still own | How it gets checked |
|---|---|---|
| Drafts the comparison, the summary, or the first-pass recommendation | Deciding which claim actually carries the decision, and whether it’s true | The interrogate-the-answer prompt above, applied to that one claim, before you send or act on it |
The tool produces the draft. Deciding what the draft is actually resting on stays yours, on anything you’re accountable for.
There’s a difference between checking an AI answer and being able to produce one yourself if the tool disappeared tomorrow, and heavy AI use can quietly erode the second without you noticing, because you’re still exercising the first. The part of the MIT work that’s actually the more concrete evidence for skill fade, more than the headline connectivity numbers, showed up in its fourth session. Participants who’d used ChatGPT across the first three sessions were then asked to write unaided. Their brain connectivity in that session measured lower than participants who’d been working without any tool the whole time, a pattern the paper describes as under-engagement.[1] That’s the closest either paper comes to direct evidence of a skill going quiet from disuse, rather than an opinion about what a fluent tool probably does to people.
It’s also, worth saying plainly, one narrow writing task over four months in a lab, not proof that this happens to every skill or that it’s permanent. What it supports is narrower and still useful: a skill you’ve stopped exercising can measurably soften faster than intuition suggests, which is a reason to keep exercising the ones your job actually depends on, not a reason to avoid AI on the rest.
None of this slows down the vast majority of tasks where handing them to AI is the right call. It’s a small, deliberate maintenance cost on the specific skill your job actually depends on.
None of this needs a bigger commitment than you already have room for:
Not because the AI version would be worse. Because that’s the only way to actually find out whether you still can.
Neither of the two major studies on this proves that. MIT’s EEG study found weaker brain connectivity and lower self-reported ownership among people using ChatGPT for essay writing, but it tracked 54 people, mostly students, on one narrow task, over four months, not a career’s worth of varied work. The Microsoft and Carnegie Mellon study found that higher confidence in an AI’s answer predicts less critical-thinking effort, but that’s a self-reported finding, not a measured drop in ability. What both studies support is narrower: thinking gets quieter when a tool takes over, and your own confidence level changes how much you check the result. That’s a habit problem, not proof of permanent decline.
It’s the term MIT’s researchers used for what showed up when participants who’d relied on ChatGPT for three sessions were then asked to write completely unaided: their brain connectivity measured lower than people who’d never used the tool, a pattern the study calls under-engagement. In practice, it means a skill you’ve stopped exercising can go quiet faster than you’d expect. The way to avoid building it up is to keep doing the core task yourself sometimes, on purpose, even when AI could do it faster, not to avoid AI on everything else.
Tasks that are reversible, cheap to check, and don’t put your own judgment on the line are safe to hand over freely, a first-pass meeting summary or a routine reformat. Anything where you’re specifically being paid for the judgment, a hiring call, a pricing recommendation, a client-facing conclusion you’re accountable for, belongs in the tier where you form your own view before you ask. The task doesn’t have to look high-stakes to belong there; it just has to be one where getting it slightly wrong wouldn’t announce itself right away.
Don’t try to verify every sentence with equal effort. Ask the tool which single claim in its own answer it’s least certain about, and what would change the conclusion if that specific claim turned out to be wrong. That tells you exactly where to spend your real checking time, instead of skimming the whole answer at the same shallow level of attention.
Yes, and that’s really the point of this piece. The research points to the risk being specific habits, not the amount of AI use: using it without noticing when you’ve stopped checking, and asking it before you’ve formed your own view. Form your own view first on anything you’re accountable for, interrogate the answer instead of accepting it, and keep one recurring task manual on purpose. None of that requires using AI any less.
Both primary studies referenced in this piece were read directly from their original sources, not secondary coverage: the MIT Media Lab paper via its arXiv listing (abstract, author list, and full methodology confirmed at arxiv.org/abs/2506.08872), and the Microsoft Research and Carnegie Mellon CHI 2025 paper via Microsoft Research’s own publication page and its DOI. The MIT study is small and short-term: 54 participants, mostly university students and postdocs in their twenties from five Boston-area schools, tested on one narrow task (essay writing) across four months. It measured brain connectivity and self-reported ownership, not long-term ability. The Microsoft and Carnegie Mellon study surveyed 319 knowledge workers about 936 real examples of their own generative-AI use; its findings on confidence and critical-thinking effort are self-reported, not a measured test of accuracy or skill. Neither study is treated here as proof that AI use permanently damages thinking ability, and both limitations are stated plainly rather than left out. Some findings from the MIT paper, including its analysis of linguistic similarity across essays and detailed activation patterns in specific brain regions, are real but too granular to turn into a reader action without overstating what they mean, so they were left out of the body by design.