Company policy covers what you can paste in. Nobody wrote you a policy for how hard to check what comes back, and that's the part that actually gets you in trouble.
The risk that lands on you personally isn’t mainly about leaking data (there’s a separate guide for that). It’s trusting a fluent wrong answer because it reads exactly like a right one, and it’s the slower habit of checking less the longer AI has been right. This piece gives you a two-second test for what not to paste, a table for matching how hard you check to what it costs to be wrong, and a plan for the day you’ve already sent something AI got wrong.
A project coordinator is three minutes from sending a client status update. She’d drafted the raw version herself, pasted it into AI to tighten the wording, and was skimming it one last time before hitting send. The draft she wrote said the integration work was “on track.” The version that came back said it was “ahead of schedule.” One word, quietly upgraded, technically still in the right neighbourhood, and completely untrue. She caught it because she happened to reread the whole paragraph instead of just the parts she’d changed. If she’d been thirty seconds more rushed, that sentence goes to a client with her name under it, and nobody at her company is the one who has to explain it to him.
She is.
That’s the risk nobody puts in the company AI policy, because it isn’t really the company’s risk. A policy can tell you which tool is approved and what data category can’t leave the building. It can’t tell you how hard to check the specific paragraph in front of you before you attach your name to it, because that decision happens in a browser tab your IT department will never see, at a speed no policy review can keep up with.
Search “AI safety at work” and almost everything that comes back is written for the organisation: data governance, vendor risk, acceptable use policies, who signs off on a new tool. All of that matters, and if you want the deep version of the data-handling half of this, our guide to using AI without leaking company data covers it properly. This piece is the other half, the one almost nothing is written for: what happens to you, specifically, when the thing you sent turns out to be wrong.
Because that’s the actual shape of the exposure. Not “did I break a rule,” but “did I put something in front of a client, a manager or a board with my judgment stamped on it that I never actually verified.” Your employer’s data policy protects the company. It was never going to protect you from being the person who signed off on a wrong number.
Strip away the specific tool and the specific task, and personal AI risk comes down to three failure points, in the order they usually happen. Only one of them is about what you type in.
Most of what gets written about AI safety spends all its time on the first one and almost none on the other two, which is backwards from how the actual damage happens. Leaking a document is a policy breach with a name attached to it, usually the company’s. A wrong number that made it into a board pack because it read convincingly is a judgment call with your name attached to it.
The rest of this piece is mostly about the second and third failure modes, because that’s where the gap in existing advice actually sits.
Match how hard you check to what it costs to be wrong, not to how confident the answer sounds.
Worth saying plainly before any of this gets operational: some tasks genuinely aren’t a good fit for AI at all, and deciding not to use it is a legitimate answer, not a failure to find the right prompt. If a task turns on a fact you can’t independently verify in the time you have, or on a judgment call that has to be defensible later and traceable to a real source, that’s often a task to do the slower, unaided way.
Sometimes the safest use of AI is not using it.
The leakage half of this gets covered in depth elsewhere, so here’s the short, usable version, because you still need to make this call several times a day and a full policy document isn’t something you can hold in your head mid-task.
Use a two-second test before you paste anything in: would you put this in an email to someone outside your team, at a company you don’t work for? If the answer is no, it doesn’t go into the AI tool either, approved or not. That single question covers most of the judgment calls that actually come up, faster than checking a policy document.
A short list is always faster than a test, though, for the cases where you shouldn’t even be pausing to think. These categories are always out, no exceptions for “just this once” or “it’s probably fine”:
That last one catches more people than it should. Someone drops a screenshot into an AI tool to ask for help formatting a chart, and the browser tab bar along the top still shows the name of the client whose deal is currently confidential. Nothing in the actual chart was sensitive. The tab bar was.
A quick check for the judgment calls, plus the list of cases that skip the judgment call entirely. For the full data-handling picture, see our guide to using AI without leaking company data.
None of this is about which AI tool your company approved. An approved enterprise tool with a clean data agreement still shouldn’t be seeing a colleague’s performance review, because the question was never really about where the data goes. It’s about what you’d be comfortable putting in front of a stranger.
Here’s the part that actually catches people out. You send AI a task, it comes back fast, confident and well formatted, and every visual signal you’d normally use to judge quality (structure, tone, specificity) is present whether the content is right or not. A fabricated figure looks exactly as polished as a real one. Nothing about the experience tells you which you’re holding.
The instinct is to check everything equally hard, which nobody actually does, or to check nothing because it’s been fine so far, which is where the real damage comes from. The better instinct is proportional: match the depth of the check to what happens if this specific piece of output is wrong.
| What the output is for | What it costs if it’s wrong | How hard to check |
|---|---|---|
| A first draft you’ll rewrite yourself before anyone sees it | Nothing. You’re the only reader | Skim for tone. Don’t verify facts you’ll rewrite anyway |
| An internal Slack summary of a meeting you also attended | Low. Colleagues can correct it in the thread | Scan for anything you don’t remember being said |
| A client-facing email with a status, date or number in it | Medium. Wrong info reaches someone outside the company | Check every number and date against the source document, not your memory of it |
| A board pack slide, a legal-adjacent summary, or anything with a figure that drives a decision | High. A wrong number changes what someone decides | Trace every figure back to its original source before it goes in, every time, no exceptions |
A worked example of the Proportional Check Rule. Copy the shape, fill in the row that matches your own recurring task.
Notice what the table isn’t saying: it isn’t telling you to check less on low-stakes work because AI is trustworthy there. It’s telling you the cost of being wrong is what should set the bar, not how confident the sentence sounds or how long the tool has been getting things right. A one-line Slack summary and a board-pack figure can come out of the same tool in the same afternoon, and they don’t deserve the same five seconds of attention.
Where AI genuinely earns a place in this kind of work, it’s worth being explicit about which part it’s actually doing, because “AI helped with this” tends to blur into “AI is responsible for this” if nobody separates the two out loud.
| What AI does | What you still own | How it gets checked |
|---|---|---|
| Drafts the first-pass client update from your raw notes | Deciding whether “on track” became something stronger than what you actually wrote | Reread the final paragraph against your original notes, word for word, before sending |
| Pulls figures out of a long report into a short summary | Confirming every figure traces back to the page it claims to come from | Click through to the source page for each number before it leaves your draft |
| Suggests which of six competitor changes matter this week | Deciding whether any of them actually change your team’s plan | Spot-checked by you before it’s shared, not after |
A worked example, not a universal rule. The split changes with the task; what doesn’t change is that something on the right-hand column stays yours.
A financial analyst and a first-line sales manager will fill that middle column in differently, and that’s the point. An analyst pulling numbers into a forecast owns tracing every figure to its source, because a wrong number compounds into a wrong decision three steps later. A sales manager using AI to draft talking points for a difficult conversation owns the judgment call about tone and timing that the draft can’t make for them. Same tool, same general shape, genuinely different thing each of them is responsible for verifying.
An operations lead had been using AI to summarise weekly vendor check-in calls for a little over a month. The first few summaries she read line by line against her own notes. By week five she was skimming them, because they’d been accurate every time and reading closely felt like it was no longer buying her anything. In week seven, a vendor had verbally agreed to extend a delivery window during the call. The summary didn’t mention it, not because the tool got it wrong exactly, more that the comment was brief and easy to miss in a long transcript. Nobody caught the gap until the vendor referenced the extension three weeks later and the operations lead had no record of agreeing to it.
Nothing about that summary looked different from the accurate ones.
That’s what makes this failure mode harder to catch than the first two. Putting the wrong thing in is a decision you make once, at the point of pasting. Trusting a specific wrong answer is a mistake you can point to afterward. Checking less over time doesn’t have a moment attached to it. It’s not a decision, it’s a habit forming, and habits don’t announce themselves.
A survey from Resume Now, fielded across just over a thousand employed US adults in December 2025, found that 35% of workers say they rarely or only occasionally review AI-generated output before using it.[1] That number on its own doesn’t tell you whether those people started out that way or drifted there. What it does confirm is that “rarely checking” is a common enough state to survey, not a rare lapse, which matches how this failure mode actually spreads: gradually, and usually without anyone deciding to let it.
Six good weeks are evidence the tool is usually right. They are not evidence this specific answer is right.
A simple way to catch the drift before it costs you something: pick one recurring task where AI has become part of your routine, and once a month, check its output as closely as you did in week one, on purpose, even though it’s been reliable. If you catch nothing, that’s genuinely useful information, not wasted time. If you catch something, you’ve found the drift while it was still cheap to find.
Worth being honest about a related failure that isn’t purely psychological: the legal world has produced the clearest public record of what happens when nobody rechecks a fluent wrong answer. Damien Charlotin, a research fellow at HEC Paris, maintains a public database tracking court cases where a party relied on AI-hallucinated material, mostly fabricated case citations. As of 1 September 2026, the database listed 2,005 such cases worldwide, with a practising lawyer the responsible party in 797 of them.[2] These are professionals whose entire training is built around checking sources, caught by exactly the failure mode this section describes: an output that read like the real thing, submitted by someone who’d stopped double-checking that specific kind of task.
Checking is a habit. Habits fade quietly.
At some point, despite all of the above, something wrong is going to get out with your name on it. What you do in the next hour matters more than anything you did to prevent it, and it’s worth having a plan before you need one, because the instinct in the moment is almost always the wrong one.
A consultant sent a client deliverable with a market-size figure that turned out to be about 40% higher than the real number, pulled from an AI summary of a report she hadn’t opened herself. She caught it the next morning, re-reading the deliverable with fresh eyes. Her first instinct was to quietly send a revised file with no explanation and hope nobody had opened the original yet. She didn’t do that. She emailed the client directly: the figure on page 4 was wrong, here’s the corrected number and where it actually comes from, and here’s what she’s changing about her own process so it doesn’t happen again. The client’s reply was two lines: thanks for catching it, appreciate the fast turnaround.
That’s the outcome fast, specific correction usually gets you.
A worked correction, not a general apology template. Steps 2 and 4 are the ones people skip under pressure, and they’re the ones that actually rebuild trust.
Two things in that sequence are easy to get backwards. Don’t lead with “AI was involved,” even though it’s tempting and even though it’s technically true. Said first, it reads as deflection, an attempt to move the blame onto the tool before the person has even heard what went wrong. Said nowhere at all, it also misses the useful part: naming the specific process gap (I didn’t open the source report myself, I trusted a summary of a summary) is what actually prevents a repeat, and that detail belongs in step four, once the correction itself is already handled.
The tone that works here is matter-of-fact, not apologetic in a way that invites more scrutiny than the mistake earned. “The figure on page 4 was wrong, here’s the corrected version” does the job. A long explanation of how sorry you are, with three paragraphs of context before the actual correction, makes the reader wait for the one thing they opened the message to find.
Disclosure is a related question worth answering honestly rather than by instinct, because the instinct varies by person and the honest answer doesn’t have to. A reasonable default: you don’t need to flag that AI helped you think through or draft something you then rewrote and stand behind, the same way you wouldn’t flag which search engine you used. Disclosure becomes necessary where a client is specifically paying for your own bespoke thinking, where the output is being presented as original research or analysis, or where your organisation’s policy already requires it. Some jurisdictions are starting to write formal disclosure obligations into law for specific use cases; our plain-English guide to the EU AI Act covers what’s actually enforced so far, rather than what’s rumoured.
| Situation | Disclose? |
|---|---|
| You used AI to draft something, then rewrote and stand behind every line | No, same as any other drafting tool |
| A client is paying specifically for your own bespoke thinking or analysis | Yes, tell them what role AI played |
| The output is being presented as original research | Yes |
| Your organisation’s policy already requires disclosure | Yes, policy overrides the default |
A default worth checking against your own organisation’s actual policy, which can set a stricter bar than this.
Everything above compresses into something you could genuinely keep taped next to a monitor, or save as a note on your phone. It’s a worked example, not a universal script. Change the specifics to match your own role and the tasks you actually reach for AI to help with, but keep the shape.
| Question | Your answer |
|---|---|
| What never goes in, no matter the tool | CVs pre-offer, named complaints, unsigned contracts, HR files, anything visible in a screenshot’s browser tabs |
| The two-second test for anything else | Would I put this in an email to someone outside my team, at a different company? |
| What gets a light check | Drafts I’m rewriting myself, internal summaries colleagues can correct |
| What gets a full trace-to-source check, every time | Anything client-facing with a number or date, anything feeding a decision |
| Monthly drift check | Pick one routine AI task, check it as closely as week one, on purpose |
| If something wrong already went out | Same channel, name the specific error, give the fix, name the process change. AI mention comes last, if at all |
A filled-in worked example, not a named model. Copy the six rows, keep the shape, and change the answers to match your own role and workload.
Name the specific error before you name what caused it.
None of this requires more time than most people already spend cleaning up AI drafts by hand. It just points the effort at the two places it’s actually needed: the output that costs something if it’s wrong, and the routine task that’s gone quiet long enough that you’ve stopped looking closely. The paste-in rules take a few seconds to apply once they’re familiar. The checking habit is the part that decays without maintenance, which is exactly why it’s the part worth building a monthly check around rather than trusting yourself to remember.
Pick one task you use AI for every week. Write down which row of the guardrails card it falls into. Check it that closely for the next output, even if it’s been fine for months. That’s the whole exercise, and it’s the one that actually tells you whether your own checking habit has quietly slipped.
Generally yes, for the large majority of everyday tasks, provided you apply two separate checks: what you put in (nothing you wouldn’t put in an email to someone outside your team) and what you accept back (checked in proportion to what it costs if the answer is wrong). The risk isn’t the tool itself, it’s skipping either of those two checks, especially the second one, which is easier to let slide over time.
A short, fixed list regardless of which tool or how well it’s covered by your employer’s licence: a candidate’s CV before an offer is made, a customer complaint that names them, an unsigned contract, a colleague’s medical or performance information, and any screenshot that still shows a client name in a browser tab, calendar invite or chat sidebar. For anything not on that list, use the two-second test: would you put this in an email to someone outside your team at a different company?
Match the depth of the check to what it costs if the answer is wrong, not to how confident the output sounds. A draft you’re rewriting yourself needs a skim. A client-facing number or a figure feeding a decision needs to be traced back to its original source every single time, with no exceptions for output that’s been reliable in the past.
Not by default, if you used it to think through or draft something you then rewrote and stand behind, the same way you wouldn’t flag which search engine you used. Disclosure becomes necessary when a client is paying specifically for your own bespoke thinking, when the output is presented as original research, or when your organisation’s policy already requires it, which some policies do.
Correct it fast, in the same channel it went out in. Name the specific thing that was wrong rather than a vague ‘an error occurred.’ Give the corrected version and where it actually comes from. Then name the process step you’re changing so it doesn’t happen again. Don’t lead with ‘AI was involved’ as the first line, it reads as deflection rather than accountability.
The two-second input test, the proportional-checking table and the personal guardrails card are Future Factors’ own synthesis, built for this piece from teaching AI use across corporate workshops, and presented as practical judgment tools rather than as research findings. Both statistics cited were checked directly against their primary source on 2 September 2026: the Resume Now figure against the publisher’s own report page, and the court-case count against Damien Charlotin’s live database rather than any secondary write-up of it. A widely repeated acceleration figure for the same database (fake-citation cases per day) could not be confirmed against the primary source’s own text and was left out rather than included on the strength of secondary coverage.