You already know roughly what you need to say. The hard part is saying it clearly, without softening it into meaninglessness or sharpening it into something you'll regret. That specific problem is one AI is unusually good at helping with.
Roughly seven in ten managers say they’re often uncomfortable communicating with employees, and 37% specifically dread giving direct performance feedback when they think it might land badly[1]. That discomfort has consequences: the classic meta-analysis of feedback interventions found that while feedback improved performance on average, over a third of interventions actually made performance worse[5]. Preparation is what separates the two outcomes. This guide covers a five-step method for using ChatGPT, Claude or Copilot to prepare: stress-testing your own reasoning, writing an opening that’s short and honest, rehearsing the reactions you’re dreading, auditing your language for bias, and knowing exactly what not to paste into a chat window.
A manager on one of our leadership programmes last year had to tell a long-serving team member that the promotion he’d been openly expecting wasn’t happening. She’d been carrying it around for nine days. When I asked what was stopping her, she said something I’ve heard almost verbatim from dozens of people since: “I know what I need to say. I just don’t know how to say it without either lying to him or destroying him.”
That gap, between knowing the substance and being able to deliver it, is where AI genuinely helps. It is not a moral advisor and it has no idea whether your decision is fair. What it can do is act as a tireless, uncritical, slightly pedantic rehearsal partner at eleven o’clock at night when the conversation is at nine tomorrow morning and you’ve read your notes eleven times.
Worth being precise about the size of the problem. A Harris Poll survey conducted for Interact, covering 2,058 U.S. adults including 616 managers, found 69% of managers said they’re often uncomfortable communicating with employees, and 37% specifically said they’re uncomfortable having to give direct feedback about performance if they think the employee might react badly[1]. So this isn’t a niche struggle among weak managers, it’s most managers.
And the avoidance has a cost that lands on everyone. The CIPD’s Good Work Index found a quarter of UK employees, roughly eight million people, had experienced workplace conflict in the past year, and the most common response was simply to “let it go” (47%)[9]. Conflict that nobody addresses does not resolve itself, it just stops being visible.
Here’s the honest counterweight, and I’d rather put it up front than bury it in a caveat at the end. Feedback is not automatically good for people. Kluger and DeNisi’s meta-analysis, covering 607 effect sizes and 23,663 observations, found feedback interventions improved performance on average (d = .41) but that over a third of them decreased performance[5]. Their explanation is the useful part: effectiveness drops as attention moves away from the task and toward the self[5]. A conversation about a missed deadline can go well. The same conversation, framed as being about who someone is, tends not to.
Almost every manager I work with wants to skip straight to “help me write what I’m going to say.” Resist that for ten minutes, because a beautifully worded message built on shaky reasoning is worse than a clumsy one built on solid reasoning.
Open a chat and describe the situation in your own words, with the identifying details removed. Then ask it to argue with you. Something like:
“I’m a manager preparing to tell a team member that their work on client reporting has been consistently late for three months. Here’s my reasoning: [your reasoning]. Before I write anything, I want you to challenge me. What am I assuming that I haven’t verified? What might be going on that I haven’t considered? What would this person’s version of events most plausibly be? Where might I be the problem? Be direct, don’t be reassuring.”
That last instruction matters more than people expect. Left to its defaults, a chatbot will validate you, because agreeable answers are what most users reward. You have to ask for friction explicitly, and sometimes twice.
What comes back is usually a mix of things you’d already considered and two or three you hadn’t. In my experience the highest-value output is the “what would their version be” question. Managers walk into hard conversations having rehearsed their own case in detail and the other person’s case not at all, and then get blindsided by an entirely predictable objection.
Acas, the UK’s workplace advisory service, makes a related point in its guidance on challenging conversations, framing the ability to talk about sensitive issues as “an integral part of effective line management” that is “critical to managing performance, promoting attendance and improving team dynamics”[8]. Preparation isn’t a nicety bolted onto the conversation. It largely determines what kind of conversation it becomes.
One question worth asking yourself before you go further, and answering honestly: is this conversation genuinely about the work, or is a chunk of it about how this person makes you feel? Both are real. Only one of them belongs in the meeting.
The opening decides the conversation. Get it wrong and you spend the rest of the meeting recovering. Two failure modes dominate, and I’ve watched both play out repeatedly in roleplay exercises during workshops.
The first is the cushion, where the manager opens with so much warm-up that the other person is genuinely confused about whether this is bad news at all, and then feels ambushed when it arrives. The second is the cold open, where the manager, terrified of waffling, leads with the harshest possible sentence and the other person stops hearing anything after it.
What works sits between: brief context, the actual point, then space. Acas’s own guidance recommends preparing a script and is unembarrassed about it: “A pre-prepared script can help you keep on track and in control of the meeting. It’s a bit like planning a few moves ahead in a game of chess,” adding “Don’t be afraid of referring to your pre-prepared script, it will help you stay in control”[8].
Draft your opening yourself first, badly, in your own words. Then hand it over:
“Here’s how I plan to open a difficult conversation with a team member: [your draft]. Cut it to under 90 seconds spoken. Keep my voice, don’t make it sound corporate. Tell me where I’m hedging, where I’m being vague to avoid discomfort, and where a reasonable person could walk away genuinely unsure what I meant. Then give me one alternative opening that’s more direct, so I can compare.”
The hedging audit is the part that earns its keep. We all do it, and we’re bad at spotting it in our own writing. “There have been some concerns raised about aspects of the reporting timeline” is a sentence a manager writes when what they mean is “the last four client reports were late.”
Draft it yourself first, though. If you let the AI write from scratch you’ll get something fluent, generic and unmistakably not you, and the person across the table will hear it. There’s now research on that instinct: a controlled study of 261 participants and 990 evaluations found that disclosing AI authorship “generally erodes perceived trustworthiness, caring, competence, and likability, with the most precipitous declines observed in social and interpersonal writing”[7]. Nobody’s going to disclose anything in a one-to-one, but the underlying signal is the same. Interpersonal words that don’t sound like you carry a cost.
This is the step people skip and the one with the most evidence behind it.
Roleplay has a slightly cringeworthy reputation in corporate training, largely because it’s usually done badly, in front of colleagues, with an audience. Done privately with an AI, the awkwardness disappears and the benefit remains. A randomised controlled study of a four-hour simulation-based training for breaking bad news in emergency medicine, comparing 37 trainees against 31 controls, found significant improvements in self-efficacy, in the process itself, and in communication skills, with a 33.3% gain on the process measure[6]. Different context, same underlying mechanism: rehearsing the hard interaction changes how you perform in it.
Acas says the same thing in one line: “Testing yourself in roleplaying can be a useful way of practising your skills”[8].
The trick is to be specific about which reaction you want to practise, because a generic roleplay will hand you a reasonable, cooperative employee, which is not the version keeping you awake.
The four reaction types recommended in this section for separate rehearsal. Author’s framework, not survey data.
Run each one as its own conversation:
“Roleplay as an employee being told their work has been consistently late. Play them as calm on the surface but quietly angry, and make them push back by blaming the handover process rather than accepting the point. Stay in character. Respond only as them, one turn at a time, and wait for me. Don’t be easy on me. After ten exchanges, drop character and tell me the three moments where my responses were weakest.”
Two things I’d flag from running this with managers. Type 4, the good argument, is the one nobody rehearses and the one that most often derails a real conversation, because managers treat “I don’t have an answer to that right now” as a defeat rather than a perfectly acceptable sentence. And the debrief instruction at the end matters. Without it you get practice; with it you get practice plus a diagnosis.
If your team runs regular one-to-ones, a lot of this pressure disappears upstream. Gallup found employees are 3.6 times more likely to strongly agree they’re motivated to do outstanding work when their manager gives daily rather than annual feedback, and that 80% of employees who received meaningful feedback in the past week are fully engaged[3]. Most genuinely difficult conversations are the accumulated interest on feedback that was never given.
This is the use case I’d argue is most underused, and it’s the one where AI does something a colleague realistically can’t, because you’d have to ask them to read your notes and then be honest about your blind spots.
Written workplace feedback carries measurable, well-documented patterns. Textio’s 2024 analysis of language bias in performance feedback found that women are seven times more likely than men to internalise negative stereotypes of themselves such as “emotional,” that men are four times more likely than people of other genders to be positively stereotyped as “likable,” and that white and Asian people are twice as likely to be positively stereotyped as “intelligent” compared with Hispanic, Latino and Black colleagues[4]. Their bluntest finding: top performers get the lowest-quality feedback, and it’s worst for high-performing women[4].
None of that shows up when you reread your own notes, because it reads as normal to you. That’s what makes it worth an external check.
“Here are my notes for a performance conversation: [paste, with names and identifying details removed]. Review the language only. Flag anything that describes personality rather than behaviour. Flag anything that would be hard to evidence if challenged. Flag adjectives that tend to be applied unevenly by gender or race in performance feedback. For each flag, give me a behaviour-based alternative. Don’t rewrite the whole thing, just show me the specific words.”
The personality-versus-behaviour distinction is doing double duty here. It reduces bias exposure, and it also lines up with what makes feedback work at all: Kluger and DeNisi found effectiveness falls as attention moves “up the hierarchy closer to the self and away from the task”[5]. “You’re not detail-oriented” is a self-directed statement with no route to improvement. “Three of the last four reports had figures that didn’t reconcile” is a task-directed one that a person can actually do something about.
Our guide on using AI for performance reviews goes deeper on the written side of this, and if the conversation you’re preparing for is heading toward a formal process, writing a performance improvement plan with AI covers what comes next.
Here’s where I get slightly stern, because this is the part that can turn a well-intentioned prep session into an HR incident.
Whether your conversation notes are safe depends on which account you’re using, and the vendors are actually quite clear about it if you read their own documentation rather than assuming.
OpenAI states that for services for individuals such as ChatGPT, “we may use your content to train our models,” with an opt-out available in the privacy portal, while “by default, we do not train on any inputs or outputs from our products for business users, including ChatGPT Team, ChatGPT Enterprise, and the API”[10]. Anthropic’s consumer terms update states that it will train new models using data from Free, Pro and Max accounts when the setting is on, and that these updates “do not apply to services under our Commercial Terms, including: Claude for Work, which includes our Team and Enterprise plans”[11]. Microsoft’s enterprise data protection documentation states plainly that “the prompts, responses, and data accessed through Microsoft Graph aren’t used to train foundation models”[12].
| Keep out entirely | Fine to include |
|---|---|
| Real names, job titles specific enough to identify someone, team names in a small org | “A team member”, “an analyst on my team of six” |
| Salary figures, bonus numbers, compensation bands | “A pay decision they will disagree with” |
| Health information, disability, medication, anything disclosed in confidence | “There are personal circumstances I’ve been made aware of and can’t discuss here” |
| Verbatim quotes from other employees, grievance or investigation content | Your own paraphrase of the behaviour pattern |
| Client names, contract values, anything commercially confidential | “A key account”, “a major deliverable” |
Summary of the data-handling guidance in this section. Use a work-tier AI account where your organisation provides one.
The practical rule I give managers: write your prep as though it might be read by someone you’ve never met, because on a consumer account that is a documented possibility rather than a paranoid hypothetical. If your organisation provides a business-tier account, use it, and if it doesn’t, that’s a conversation worth having with whoever owns the AI policy.
There’s a second reason to keep specifics out that has nothing to do with data policy. Formal processes have evidential weight. The Acas Code of Practice on disciplinary and grievance procedures notes that “employers would be well advised to keep a written record of any disciplinary or grievances cases they deal with,” and that tribunals “will be able to adjust any awards made in relevant cases by up to 25 per cent for unreasonable failure to comply with any provision of the Code”[13]. If you’re anywhere near a formal process, your notes need to be accurate, contemporaneous and human-authored, and your HR team needs to see them before your chatbot does.
Everything above is preparation. The conversation itself is unassisted, and it should be.
What makes these conversations land is not word choice. It’s that the person across the table believes you’re being straight with them and that you’ll still be reasonable tomorrow. Amy Edmondson’s foundational work on psychological safety, a study of 51 work teams, established it as “a shared belief held by members of a team that the team is safe for interpersonal risk taking,” and found learning behaviour mediates between psychological safety and team performance[2]. Google’s own Project Aristotle research across 180 teams reached a similar conclusion, finding that what mattered “was less about who is on the team, and more about how the team worked together,” with psychological safety ranking first among five dynamics[14].
You cannot outsource being trusted. A perfectly scripted conversation delivered by someone who has ducked every hard truth for two years will not work, and an awkward one delivered by someone with a track record of straightness usually will.
The World Economic Forum’s skills research makes the same point from the labour-market side. Analysing which skills are exposed to AI substitution, it found “skills rooted in human interaction, including empathy and active listening, and sensory processing abilities” currently show no substitution potential “due to their physical and deeply human components,” with 69% of examined skills having low or very low capacity for substitution[15]. Meanwhile leadership and social influence saw one of the largest jumps in reported importance, up 22 percentage points on the previous edition[15].
Reading that alongside the SHRM finding that nearly 45% of U.S. workers now use AI in their jobs, with 74% agreeing AI should complement human talent rather than replace it[16], the shape of the thing is fairly clear. AI is getting good at the preparation. The conversation is still yours.
My honest view, having watched a lot of managers go through this: the preparation isn’t really about producing better sentences. It’s that walking in having already heard the worst version of their response, out loud, twice, means you’re no longer bracing. And people can tell.
Two things happen in the twenty minutes after a hard conversation. You feel enormous relief, and you start forgetting details at speed. Write it down before you do anything else.
Acas’s guidance is direct on this: “Document any agreement and give a copy to the employee”[8], and the Code of Practice adds that employees should be informed “of the basis of the problem and given an opportunity to put their case in response before any decisions are made”[13].
Write the record yourself, from memory, immediately. Facts, what was agreed, dates, and what the other person said in their own words as closely as you can recall. This is not a document to generate. Once it’s drafted, AI is useful for one narrow job: checking it for language that describes a person rather than an event, and for anything that reads as a conclusion you didn’t actually reach in the room.
“Here’s my note of a performance conversation, with identifying details removed: [paste]. Don’t rewrite it. Tell me: where have I recorded an interpretation as if it were a fact? Where is a commitment vague enough that two people could read it differently? What did I fail to record that a reader in three months would need?”
That third question catches things surprisingly often. The commitment you both understood perfectly in the room is routinely the one that turns out, three months later, to have had no date attached.
Then book the follow-up before you leave your desk. Given how strongly Gallup’s data ties frequent feedback to engagement[3], and given that managers account for at least 70% of the variance in employee engagement scores across business units[17], the conversation you just had matters considerably less than whether anything changes in the six weeks after it.
If you want the underlying prompting habits rather than the specific ones above, ChatGPT prompts for managers covers the wider set, and using AI for employee disciplinary write-ups handles the formal documentation side properly.
It depends entirely on which account you use and what you paste in. OpenAI’s own documentation states that content from its individual consumer services may be used to train models unless you opt out, while ChatGPT Team, Enterprise and API accounts are excluded by default. Anthropic and Microsoft publish similar distinctions between consumer and commercial tiers. The practical rule is to use a business-tier account where your organisation provides one, and to strip out names, job titles specific enough to identify someone, salary figures, health information and anything disclosed in confidence regardless of which account you use. You can describe a situation accurately without describing a person identifiably.
No, and treating it as though it can is the main risk in this whole approach. AI has no access to the context, the history, the previous conversations, or the things you know and haven’t written down. What it can do usefully is challenge your reasoning: surface assumptions you haven’t verified, construct the most plausible version of the employee’s account, and point out where your evidence is thinner than you think. That’s genuinely valuable, and it’s a different thing from a fairness verdict. If the fairness question is live, that’s a conversation for your HR partner or a trusted peer, both of whom carry accountability that a chatbot does not.
Five things, in order. First, argue against your reasoning and construct the employee’s likely version of events. Second, tighten an opening you drafted yourself down to under ninety seconds and flag where you’re hedging. Third, roleplay the specific reactions you’re dreading, one at a time, with instructions to stay in character and debrief you afterwards. Fourth, audit your written notes for language that describes personality rather than behaviour. Fifth, after the conversation, check your written record for interpretations recorded as facts and for commitments vague enough to be read two ways. Notice that none of these ask it to write the conversation for you.
The evidence on behavioural rehearsal generally is strong. A randomised controlled trial of simulation-based training for breaking bad news in emergency medicine found significant improvements in self-efficacy, process and communication skills, including a 33.3% gain on the process measure, against a control group. The mechanism is the same whether your rehearsal partner is a trained actor or a chatbot: you’ve already encountered the difficult moment once, so you’re not encountering it cold. The procrastination risk is real though. If you’ve run four roleplays and you’re setting up a fifth, you’re no longer preparing, you’re avoiding. Two or three specific reactions, then stop.
There’s no obligation to, and in most cases it would be an odd thing to raise, in the same way you wouldn’t announce that you’d rehearsed with a colleague or made notes beforehand. Preparing thoroughly is part of the job. What matters is that the words you say are yours and the decisions are yours. Research on AI authorship disclosure found it generally erodes perceived trustworthiness, caring and competence, with the sharpest declines in social and interpersonal writing, which is worth knowing if you’re tempted to send an AI-drafted follow-up email. Write the follow-up yourself. Use AI to check it, not to compose it.