Asking a room what might go wrong gets you polite hedging. Telling them it already failed gets you the thing somebody has been quietly worried about for a fortnight.
The pre-mortem comes from Gary Klein, published in Harvard Business Review in 2007. You gather the people who know the project, tell them it has failed spectacularly, and ask each of them to write down why. The reason it works better than a normal risk review is that the past tense removes the social cost of pessimism: you’re explaining a fact, not predicting a disaster. Klein designed it as a team exercise. The version in this article uses an AI assistant to run it on your own, which is a real adaptation with real limits, and it’s most useful when the team version isn’t going to happen. The prompt is tool-agnostic and ready to paste.
Every kickoff meeting has the same dead patch in it. Someone says “any concerns?”, there’s a pause, one person raises a scheduling question, everybody nods, and the meeting moves on. And then eight weeks later, when the thing has gone sideways, you find out that two people in that room had a bad feeling about the same thing and neither said it.
A pre-mortem fixes that with one change to the wording. Instead of asking what might go wrong, you tell everyone the project has already failed, and ask them to explain why.
That’s it. The past tense is the technique.
It comes from Gary Klein, a research psychologist, who published it in Harvard Business Review in September 2007 under the title “Performing a Project Premortem.”[1] His description of the mechanism is worth having exactly: a postmortem in a medical setting lets people learn what caused a patient’s death, and everyone benefits except the patient. A pre-mortem comes at the beginning instead, so the project can be improved rather than autopsied.
The process he describes is short enough to run properly in half an hour:
One thing to be straight about before we go further, because it shapes everything below. Klein designed this as a group exercise, and the group is doing real work in it: the point is partly to make it socially safe for the quiet person to say the awkward thing out loud. The version in this article runs it on your own with an AI assistant. That’s an adaptation, not the original, and it trades away the thing the original was best at.
Ask what did go wrong, not what might. The tense is doing the work.
The obvious explanation is that people are reluctant to be the pessimist in a room full of enthusiasm, and there’s a lot in that. Saying “I think this might fail” costs you something socially. Saying “here’s why it failed” when everyone has agreed to pretend it did costs you nothing, because you’re just playing along with the exercise.
There’s a second thing going on though, and it’s about how explanations work. Once an outcome is settled, your brain reaches for causes differently. You stop generating vague categories of risk and start telling a specific story, because a story is what an explanation needs.
Watch what that does to the same underlying worry:
| Asked “what might go wrong?” | Asked “why did it fail?” |
|---|---|
| “There’s some dependency risk on the agency side.” | “The agency’s creative lead went on leave in week three and the person covering had never seen the brand guidelines. We got two rounds of work back that we couldn’t use and lost a fortnight.” |
| “We should keep an eye on stakeholder alignment.” | “Finance never actually signed off on the media budget. They said they’d ‘come back on it’ in the kickoff and nobody chased. In week six they came back on it.” |
| “Timelines are tight.” | “We launched the week of the trade show, so nobody in sales was available to follow up the leads for eight days, and by the time they called, the leads had gone cold.” |
Illustrative examples of the shift the past tense produces. The left column is a category of risk. The right column is something you can put a name and a date against.
The right-hand column is actionable and the left isn’t, and it’s the same person with the same information both times. Nothing new was learned between the two columns. The question just gave them somewhere to put what they already knew.
Klein anchors this in research he cites from 1989, by Deborah Mitchell at Wharton, Jay Russo at Cornell and Nancy Pennington at the University of Colorado. As he reports it, prospective hindsight, meaning imagining that an event has already occurred, increases the ability to correctly identify reasons for future outcomes by 30%.[1] Worth holding that number carefully: I’ve read Klein’s article in full but the underlying paper sits behind a journal paywall, so I’m passing on his characterisation of it rather than something I’ve checked myself.[2] It’s also a lab finding about explaining events, from 1989, not a measurement of what pre-mortems do to real projects. Nobody has that number.
The examples in Klein’s own article are the more persuasive evidence, honestly. In one session, a team member who’d stayed silent through the entire kickoff volunteered that one of the algorithms wouldn’t fit easily on the laptops being used in the field, so the software would take hours to run when users needed fast results. It turned out the developers already had a shortcut and had been reluctant to mention it.[1] That’s not a risk a register would have caught. It was sitting in someone’s head, unspoken, for reasons that had nothing to do with the risk itself.
This is written to work in any assistant, so there’s nothing Copilot-specific or Claude-specific in it. Paste it, fill the square brackets, and give it enough real detail that the output is about your project rather than about projects in general. That last part is where most people undercook it.
| You are running a pre-mortem with me, using Gary Klein’s technique. Here is the plan: [WHAT WE’RE DOING, in 4 to 6 sentences. Include the actual dates, the budget if there is one, who is involved by role, and what we’re assuming will be true.] It is now [DATE, 3 to 6 months after the plan ends]. This project failed. Not slightly disappointing: it clearly failed, and everyone involved agrees it did. Write 12 separate explanations of why it failed. Rules: – Each one is a short specific story with a rough date, not a category of risk. “Stakeholder misalignment” is not an answer. “Finance never signed off the budget and came back on it in week six” is. – At least three of them must be things people would have been reluctant to say out loud at the kickoff, for political or personal reasons rather than technical ones. – At least two must be about something going right in a way that caused a problem. – Do not repeat the same underlying cause with different wording. Then, separately: tell me which of your 12 you think I am least likely to have already considered, and why. Finally, list what you would have needed to know about this project to do this better. Be specific about what I did not tell you. |
Tool-agnostic. The three constraints in the middle are what stop it returning a generic risk register with your project name at the top.
Three notes on adapting it.
The “reluctant to say out loud” constraint is the important one. Take it out and you get a risk register. Left in, you get the things that actually sink projects: someone senior having a pet view, a team being under-resourced in a way nobody wants to name, a supplier relationship everyone knows is strained.
The “something going right” constraint catches a whole category people forget. The campaign worked and support couldn’t cope. The hiring push landed and there was nowhere to put anyone. In marketing especially, success is a load you have to plan for.
The last instruction is a quality check on yourself. When it tells you what it needed to know, that’s usually a list of things you also haven’t written down anywhere, which is its own finding.
Say you’re relaunching your pricing page, moving from three tiers to two, with a supporting email sequence and a paid push. Six weeks of build, launch on 1 October, roughly forty thousand in media behind it. Everybody’s keen. The kickoff had no concerns in it.
Here’s a representative slice of what comes back from the prompt above, and I’ve kept both the good and the useless so you can see the ratio you should expect.
| Verdict | The explanation it generated | What you do with it |
|---|---|---|
| Useful | “Existing customers on the middle tier saw the new page before anyone told them what happened to their plan. Support took 200 tickets in four days and the sales team found out from a customer.” | Real. Nobody had a customer communication plan. This became a hard dependency on launch. |
| Useful | “The paid campaign drove traffic to the new page while the old pricing was still cached and indexed, so some visitors compared two versions and rang to ask which was true.” | Real and cheap to prevent. Went on the launch checklist. |
| Useful, and nobody would have said it | “The two-tier structure was decided by the founder in a meeting nobody minuted, and the team spent six weeks building it while privately thinking three tiers was right. Nobody re-opened it because it felt settled.” | Uncomfortable, and worth one direct conversation before build starts rather than after. |
| Generic | “Insufficient stakeholder alignment led to delays.” | Nothing. This is the category-not-story failure the prompt is designed to reduce, and some still gets through. |
| Wrong for us | “The pricing change triggered a contractual renegotiation with enterprise accounts.” | Discard. We don’t have enterprise contracts. Useful signal that it’s guessing where I gave it nothing. |
A composite illustration of a typical output, not a transcript from a named company. The ratio is realistic: roughly a quarter genuinely useful, and the useful ones are worth the whole exercise.
Three out of twelve was a good run. That sounds like a poor hit rate until you compare it to the kickoff meeting, which produced nothing at all.
The third row is the one that justifies doing this. An assistant with no career at your company will cheerfully write down “the founder decided this and nobody wanted to reopen it,” which is precisely the sentence a human on the team has every reason not to write. That’s a genuine and slightly odd advantage of running it solo, and it’s the opposite of what I expected the first time.
What it can’t do is know your business. Every generic and wrong item in that list came from the same place: I didn’t give it enough context, and it filled the gap. More detail in the brackets, fewer of those.
Running one on everything is how a good technique becomes a thing people roll their eyes at in your team, so it’s worth having a threshold you can say out loud.
Mine is roughly: run one when the decision is hard to reverse, or when the cost of finding out late is much bigger than the cost of finding out now. Those two conditions catch most of what matters and exclude most of what doesn’t.
| Worth running one | Probably not |
|---|---|
| A rebrand, a pricing change, a website migration. Hard to reverse, visible to customers. | A single campaign in a channel you run every month. You already know how this goes wrong. |
| Anything where the plan depends on a third party doing something on time. | Reversible experiments with a small budget. Just run it and find out. |
| A project everyone is unusually enthusiastic about. Enthusiasm is a risk factor, not a good sign. | Decisions you have to make today. A pre-mortem you rush is theatre. |
| Anything you’re doing for the first time, especially if a competitor has done it and you haven’t asked how it went for them. | Work where the failure mode is genuinely known and already mitigated. |
Worked examples of where the threshold falls in practice, not a scoring model. Your own line will sit somewhere slightly different.
Where the threshold sits also depends on what your job is. A marketing lead should be running these on anything customer-visible, because the failure is public and permanent. A product manager cares more about the sequencing failures, the ones where two dependencies collide in week five. A finance or operations lead usually has the opposite problem: their risks are well documented already, and the ones a pre-mortem adds are the political ones nobody minutes. Same technique, different thing it’s catching for you.
On solo versus team, they’re different exercises and I’d use both.
The team version is the better one where you can get it, because half of what it does is social. It gives the person who’s been worried since week one a legitimate reason to say so, and it lets everyone else discover that three of them were worried about the same thing. An AI assistant can’t manufacture that. What the solo version has going for it is that it takes ten minutes and needs nobody’s calendar, which means it happens.
Run the solo version when the team version isn’t going to happen. It’s a floor, not a replacement.
The combination I like best: run the solo one first, then take the three or four uncomfortable items into the team session as opening questions. You’ve done the awkward part already, so nobody has to be the one who raised it.
This is where pre-mortems usually die. You get twelve explanations, they’re interesting, everyone agrees they’re interesting, the document goes in the project folder, and the project proceeds exactly as planned.
The fix is to stop sorting by likelihood. Likelihood is the instinct and it’s the wrong axis, because the likely risks are the ones you’ve already thought about. Sort by how late you’d find out.
| Score each risk | 1 | 2 | 3 |
|---|---|---|---|
| How late would you find out? | Immediately, it’s obvious | Within a week or two | Only after it has already cost you |
| How hard is it to undo? | Easy, just change it back | Costly but possible | Public, or contractual, or both |
| How cheap is the prevention? | A project in itself | A few days’ work | One conversation or one line on a checklist |
Multiply the three. Anything scoring 18 or above gets an owner and a date this week. In a normal twelve-item output, that’s usually two or three items, which is the number a team can actually act on.
The risk worth acting on isn’t the likeliest. It’s the one you’d only notice too late.
Then be explicit about who is doing what, including in the bit where AI was involved, because “we ran a pre-mortem” can quietly turn into “the model told us the risks” and those are very different things.
| What the AI does | What you still own | How it gets checked |
|---|---|---|
| Generates twelve specific failure stories, including ones a colleague would be reluctant to raise | Which of them are real for your business, and which are it guessing at gaps in your brief | Anything you can’t connect to something true about your company gets discarded, not softened. |
| Points out which explanations you were least likely to have considered | Whether that’s genuine insight or just the most unusual-sounding item | You test one of them with a person who’d know, before it goes on the plan. |
| Nothing at the sorting stage | The scoring, the threshold, and the decision about what changes | Every item above threshold gets a named owner and a date, in the plan itself rather than in the pre-mortem document. |
Fill this in before you run the exercise, so the boundary is set while the output is still hypothetical.
For each item that clears the threshold, four things go into the project plan itself rather than into the pre-mortem document, which is the step that decides whether any of this was worth doing:
One thing worth measuring, and it isn’t how many pre-mortems you ran. Six months on, look at whatever did go wrong and check it against the list. Was it on there? If it was on there and you didn’t act, your problem is the sorting step rather than the technique, and if it wasn’t on there at all, look at what you didn’t tell the model, because that’s usually the answer.
Next time you’ve got something on the calendar that’s hard to undo, block ten minutes the day before the kickoff. Paste the prompt, fill in the brackets properly, and take the three most uncomfortable items into the room as your opening questions. That’s the whole thing. If you want more of these, we keep a running set in our piece on building an AI prompt library for your team, which is also where a prompt like this should end up once it’s earned its place.
A pre-mortem asks you to assume the project has already failed and explain why, rather than asking what might go wrong. The difference sounds cosmetic and isn’t. A risk assessment produces categories (“dependency risk”, “stakeholder alignment”) because that’s what the future tense invites. A pre-mortem produces stories with dates and names in them, because explaining a settled outcome requires a specific cause. It also removes the social cost of pessimism: raising a concern makes you the difficult one, while explaining a failure everyone has agreed to pretend happened makes you a good sport. Gary Klein, who published the technique in Harvard Business Review in 2007, describes it as the hypothetical opposite of a postmortem.
Gary Klein, a research psychologist, published “Performing a Project Premortem” in Harvard Business Review in September 2007. At the time he was chief scientist of Klein Associates, a division of Applied Research Associates. He attributes the underlying cognitive mechanism to 1989 research by Deborah Mitchell of the Wharton School, Jay Russo of Cornell and Nancy Pennington of the University of Colorado, on what they called prospective hindsight. As Klein reports it, imagining that an event has already occurred increases the ability to correctly identify reasons for future outcomes by 30%. That paper is paywalled, so this article passes on Klein’s characterisation of it rather than a first-hand reading.
Give the assistant your plan in four to six sentences including real dates, budget and roles, tell it a date several months after the plan ends, and state that the project clearly failed. Then ask for twelve separate explanations with three constraints: each must be a specific story with a rough date rather than a category of risk, at least three must be things people would have been reluctant to say out loud for political rather than technical reasons, and at least two must involve something going right in a way that caused a problem. Finish by asking which explanations you were least likely to have considered, and what it would have needed to know to do better. The full version is in this article, ready to paste.
Klein’s team version fits comfortably in half an hour: brief the plan, announce the failure, a few minutes of everyone writing independently, then one round of reading out different reasons, then the project manager reviews the list. The solo AI-assisted version takes about ten minutes, most of which is writing a good enough description of the plan, and that’s the part worth not rushing. The sorting step afterwards takes another ten. What makes the difference isn’t the length of the exercise, it’s whether anything goes into the actual project plan with an owner and a date attached, which is where most pre-mortems quietly stop.
With a team where you can get one, because roughly half of what the technique does is social: it gives the person who’s been uneasy since week one a legitimate way to say so, and it lets several people discover they’re worried about the same thing. An AI assistant can’t reproduce that. The solo version has one genuine advantage though, which is that an assistant with no stake in your company will happily write down uncomfortable explanations involving senior people that a colleague has every reason not to raise. The combination that works best is to run the solo version first, then take the three most uncomfortable items into the team session as opening questions.
Gary Klein’s Harvard Business Review article was read in full for this piece rather than summarised from secondary coverage, which matters because a lot of what circulates about pre-mortems garbles the process and drops the attribution. The process steps, the medical analogy, and the field example about algorithms not fitting on laptops all come from Klein’s own text. The 30% figure is reported as Klein reports it: the underlying 1989 paper by Mitchell, Russo and Pennington sits behind a journal paywall that couldn’t be read from here, so this article passes on his characterisation rather than pretending to a first-hand reading, and notes that it is a lab finding about explaining events rather than a measurement of pre-mortems in business. Nobody appears to have that second number. The worked example is a composite illustration of the shape and quality of a typical output, not a transcript from a named company, and the useful-to-generic ratio shown is realistic rather than flattering.