You launched it, it did fine, you moved on to the next one. That habit is costing you more than any channel test you have ever run.
Debriefs work: Tannenbaum and Cerasoli’s meta-analysis of 46 samples found properly conducted debriefs improve effectiveness over a control group by approximately 25%[3], and a later meta-analysis of 83 studies found the effect is largest for complex, ambiguous tasks that give no natural feedback, which describes a marketing campaign exactly[4]. The catch is that knowing how something turned out actively distorts your judgement of it[5]. This guide covers why most retros are useless, the four-question structure that fixes it, the exact AI prompts for each stage, the data you need to grab before Google deletes it, and how to end with decisions rather than a document.
I ran a campaign a few years back that beat its lead target by about 40%. We had a debrief. It lasted twenty-five minutes, everyone agreed the creative had been strong, and we went to lunch.
Six months later we reused that creative approach on a different product and it did nothing at all. Which is when I went back and looked properly, and found that the original campaign had launched three days after a competitor’s very public outage. We hadn’t won on creative. We’d won on timing we didn’t plan and hadn’t noticed.
That’s the whole problem with retrospectives in one story. When you know the outcome, you build a tidy explanation for it, and the explanation feels obvious.
This is a well-documented cognitive effect rather than a personal failing. Fischhoff’s 1975 experiments found that simply telling people how something turned out increased its perceived likelihood in every one of 24 tests, roughly doubling how inevitable the outcome seemed[5]. His conclusion is worth reading slowly: “the very outcome knowledge which gives us the feeling that we understand what the past was all about may prevent us from learning anything from it”[5]. And people don’t notice it happening. The same work found judges believed the inevitability had been apparent in foresight all along.
There’s a related trap. Outcome bias means a good result makes the decision behind it look good, even when the decision was poor and you got lucky. A large pre-registered replication of Baron and Hershey’s classic study found the effect held with stronger effect sizes than the original, and persisted even among participants who explicitly said outcomes shouldn’t matter when evaluating decisions[6].
A typical marketing retro, the “let’s look at the dashboard and talk about what happened” kind, makes it surprisingly easy to build a confident story after the fact.
How common is this? Less than you’d hope, going by PMI’s numbers.
| Metric | Figure |
|---|---|
| Organizations highly effective at knowledge transfer | 14% |
| Original goals met, with a formal process | 62% |
| Original goals met, without one | 48% |
PMI survey of more than 2,800 project leaders. High performers were more than twice as likely to have a formal knowledge-transfer process in place[7].
I’d note honestly that this is project data rather than marketing data, and it’s from 2015. I looked for a credible figure on what share of marketing teams never run a post-campaign review and couldn’t find one that traced back to anything solid. Every version I chased ended at a vendor blog quoting another vendor blog. So I’m not going to give you a percentage I can’t stand behind.
What I will say from ten years of running marketing teams: the review gets skipped because the next campaign is already late, and because a badly-run one genuinely is a waste of an hour, so everyone has learned to protect their calendar from it.
The best-developed version of this practice doesn’t come from marketing or from agile software teams. It comes from the US Army, which has been formalising after-action reviews since the early nineties and has thought harder about the failure modes than anyone in our industry.
The US Government’s official adaptation of that doctrine states it plainly: an after-action review answers four questions. What was expected to happen? What actually occurred? What went well, and why? What can be improved, and how?[2]
That first question is the one marketing teams skip, and skipping it is what lets hindsight bias walk in the door. If you don’t state what you expected before you look at what happened, you make it much easier to reconstruct your expectations around the result.
The four AAR questions as stated in US Government guidance derived from Army doctrine, adapted here for marketing campaigns. The marketing framing is the author’s.
Two more things from that doctrine that translate directly, and that I’d argue matter more than the questions.
It is not a critique. The guidance is explicit: “An AAR is not a critique or a complaint session. No one, regardless of rank, position, or strength of personality has all of the information or answers”[2]. The moment a retro becomes performance review, people stop telling you things.
Statistics can wreck it. Army doctrine warns that piling on charts and performance ratios turns the session into a graded scorecard, which shuts down discussion instead of opening it up[1]. Written in 1993, and it’s a perfect description of every campaign retro that opens with someone screen-sharing a Looker Studio dashboard.
On timing and length, the official guidance says hold it as soon as possible and within two weeks, and cap it at 90 minutes[2]. I’d aim to run it within two weeks and cap it around 90 minutes for marketing too. Beyond that, memory gets fuzzier and the discussion usually starts to lose energy.
And the reason to bother at all is straightforward: the doctrine is clear that the real payoff comes from applying what you found to the next round of work, not from the document itself[1]. Swap “training” for “campaigns” and that’s your success criterion.
Before anything else, and this catches people out badly: your campaign data has an expiry date, and it’s shorter than you think.
Google Analytics 4 lets you set user-level and event-level data retention to 2 months or 14 months, with longer options only on the 360 paid tier[8], and once that window closes the data is gone for good. The catch most people miss: that limit only touches explorations and funnel reports, not your standard reports[8]. So your headline numbers survive, and the granular path and funnel analysis, which is exactly where a retrospective gets its interesting findings, quietly does not.
Pull this into one folder. It takes about twenty minutes and it’s the difference between a real retro and a vibes conversation.
One honest note on measurement generally: even well-resourced teams struggle here, so if yours is messy, you’re not the exception.
| Metric | Figure |
|---|---|
| Global marketers confident in full-funnel measurement | 54% |
| Using multiple measurement tools for one campaign | 62% |
| “Demonstrating ROI” score, out of 7 | 4.2 |
| “Generating ROI” score, out of 7 — for comparison | 4.5 |
Nielsen[10] and The CMO Survey’s 2026 edition, 308 US marketing leaders, Duke Fuqua School of Business[11], which describes the ROI gap as pointing to “persistent measurement and attribution challenges.”
So if your data is messier than you’d like, that’s the norm rather than an indictment. Run the retro anyway. An honest conversation about imperfect data beats no conversation about perfect data, and you’ll never have perfect data.
AI is useful in a campaign retrospective in three specific places, and actively harmful in a fourth. Let me be clear about the harmful one first: do not ask AI to tell you why the campaign performed the way it did. It will produce a fluent, plausible narrative built on the same hindsight bias you’re trying to escape, and because it sounds authoritative it’ll anchor the whole room. It doesn’t know what happened. It’s pattern-matching your framing back at you.
Adoption of AI for this kind of work has moved fast, incidentally. The CMO Survey shows marketing teams’ use of AI for “data analysis and reporting: to measure performance, track metrics, and generate reports” has nearly doubled in about two years[11]. Which makes it worth being precise about what it’s good for.
This is the least glamorous use and the most valuable. You’ve got a Slack channel, an email thread, a project board, and four platform exports. Getting those into one chronology by hand takes an hour and everyone hates it.
Paste your exports and message history into a document and use something like this:
That last sentence matters more than it looks. Without it you get a smoothed narrative. With it you get the contradictions, and the contradictions are usually where the interesting stuff is hiding.
Every team has questions it politely avoids. AI can be useful for generating neutral challenge questions because it isn’t participating in the team’s dynamics or hierarchy.
Run this before the session and put the questions on the agenda. It takes the personal sting out of asking them, because nobody in the room chose them.
Record the session (with everyone’s agreement) and use the transcript.
The third category is the one I care about most. Unresolved disagreements are usually the most valuable output of a retro and they’re exactly what gets sanded off in a tidy summary. If two experienced people disagree about why something worked, you’ve found the thing worth testing next.
| AI can help with | Humans still own |
|---|---|
| Reconstructing the timeline | Deciding what matters |
| Finding contradictions | Interpreting why they matter |
| Generating challenge questions | Having the uncomfortable conversation |
| Extracting decisions and disagreements | Choosing what to change |
| Organizing evidence | Deciding what is a lesson vs. a hypothesis |
AI structures the material a retrospective runs on. It doesn’t make the judgement calls that make the retrospective worth having.
Campaign data often contains customer information, and internal messages contain things people said assuming a small audience. Before uploading campaign data, transcripts, or internal messages, check your organisation’s approved AI environment and data-handling policy. Enterprise products may provide stronger protections, but approval depends on your specific setup. Our guide to summarising long documents with ChatGPT covers the practical side of getting good output from large pasted material.
You can have perfect data and the right four questions and still get nothing, because the person who knows what actually went wrong has decided it’s not worth saying.
Amy Edmondson’s work on this is the reference point. Her study of 51 work teams found that team psychological safety, the shared belief that the team is safe for interpersonal risk-taking, was associated with learning behaviour, while team efficacy was not once you controlled for psychological safety[12]. Confidence doesn’t substitute for safety. And she identifies the leader as the decisive signal: where leaders act in “authoritarian or punitive ways, team members may be reluctant to engage in the interpersonal risk involved in learning behaviors such as discussing errors”[12].
Which means how you open the session does most of the work.
Go first, and go badly. If you ran the campaign, name your own worst call before anyone else speaks. Not a humble-brag about caring too much. An actual mistake with a consequence. Everything after that is calibrated to what you just modelled.
Ask open questions, not “why didn’t you” questions. Army doctrine makes this point with an example that translates perfectly: better to ask what happened at a specific moment than to ask why someone didn’t do the thing you’d have done, because the second version “put[s] him on the defensive” and you get less of the truth.
Watch the seniority ordering. The Army physically seats junior people at the front and commanders behind them, explicitly so that “soldiers taking part in the AAR feel that their comments are valued as much as those of senior leaders”[1]. You can’t rearrange a Zoom call, but you can make the most senior person speak last on every question. It’s a small thing that changes what gets said.
Do question one with the dashboards closed. Genuinely closed. Read the original brief out loud, have people say what they expected, and only then open the numbers. This one change does more for the quality of a retro than any tool.
If it helps, this is roughly what I run:
Notice there isn’t a long standalone results presentation. The numbers should support the discussion, not become the meeting.
If the retrospective ends only in a document, very little has changed. I’ve written plenty of those documents. Nobody has ever opened one.
The failure is structural rather than lazy. The lessons get captured somewhere nobody looks at the moment when they’d be useful, which is when the next campaign brief is being written.
So the only output that matters is a small number of changes attached to the next campaign’s process, not to a lessons-learned register.
My rule is a maximum of three changes. A list of eleven improvements is a list of zero improvements. Pick the three with the best ratio of impact to effort and let the rest go. You’ll get another retro next campaign.
Each change edits an artefact, not an intention. “Be more careful about audience overlap” is an intention and will evaporate. “Add an audience overlap check to the pre-launch checklist, owned by Priya, before the September campaign” edits a real thing that a real person will open. If a change doesn’t correspond to an edit in a brief template, a checklist, a QA step or a calendar entry, it isn’t a change.
Separate lessons from hypotheses, and be strict about it. A lesson is something you now know. A hypothesis is something you suspect and are going to test. Most retro output is hypotheses wearing a lesson’s clothes, and treating a hypothesis as settled is precisely how a team convinces itself it understands its own market.
My rough test: if you can explain the mechanism, point to evidence, and would expect it to repeat, treat it as a lesson. Otherwise it goes on the test list.
At the start of the next campaign’s kickoff, spend five minutes reading out the three changes from last time and confirming each one actually happened. Five minutes. That single habit is what converts a retrospective from a ritual into a system, and it’s the thing almost nobody does.
PMI’s report quotes a practitioner on the cost of not doing it, and it’s blunter than anything I’d write: “We would find on similar projects that we repeated the same mistakes”[7].
The evidence that this is worth the hour is genuinely strong, and it’s worth ending on. Tannenbaum and Cerasoli’s meta-analysis of 46 samples concluded that “organizations can improve individual and team performance by approximately 20% to 25% by using properly conducted debriefs”[3], and the follow-up meta-analysis of 83 studies found the effect strongest “in task environments that are characterized by a combination of high complexity and ambiguity in terms of offering no intrinsic feedback,” precisely the industries that don’t traditionally run one[4].
Complex, ambiguous, no intrinsic feedback, and not traditionally doing this. That’s marketing. The research suggests well-run debriefs can materially improve team effectiveness, and marketing happens to be exactly the kind of complex, ambiguous work where structured reflection is useful. The opportunity is not to squeeze 20% more out of every campaign. It is to stop paying for the same lesson twice.
If you want to strengthen the measurement side before your next retro, our guide to using AI for marketing attribution covers working out what actually drove the result, and the GA4 reporting walkthrough will help you pull cleaner exports without a data analyst. For the wider question of proving value, the honest guide to measuring AI ROI is a good companion.
A results report presents what happened. A retrospective works out what to do differently, which is a different activity with a different structure. The best-developed version comes from US Army after-action review doctrine, and the US Government guidance derived from it states four questions: what was expected to happen, what actually occurred, what went well and why, and what can be improved and how. The first question is the one marketing teams skip, and skipping it is what allows people to reconstruct their expectations to match the result.
Within two weeks, and cap the session at 90 minutes. That is the guidance in official after-action review documentation, on the reasoning that participants get better feedback and remember the lessons longer when the review is timely but not rushed. For marketing specifically there is a second reason to move fast: your granular analytics data may be on a short clock. GA4 lets you set event and user-level retention to as little as 2 months, and that limit affects explorations and funnel reports, which is exactly where retrospective analysis happens.
Use it for three things: building a single chronological timeline from scattered exports and message history, generating the sceptical questions nobody in the room wants to ask, and extracting decisions, open questions and unresolved disagreements from the session transcript. Do not ask it to explain why the campaign performed as it did. It will produce a fluent narrative built from your own framing, it will sound authoritative, and it will anchor the whole discussion around a cause it has no way of knowing.
Because knowing the outcome changes how you see everything that led to it. Fischhoff’s experiments found that reporting an outcome increased its perceived likelihood in every one of 24 tests, roughly doubling how inevitable it seemed, and that people are unaware this is happening to them. Outcome bias compounds it: a pre-registered replication of Baron and Hershey’s study found people rate decisions more favourably when they happened to succeed, even when they had explicitly said outcomes should not affect the evaluation. The defence is stating your expectations before you look at results.
The evidence is unusually good for a management practice. A meta-analysis of 46 samples covering 2,136 people found properly conducted debriefs improve effectiveness over a control group by around 25%, with similar results for teams and individuals across simulated and real settings. A later meta-analysis of 83 studies found the benefit is largest for complex, ambiguous tasks that offer no natural feedback, and noted those tasks are common in industries that do not traditionally run after-action reviews. Marketing campaigns fit that description precisely.
This guide draws on ten years of running marketing campaigns, combined with primary research on debriefing, hindsight bias and psychological safety, and official after-action review guidance from US Government sources. Statistics were traced to the original study or report rather than to secondary coverage. A widely-quoted figure for the share of marketing teams that never run a post-campaign review was deliberately excluded because it could not be traced to any real survey; the PMI knowledge-transfer figure is used instead, with its limitations stated.