Revenue is up, the team is using the tool, and you still can't answer the only question that matters, which is whether one caused the other.
You can tell whether people are using an AI tool. You almost certainly cannot tell whether it caused the change you’re pointing at, and those are different claims. The gap shows up clearly in the research: in one 2025 survey of large US enterprises, 72% of leaders said they track structured business-linked ROI metrics, while a widely circulated report from the same year claimed 95% of organisations were getting nothing back. Both can’t be right. This piece is about the difference between counting a change and proving you caused it: how to capture a baseline in an afternoon, how to run a comparison group with one team and a spreadsheet, why writing down your success definition in advance is the single highest-value thing here, and how to handle the argument that it just needs more time.
Picture the slide. Support ticket resolution time down 23% since the AI assistant went in, with a nice downward line. The head of support presents it, everyone nods, someone asks a question, and the question is: how do we know that’s the AI?
And the honest answer, in most companies, is that you don’t. In the same six months you also hired two people, changed the ticket categories, ran a product release that removed a whole class of complaints, and went through a quiet period in August. Any of those moves the number. The AI landed in the middle of all of it.
Nobody is lying on that slide. The number is real. It just isn’t evidence of what it’s being used to argue, and everyone in the room half knows it, which is why the conversation moves on quickly.
There’s a strange pair of findings that captures this. Wharton’s 2025 survey of around 800 decision-makers at large US enterprises found 72% saying they track structured, business-linked ROI metrics for generative AI, things like profitability and workforce productivity rather than adoption counts. [1] A widely circulated report from the same year claimed 95% of organisations were getting zero return.
Those two can’t both be describing the same reality. Either most of that first group is measuring something that isn’t causal attribution, or one of the samples is badly unrepresentative. I lean towards the first, and I’d include a fair amount of my own past reporting in that.
Worth separating two claims that get treated as one, because everything below depends on it. People are using this is a usage claim, and it’s easy to support. This caused that result is a causal claim, and it needs a completely different kind of evidence that has to be set up in advance. Most AI business cases quietly swap the first for the second.
If what you need is the categories of value and how to build one, we’ve covered that separately in the guide to measuring AI ROI. Track it there. This is about proving it.
Every AI tool ships with a dashboard now. Active users, prompts per week, most-used features, adoption by department. It’s genuinely useful for one thing, which is knowing whether the rollout is alive.
It tells you nothing about value, and the reason is worth being precise about. Usage is an input. You bought it hoping it would produce an output. Reporting the input as though it were the output is how a rollout stays funded for two years without anyone establishing whether it did anything.
Here’s the widely repeated figure on the other side, and it needs handling carefully because it’s become shorthand for something it doesn’t actually say. A July 2025 report from a project at MIT stated that “95% of organizations are getting zero return” on enterprise generative AI. [2] It went everywhere.
Read the document itself and it’s narrower than the headline. The 95% refers specifically to custom and task-specific enterprise GenAI tools reaching sustained production impact, not to all AI pilots, and general-purpose tools in the same work had a 40% implementation rate. The authors label their own document “preliminary findings,” say the figures are “directionally accurate based on individual interviews rather than official company reporting,” and note that their six-month observation window “may be insufficient” and could be understating success. The evidence base is 52 interviews, 153 questionnaire responses collected at industry conferences, and a review of public disclosures. It isn’t peer reviewed, and the project’s own conclusion happens to be that the fix is the kind of architecture that project builds.
None of that makes it worthless. It makes it a directional signal from a convenience sample, which is a different thing from the settled fact it gets quoted as. I’m including it partly because the way it travelled is the exact failure this article is about: a number with heavy caveats attached at source, repeated until the caveats fall off.
A number nobody can reproduce is an anecdote in a spreadsheet.
A baseline is the boring one and it’s the one that decides everything. It’s the number as it stood before you changed anything, captured in a way you could show someone.
The reason it has to happen first is that a baseline reconstructed afterwards is not a baseline. Once you know the result, you’ll choose the comparison window that makes the story coherent, and you won’t notice yourself doing it. Six months ago was a bad quarter. Last year had that one enormous client. Everyone does this. It’s not fraud, it’s memory doing what memory does.
It takes an afternoon. This is what one looks like filled in, for a marketing team putting AI into their monthly campaign reporting:
| Field | Monthly campaign report |
|---|---|
| What we’re measuring | Hours from data available to report signed off, per monthly cycle |
| The number, before | 11.5 hours average across the last six cycles (Feb to Jul). Range 9 to 16 |
| How we got it | Two people reconstructed it from calendar blocks and the file version history. Rough, and written down as rough |
| What else moves this number | Campaign count that month, whether the analyst is on leave, whether the CFO asks for a rebuild |
| Captured on | 4 August 2026, three weeks before the tool went in. Signed off by the head of marketing ops |
A worked example. The fourth row is the one people skip, and it is the one that stops the final number from being embarrassing.
Two things about that card. The third row admits the measurement is crude, which is what makes it credible rather than what undermines it. A stated-as-rough number that was captured before you started beats a precise number assembled afterwards, every time, because the second one had a result to live up to.
And the fourth row is the whole game. Writing down what else moves your number, before you know the answer, is how you avoid claiming credit in six months for something a seasonal dip did.
If you didn’t write the before-number down before you started, you don’t have one.
If you’re reading this and the tool went in four months ago, you’re not stuck. You can start a baseline today for the next thing, and for the current one you can be straight about what you have: a plausible improvement you can’t attribute. Saying that out loud costs less credibility than people expect, and considerably less than being unpicked later.
The reason a proper experiment works is that it watches two nearly identical groups, changes one thing for one of them, and compares. You don’t have a budget for a real experiment. You can still hold something back.
Concretely, that means one of these:
This is deliberately cruder than a real experiment. The groups aren’t randomised, they aren’t identical, and the sample is small. It’s still enormously better than nothing, because it gives you something to point at when someone asks whether the number would have moved anyway.
Worth saying plainly what the staggered version costs you: teams talk to each other, so the later groups aren’t clean. That’s fine. It weakens the comparison, it doesn’t void it, and you should say so in the write-up rather than hoping nobody asks.
There’s a live example of why this discipline matters, and it’s a slightly awkward one. DBS Bank publishes an AI value figure in its statutory annual report: approximately SGD 1 billion of economic value from data analytics and AI in FY2025, across more than 2,000 models and 430 use cases, up from around SGD 750 million the previous year. [3] Publishing a number like that, audited, two years running, is rare enough to be the interesting part.
What’s also true is that the line most often repeated about DBS, that it validates this by comparing outcomes against matched control groups, is not something DBS itself says anywhere in that document. It comes from an analyst blog characterising an interview, and was then syndicated word for word by another outlet, which makes one source look like two. I went looking for it because I wanted to use it, and this article would have been easier to write if it had been there.
So the defensible version is narrower and still useful: a bank that puts an AI value figure into its published accounts has to stand behind it in a way that an internal slide never does. Publication is itself a form of discipline.
Hold something back, or you have a story with no counterfactual.
This is the cheapest item on the list and the one that changes the most. It takes about fifteen minutes and it has to happen before the tool arrives.
The reason it works is uncomfortable. If you haven’t defined success in advance, you will define it afterwards using whatever moved. That’s not cynicism about your colleagues, it’s how anyone reads a mixed result. Three numbers went up, two went down, one didn’t move; the write-up features the three that went up.
Here’s the whole artifact. Copy it, fill it in, and send it to someone before you start so it’s timestamped in an inbox you don’t control:
| Line | Filled in |
|---|---|
| We are putting AI into | The monthly campaign report, for the marketing ops team of four |
| We will call this a success if | Median cycle time drops below 8 hours by the December cycle, and the head of marketing still signs it off without a rework round |
| We will call it a failure if | Cycle time is above 10 hours in December, or rework rounds go up at all |
| We will stop early if | Two consecutive cycles contain a factual error that reached the leadership pack |
| Things that could produce this result instead | Fewer campaigns in Q4, the new template we introduced in September, one analyst getting faster with practice |
| Decision date | 15 January 2027. Not “when we have enough data” |
A worked example of a pre-commitment note. The failure line and the decision date are the two that do the work.
The failure line is the one people resist writing, and it’s the reason the whole thing functions. A success criterion with no matching failure criterion isn’t a test, because there’s no result that would count as a no.
The decision date matters for a related reason. “We’ll review it when we’ve got enough data” means the review happens when someone remembers, which is usually when the budget conversation forces it, by which point the answer is already politically determined.
Someone in the room always says it, usually the person who championed the tool. It needs longer. These things take time to bed in.
They have a genuine point, and it’s stronger than the eye-rolling it usually gets.
US Census Bureau researchers, including Erik Brynjolfsson and Kristina McElheran, studied AI adoption across manufacturing firms using government data and found what they describe as causal evidence of J-curve-shaped returns, where short-term performance losses come before longer-term gains. [4] Firms adopting AI saw productivity and profitability fall in the short run while they absorbed the adjustment costs. Firms that had adopted by 2017 showed stronger growth in revenue, labour productivity and employment over the following four years.
So the dip is real and it’s measurable, in official data, and the authors went to some trouble to establish causation rather than correlation. Anyone dismissing “give it time” as an excuse is arguing against reasonably good evidence.
The same paper contains the detail that turns this from a defence into a warning. It found that older firms in particular struggled to maintain basic production management practices during the transition, specifically monitoring their KPIs and production targets, and that this collapse in structured management accounted for around a third of their productivity loss.
Read that again in the context of this article. A meaningful chunk of the dip came from companies letting go of their own measurement discipline while they adopted the new thing.
Which gives you the actual position. “Give it time” is a legitimate argument and it becomes an excuse at one specific moment: when nobody has written down what recovery would look like or when it should arrive. Those are different sentences.
Give it time is a plan only if you wrote down what recovery looks like, and when.
Everything above is in service of one moment. You’re in a room, you’ve said the number out loud, and a sceptical person who is good at their job starts asking.
Four questions do most of the damage, and you can test your own number against them before anyone else does:
| The question | A passing answer | A failing answer |
|---|---|---|
| “Compared to what?” | “Six cycles before we started, captured in August, plus the two teams that didn’t get it” | “Compared to before” |
| “What else changed?” | “Three things, here they are, and here’s why I don’t think they explain it” | “Nothing significant” |
| “Would you have called this a win in advance?” | “Here’s the note I sent in August saying under 8 hours was the bar” | “Well, it’s clearly better” |
| “Could someone else get this number from the same data?” | “Yes, it’s in this sheet, here’s the method” | “I pulled it together from a few places” |
A self-check before presenting. Failing any single one is survivable if you say so first. Failing three is where credibility goes.
Reproducibility, that last row, is the one that gets least attention and does the most long-term damage when it’s missing. If the only person who can produce the number is you, and you produced it by hand from four sources, then it isn’t a measurement, it’s a claim. Everyone senior has been burned by one of those before and they’re reading yours through that.
The measurement work itself is a reasonable thing to hand partly to AI, as long as the split is explicit:
| What AI does | What you still own | How it gets checked |
|---|---|---|
| Pulls the cycle times into one sheet, computes medians, drafts the write-up in your standard format | Whether the comparison is fair, which confounders are live, and the actual causal claim. Your name goes on that | Recompute one month by hand against source before presenting. Every time, not when something looks off |
The split for the measurement work. The arithmetic is delegable. The attribution is not.
If your current AI programme has none of this, the honest read is that you’re in the position most people are in, and it’s recoverable. The executives who say they’ve been disappointed by AI are largely describing this exact problem rather than a technology failure. And when the number does have to go in front of a board, how you frame it matters nearly as much as how you built it.
The move for this week is small and it isn’t the current project. Pick the next thing you’re about to roll out, spend fifteen minutes writing the six lines from section five, and email it to someone. That’s the entire intervention. Everything else here is easier once that note exists.
Because usage is an input and value is an output, and a dashboard only measures the first. Knowing that 240 people ran 9,000 prompts last month tells you the rollout is alive, which is genuinely worth knowing, but it says nothing about whether any work got better. The trap is that usage numbers are easy to pull and always available, so they end up standing in for the harder measurement nobody set up. A useful test: if usage doubled and the business outcome didn’t move at all, would your dashboard show you that? If the answer is no, it isn’t measuring value. Usage is a reasonable first signal, and an argument built only on it will not survive the first person who asks what changed as a result.
A baseline is the number as it stood before you changed anything, recorded in a way you could show someone, alongside a note of what else moves that number. It has to be captured before the change, because a baseline reconstructed afterwards is unconsciously selected to fit the result you already know. If you never captured one, don’t fabricate a comparison window. You have two honest options. Start a proper baseline today for the next rollout, which is where the value is anyway. And for the current one, say plainly that you have a plausible improvement you can’t attribute, and name the other things that changed in the same period. That costs far less credibility than a number that gets unpicked in the meeting.
Hold something back. One of two similar teams, one customer segment, one campaign type, or a rollout staggered over several months so each wave is compared against the ones that haven’t started yet. None of these are randomised and the groups won’t be identical, so this is deliberately cruder than a real experiment. It still gives you something to answer the question ‘would that have happened anyway’ with, which is the question that ends most AI business cases. The main thing to be honest about is contamination: people talk to each other, so a held-back team often picks things up informally. That weakens the comparison rather than voiding it, and stating it yourself is much better than having it pointed out.
Decide the date in advance and write it down, because the length matters less than the pre-commitment. There is real evidence for a dip before recovery: US Census Bureau research found causal evidence of J-curve returns from AI adoption in manufacturing, with short-term productivity and profitability losses preceding longer-term gains. So patience is defensible. What turns it into an excuse is the absence of a stated expectation. The same research found that a large part of the short-run loss at older firms came from those firms abandoning their own KPI and target monitoring during the transition, which is the opposite of what you want. A specific date with a specific expected number, agreed before launch, is the whole difference between waiting and drifting.
Four things, and none of them require a data team. A before-number captured before you started, with a note of how you got it and how rough it is. Something held back to compare against, even if it’s just one team or one segment. A written success definition, sent to someone else and timestamped, that includes what would have counted as failure. And a method someone else could follow to reproduce your figure from the same data. If you have all four, you can survive a genuinely sceptical reading. If you have none, you have a plausible story, and the difference between those two only becomes visible at the exact moment you most need it not to.
The four figures in this article come from their original publishers, all checked on 29 August 2026, and one of them required a correction worth stating openly. The DBS figures are from DBS Group’s own FY2025 annual report chapters, which are freely readable; the consolidated PDF is bot-blocked and was not used. The widely repeated claim that DBS validates its AI value figure against matched control groups is not in DBS’s report. It originates in a Forrester analyst blog characterising an interview, and was syndicated verbatim elsewhere, which makes a single source look like corroboration. We went looking for it in order to use it, could not verify it at source, and have said so in the body rather than quietly citing the analyst version. The MIT-affiliated 95% figure is quoted from the report document itself along with the authors’ own stated limitations; note that no copy is hosted on an mit.edu domain, which is a real provenance weakness. The Census working paper carries the standard notice that it has not undergone Census Bureau review. The Wharton figure is self-reported, from large US enterprises only, and does not tell you those organisations captured a baseline. Three further statistics were rejected during research for tracing back to content-marketing sites citing unlocatable studies, and no McKinsey figure appears here because no McKinsey primary source was obtained. The baseline card, the pre-commitment note and the four-question test are Future Factors’ own.