Explore our AI courses, practical training for non-technical teamsExplore courses Explore AI courses
Finance & OpsAI ROIMeasurement

Your AI Might Be Working. Right Now You Can't Prove It.

Revenue is up, the team is using the tool, and you still can't answer the only question that matters, which is whether one caused the other.

TLDR: Most AI business cases are assembled after the decision has already been made, from numbers that moved for six different reasons. That isn’t dishonesty, it’s what happens when nobody captured a before-number. Proving AI worked needs three unglamorous things, and all of them have to happen before you start: a baseline, something held back to compare against, and a written definition of what success would look like. None of it requires a data team. All of it requires deciding in advance, which is the part almost nobody does.
72%Of business leaders at large US enterprises say they track structured, business-linked ROI metrics for generative AI (Wharton Human-AI Research with GBK Collective, around 800 respondents at US firms with 1,000+ employees, fielded June to July 2025)
SGD 1bnEconomic value DBS attributes to data analytics and AI in FY2025, across more than 2,000 models and 430+ use cases, published in its own annual report. The unusual part is publishing a number at all
3Things that have to exist before you start, not after: a before-number, something held back to compare against, and a written definition of what success would look like

Share this article

The Short Version

You can tell whether people are using an AI tool. You almost certainly cannot tell whether it caused the change you’re pointing at, and those are different claims. The gap shows up clearly in the research: in one 2025 survey of large US enterprises, 72% of leaders said they track structured business-linked ROI metrics, while a widely circulated report from the same year claimed 95% of organisations were getting nothing back. Both can’t be right. This piece is about the difference between counting a change and proving you caused it: how to capture a baseline in an afternoon, how to run a comparison group with one team and a spreadsheet, why writing down your success definition in advance is the single highest-value thing here, and how to handle the argument that it just needs more time.

The question you can't answer: did the AI do that?

Picture the slide. Support ticket resolution time down 23% since the AI assistant went in, with a nice downward line. The head of support presents it, everyone nods, someone asks a question, and the question is: how do we know that’s the AI?

And the honest answer, in most companies, is that you don’t. In the same six months you also hired two people, changed the ticket categories, ran a product release that removed a whole class of complaints, and went through a quiet period in August. Any of those moves the number. The AI landed in the middle of all of it.

Nobody is lying on that slide. The number is real. It just isn’t evidence of what it’s being used to argue, and everyone in the room half knows it, which is why the conversation moves on quickly.

There’s a strange pair of findings that captures this. Wharton’s 2025 survey of around 800 decision-makers at large US enterprises found 72% saying they track structured, business-linked ROI metrics for generative AI, things like profitability and workforce productivity rather than adoption counts. [1] A widely circulated report from the same year claimed 95% of organisations were getting zero return.

Those two can’t both be describing the same reality. Either most of that first group is measuring something that isn’t causal attribution, or one of the samples is badly unrepresentative. I lean towards the first, and I’d include a fair amount of my own past reporting in that.

Worth separating two claims that get treated as one, because everything below depends on it. People are using this is a usage claim, and it’s easy to support. This caused that result is a causal claim, and it needs a completely different kind of evidence that has to be set up in advance. Most AI business cases quietly swap the first for the second.

If what you need is the categories of value and how to build one, we’ve covered that separately in the guide to measuring AI ROI. Track it there. This is about proving it.

Why your usage dashboard settles nothing

Every AI tool ships with a dashboard now. Active users, prompts per week, most-used features, adoption by department. It’s genuinely useful for one thing, which is knowing whether the rollout is alive.

It tells you nothing about value, and the reason is worth being precise about. Usage is an input. You bought it hoping it would produce an output. Reporting the input as though it were the output is how a rollout stays funded for two years without anyone establishing whether it did anything.

Here’s the widely repeated figure on the other side, and it needs handling carefully because it’s become shorthand for something it doesn’t actually say. A July 2025 report from a project at MIT stated that “95% of organizations are getting zero return” on enterprise generative AI. [2] It went everywhere.

Read the document itself and it’s narrower than the headline. The 95% refers specifically to custom and task-specific enterprise GenAI tools reaching sustained production impact, not to all AI pilots, and general-purpose tools in the same work had a 40% implementation rate. The authors label their own document “preliminary findings,” say the figures are “directionally accurate based on individual interviews rather than official company reporting,” and note that their six-month observation window “may be insufficient” and could be understating success. The evidence base is 52 interviews, 153 questionnaire responses collected at industry conferences, and a review of public disclosures. It isn’t peer reviewed, and the project’s own conclusion happens to be that the fix is the kind of architecture that project builds.

None of that makes it worthless. It makes it a directional signal from a convenience sample, which is a different thing from the settled fact it gets quoted as. I’m including it partly because the way it travelled is the exact failure this article is about: a number with heavy caveats attached at source, repeated until the caveats fall off.

The Attribution Rule

A number nobody can reproduce is an anecdote in a spreadsheet.

What a baseline actually is, and why it has to exist first

A baseline is the boring one and it’s the one that decides everything. It’s the number as it stood before you changed anything, captured in a way you could show someone.

The reason it has to happen first is that a baseline reconstructed afterwards is not a baseline. Once you know the result, you’ll choose the comparison window that makes the story coherent, and you won’t notice yourself doing it. Six months ago was a bad quarter. Last year had that one enormous client. Everyone does this. It’s not fraud, it’s memory doing what memory does.

It takes an afternoon. This is what one looks like filled in, for a marketing team putting AI into their monthly campaign reporting:

A baseline capture card, filled in

FieldMonthly campaign report
What we’re measuringHours from data available to report signed off, per monthly cycle
The number, before11.5 hours average across the last six cycles (Feb to Jul). Range 9 to 16
How we got itTwo people reconstructed it from calendar blocks and the file version history. Rough, and written down as rough
What else moves this numberCampaign count that month, whether the analyst is on leave, whether the CFO asks for a rebuild
Captured on4 August 2026, three weeks before the tool went in. Signed off by the head of marketing ops

A worked example. The fourth row is the one people skip, and it is the one that stops the final number from being embarrassing.

Two things about that card. The third row admits the measurement is crude, which is what makes it credible rather than what undermines it. A stated-as-rough number that was captured before you started beats a precise number assembled afterwards, every time, because the second one had a result to live up to.

And the fourth row is the whole game. Writing down what else moves your number, before you know the answer, is how you avoid claiming credit in six months for something a seasonal dip did.

The Baseline Rule

If you didn’t write the before-number down before you started, you don’t have one.

If you’re reading this and the tool went in four months ago, you’re not stuck. You can start a baseline today for the next thing, and for the current one you can be straight about what you have: a plausible improvement you can’t attribute. Saying that out loud costs less credibility than people expect, and considerably less than being unpicked later.

Comparison without a research budget

The reason a proper experiment works is that it watches two nearly identical groups, changes one thing for one of them, and compares. You don’t have a budget for a real experiment. You can still hold something back.

Concretely, that means one of these:

  • One team, not all of them. Two similar sales pods, the tool goes to one for a quarter. The other keeps working as it was.
  • One segment. AI-assisted follow-up for inbound leads from paid search only, not from organic, when the two normally behave similarly.
  • One campaign type. The tool writes the subject lines for the newsletter and not for the product announcements, and you compare against those same two streams last quarter.
  • Staggered rollout. Department by department over four months, which gives you a natural comparison at every step and is usually easier to get agreed than withholding anything.

This is deliberately cruder than a real experiment. The groups aren’t randomised, they aren’t identical, and the sample is small. It’s still enormously better than nothing, because it gives you something to point at when someone asks whether the number would have moved anyway.

Worth saying plainly what the staggered version costs you: teams talk to each other, so the later groups aren’t clean. That’s fine. It weakens the comparison, it doesn’t void it, and you should say so in the write-up rather than hoping nobody asks.

There’s a live example of why this discipline matters, and it’s a slightly awkward one. DBS Bank publishes an AI value figure in its statutory annual report: approximately SGD 1 billion of economic value from data analytics and AI in FY2025, across more than 2,000 models and 430 use cases, up from around SGD 750 million the previous year. [3] Publishing a number like that, audited, two years running, is rare enough to be the interesting part.

What’s also true is that the line most often repeated about DBS, that it validates this by comparing outcomes against matched control groups, is not something DBS itself says anywhere in that document. It comes from an analyst blog characterising an interview, and was then syndicated word for word by another outlet, which makes one source look like two. I went looking for it because I wanted to use it, and this article would have been easier to write if it had been there.

So the defensible version is narrower and still useful: a bank that puts an AI value figure into its published accounts has to stand behind it in a way that an internal slide never does. Publication is itself a form of discipline.

The Comparison Rule

Hold something back, or you have a story with no counterfactual.

Writing down what success looks like before you build

This is the cheapest item on the list and the one that changes the most. It takes about fifteen minutes and it has to happen before the tool arrives.

The reason it works is uncomfortable. If you haven’t defined success in advance, you will define it afterwards using whatever moved. That’s not cynicism about your colleagues, it’s how anyone reads a mixed result. Three numbers went up, two went down, one didn’t move; the write-up features the three that went up.

Here’s the whole artifact. Copy it, fill it in, and send it to someone before you start so it’s timestamped in an inbox you don’t control:

The note you write before you start

LineFilled in
We are putting AI intoThe monthly campaign report, for the marketing ops team of four
We will call this a success ifMedian cycle time drops below 8 hours by the December cycle, and the head of marketing still signs it off without a rework round
We will call it a failure ifCycle time is above 10 hours in December, or rework rounds go up at all
We will stop early ifTwo consecutive cycles contain a factual error that reached the leadership pack
Things that could produce this result insteadFewer campaigns in Q4, the new template we introduced in September, one analyst getting faster with practice
Decision date15 January 2027. Not “when we have enough data”

A worked example of a pre-commitment note. The failure line and the decision date are the two that do the work.

The failure line is the one people resist writing, and it’s the reason the whole thing functions. A success criterion with no matching failure criterion isn’t a test, because there’s no result that would count as a no.

The decision date matters for a related reason. “We’ll review it when we’ve got enough data” means the review happens when someone remembers, which is usually when the budget conversation forces it, by which point the answer is already politically determined.

The "give it time" argument, taken seriously

Someone in the room always says it, usually the person who championed the tool. It needs longer. These things take time to bed in.

They have a genuine point, and it’s stronger than the eye-rolling it usually gets.

US Census Bureau researchers, including Erik Brynjolfsson and Kristina McElheran, studied AI adoption across manufacturing firms using government data and found what they describe as causal evidence of J-curve-shaped returns, where short-term performance losses come before longer-term gains. [4] Firms adopting AI saw productivity and profitability fall in the short run while they absorbed the adjustment costs. Firms that had adopted by 2017 showed stronger growth in revenue, labour productivity and employment over the following four years.

So the dip is real and it’s measurable, in official data, and the authors went to some trouble to establish causation rather than correlation. Anyone dismissing “give it time” as an excuse is arguing against reasonably good evidence.

The same paper contains the detail that turns this from a defence into a warning. It found that older firms in particular struggled to maintain basic production management practices during the transition, specifically monitoring their KPIs and production targets, and that this collapse in structured management accounted for around a third of their productivity loss.

Read that again in the context of this article. A meaningful chunk of the dip came from companies letting go of their own measurement discipline while they adopted the new thing.

Which gives you the actual position. “Give it time” is a legitimate argument and it becomes an excuse at one specific moment: when nobody has written down what recovery would look like or when it should arrive. Those are different sentences.

  • “It needs longer” is an excuse.
  • “We expected a dip through Q3 and a return to baseline by the December cycle, and we’re on track” is a plan.
  • “We expected recovery by December, it hasn’t come, so here’s what we’re changing” is a functioning programme.
The Patience Rule

Give it time is a plan only if you wrote down what recovery looks like, and when.

Building a number that survives the first hard question

Everything above is in service of one moment. You’re in a room, you’ve said the number out loud, and a sceptical person who is good at their job starts asking.

Four questions do most of the damage, and you can test your own number against them before anyone else does:

The four questions, and what a passing answer sounds like

The questionA passing answerA failing answer
“Compared to what?”“Six cycles before we started, captured in August, plus the two teams that didn’t get it”“Compared to before”
“What else changed?”“Three things, here they are, and here’s why I don’t think they explain it”“Nothing significant”
“Would you have called this a win in advance?”“Here’s the note I sent in August saying under 8 hours was the bar”“Well, it’s clearly better”
“Could someone else get this number from the same data?”“Yes, it’s in this sheet, here’s the method”“I pulled it together from a few places”

A self-check before presenting. Failing any single one is survivable if you say so first. Failing three is where credibility goes.

Reproducibility, that last row, is the one that gets least attention and does the most long-term damage when it’s missing. If the only person who can produce the number is you, and you produced it by hand from four sources, then it isn’t a measurement, it’s a claim. Everyone senior has been burned by one of those before and they’re reading yours through that.

The measurement work itself is a reasonable thing to hand partly to AI, as long as the split is explicit:

Handing the analysis over, without handing over the claim

What AI doesWhat you still ownHow it gets checked
Pulls the cycle times into one sheet, computes medians, drafts the write-up in your standard formatWhether the comparison is fair, which confounders are live, and the actual causal claim. Your name goes on thatRecompute one month by hand against source before presenting. Every time, not when something looks off

The split for the measurement work. The arithmetic is delegable. The attribution is not.

If your current AI programme has none of this, the honest read is that you’re in the position most people are in, and it’s recoverable. The executives who say they’ve been disappointed by AI are largely describing this exact problem rather than a technology failure. And when the number does have to go in front of a board, how you frame it matters nearly as much as how you built it.

The move for this week is small and it isn’t the current project. Pick the next thing you’re about to roll out, spend fifteen minutes writing the six lines from section five, and email it to someone. That’s the entire intervention. Everything else here is easier once that note exists.

Hina Mian
Hina Mian, Co-Founder of Future Factors AI

Hina is a marketing strategist with over a decade of hands-on campaign experience across B2B and consumer brands. She writes about using AI to run leaner, sharper marketing without losing the human touch. Future Factors helps professionals and teams build practical AI capability through role-based training, workflow design, and hands-on adoption.

More about Hina →

Frequently Asked Questions

Why can't I just use my AI tool's usage dashboard to prove ROI?

Because usage is an input and value is an output, and a dashboard only measures the first. Knowing that 240 people ran 9,000 prompts last month tells you the rollout is alive, which is genuinely worth knowing, but it says nothing about whether any work got better. The trap is that usage numbers are easy to pull and always available, so they end up standing in for the harder measurement nobody set up. A useful test: if usage doubled and the business outcome didn’t move at all, would your dashboard show you that? If the answer is no, it isn’t measuring value. Usage is a reasonable first signal, and an argument built only on it will not survive the first person who asks what changed as a result.

What is a baseline, and what do I do if I never captured one?

A baseline is the number as it stood before you changed anything, recorded in a way you could show someone, alongside a note of what else moves that number. It has to be captured before the change, because a baseline reconstructed afterwards is unconsciously selected to fit the result you already know. If you never captured one, don’t fabricate a comparison window. You have two honest options. Start a proper baseline today for the next rollout, which is where the value is anyway. And for the current one, say plainly that you have a plausible improvement you can’t attribute, and name the other things that changed in the same period. That costs far less credibility than a number that gets unpicked in the meeting.

How do I run a control group without a data team or a research budget?

Hold something back. One of two similar teams, one customer segment, one campaign type, or a rollout staggered over several months so each wave is compared against the ones that haven’t started yet. None of these are randomised and the groups won’t be identical, so this is deliberately cruder than a real experiment. It still gives you something to answer the question ‘would that have happened anyway’ with, which is the question that ends most AI business cases. The main thing to be honest about is contamination: people talk to each other, so a held-back team often picks things up informally. That weakens the comparison rather than voiding it, and stating it yourself is much better than having it pointed out.

How long should I wait before judging whether an AI rollout worked?

Decide the date in advance and write it down, because the length matters less than the pre-commitment. There is real evidence for a dip before recovery: US Census Bureau research found causal evidence of J-curve returns from AI adoption in manufacturing, with short-term productivity and profitability losses preceding longer-term gains. So patience is defensible. What turns it into an excuse is the absence of a stated expectation. The same research found that a large part of the short-run loss at older firms came from those firms abandoning their own KPI and target monitoring during the transition, which is the opposite of what you want. A specific date with a specific expected number, agreed before launch, is the whole difference between waiting and drifting.

What is the minimum I need to defend an AI number to my boss or board?

Four things, and none of them require a data team. A before-number captured before you started, with a note of how you got it and how rough it is. Something held back to compare against, even if it’s just one team or one segment. A written success definition, sent to someone else and timestamped, that includes what would have counted as failure. And a method someone else could follow to reproduce your figure from the same data. If you have all four, you can survive a genuinely sceptical reading. If you have none, you have a plausible story, and the difference between those two only becomes visible at the exact moment you most need it not to.

About This Article

The four figures in this article come from their original publishers, all checked on 29 August 2026, and one of them required a correction worth stating openly. The DBS figures are from DBS Group’s own FY2025 annual report chapters, which are freely readable; the consolidated PDF is bot-blocked and was not used. The widely repeated claim that DBS validates its AI value figure against matched control groups is not in DBS’s report. It originates in a Forrester analyst blog characterising an interview, and was syndicated verbatim elsewhere, which makes a single source look like corroboration. We went looking for it in order to use it, could not verify it at source, and have said so in the body rather than quietly citing the analyst version. The MIT-affiliated 95% figure is quoted from the report document itself along with the authors’ own stated limitations; note that no copy is hosted on an mit.edu domain, which is a real provenance weakness. The Census working paper carries the standard notice that it has not undergone Census Bureau review. The Wharton figure is self-reported, from large US enterprises only, and does not tell you those organisations captured a baseline. Three further statistics were rejected during research for tracing back to content-marketing sites citing unlocatable studies, and no McKinsey figure appears here because no McKinsey primary source was obtained. The baseline card, the pre-commitment note and the four-question test are Future Factors’ own.

Sources

  1. Wharton Human-AI Research (WHAIR), The Wharton School, University of Pennsylvania, with GBK Collective. Accountable Acceleration: Gen AI Fast-Tracks Into the Enterprise. Published October 2025. 15-minute online tracking survey of approximately 800 senior decision makers at US-based enterprises with 1,000+ employees and over $50m revenue, fielded 26 June to 11 July 2025. Self-reported. Read 29 August 2026. https://ai.wharton.upenn.edu/wp-content/uploads/2025/10/2025-Wharton-GBK-AI-Adoption-Report_Full-Report.pdf
  2. Challapally A, Pease C, Raskar R, Chari P. The GenAI Divide: State of AI in Business 2025. Project NANDA, MIT. Dated July 2025, research period January to June 2025. Multi-method design: review of over 300 publicly disclosed AI initiatives, structured interviews with 52 organisations, and 153 survey responses collected at four industry conferences. The document labels itself preliminary findings and states its figures are directionally accurate based on interviews rather than official company reporting. Not peer reviewed. No copy is hosted on an mit.edu domain; read from a third-party mirror on 29 August 2026. https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf
  3. DBS Group Holdings. Annual Report 2025, Letter from Chairman and CEO, and CIO statement. Figures of approximately SGD 1 billion in economic value from data analytics and AI/ML, over 2,000 models and more than 430 use cases, are stated in both chapters. The FY2024 comparison figure of over SGD 750 million is from the equivalent FY2024 chapter. Read 29 August 2026. DBS’s own report does not describe a control-group methodology; see the About This Guide note. https://www.dbs.com/annualreports/2025/letter-from-chairman-ceo.html
  4. McElheran K, Yang M-J, Kroff Z, Brynjolfsson E. The Rise of Industrial AI in America: Microfoundations of the Productivity J-curve(s). US Census Bureau, Center for Economic Studies Working Paper 25-27, April 2025. Built on the 2021 Management and Organizational Practices Survey (68% response rate) and a panel of approximately 55,000 manufacturing firms from the 2018 Annual Business Survey and Economic Census, linked to Census and IRS records. Carries the standard notice that it has not undergone the review accorded Census Bureau publications. Read 29 August 2026. https://www2.census.gov/library/working-papers/2025/adrm/ces/CES-WP-25-27.pdf

Psst, Hey You!

(Yeah, You!)

Want helpful AI tips flying Into your inbox?

Weekly tips. Real examples. Practical help for busy professionals.

We care about your data, check out our privacy policy.