Most AI pilots don't fail. They just never get a decision made about what happens after they work.
A pilot that never scales usually wasn’t designed to. The fix isn’t a better pilot, it’s deciding four things before the pilot launches: what decision it’s meant to inform, who owns that decision, the specific threshold that triggers scale or kill, and the budget and mandate waiting if it clears that bar. This piece covers the specific, non-obvious reasons pilots stall, a simple framework for setting that threshold up front, why the right decision-maker changes depending on whether the pilot sits inside one function or crosses several, and a six-step path from a cleared pilot to an actual company-wide rollout.
Say a VP of customer experience greenlights a chatbot pilot on the smallest support queue in the building, the one nobody would notice if it broke. Eight weeks in, the numbers look decent: response time is down, a few agents are genuinely relieved, and deflection is ticking up on the low-stakes tickets. She puts together a tidy slide for the quarterly leadership review. The room nods. Someone says “great, let’s keep an eye on it.” The meeting moves to the next agenda item.
Eighteen months later, the same bot is still running on the same queue. Nobody killed it and nobody expanded it. It just sits there, a permanent pilot, and if you asked five people in that company why it never moved past its original scope, you’d get five different, vague answers about bandwidth and priorities.
That’s not really a failure story, and treating it as one misses the point. The pilot did what it was built to do. What never existed was a mechanism connecting “the pilot worked” to “now we do this everywhere.” Nobody had written down what winning looked like, who got to declare it, or what happened next if it did. So when it worked, there was no next step waiting, just a slide and a nod.
This is a strange thing to keep getting wrong, because piloting is now standard practice almost everywhere AI shows up. Most companies test something small before rolling it out further. But testing and scaling are two separate muscles, and plenty of organizations have gotten fairly good at the first one without building the second at all.
The size of that gap is bigger than most leaders assume. In a 2024 survey of 1,000 CxOs and senior executives across 59 countries, BCG found that only 26% of companies had built the capabilities needed to move beyond proof-of-concept work and generate tangible business value from AI.[1] Put the other way round, roughly three in four companies in that survey were stuck somewhere between “we tried it” and “it changed how we work.” Rather than a talent problem or a technology problem, it’s what happens when an organization gets good at starting pilots and never builds the habit of finishing them.
If nobody can name the number that ends the pilot, the pilot was never going to graduate.
The rest of this piece is about that missing mechanism: what specifically breaks it, what a pilot needs built in from day one, who should own the decision to scale, and a workable path from a good pilot to something that runs across the whole company.
Ask why a specific AI pilot didn’t go anywhere and you’ll usually get one of a handful of stock answers: the technology wasn’t ready, the team lost interest, there wasn’t enough buy-in. These explanations are comfortable because they don’t point at a decision anyone actually made. They also don’t hold up once you look closely at pilots that stalled.
Say a regional sales director runs a pilot letting her team draft proposals with AI. It goes well. Reps like it, and proposal turnaround drops from four days to one. She wants to roll it out to the other four regions. But she doesn’t own the budget for the tool license at that volume, and the decision about company-wide sales tooling sits with a VP two levels up who was never part of the pilot and has never seen the results. Nine months pass. The pilot region keeps using it quietly. Everyone else doesn’t.
That’s the first specific, mechanical reason pilots stall, not a vague culture problem: the person running the pilot had no mandate or budget to take it past pilot stage. Enthusiasm and good results don’t substitute for spending authority.
Three more specific reasons show up often enough to name on their own:
| What gets said | What’s actually going on |
|---|---|
| “There wasn’t enough buy-in” | No one with budget authority ever co-owned the pilot |
| “The tech wasn’t ready” | Success was never defined, so nothing could prove readiness either way |
| “We lost momentum” | The pilot proved a low-stakes case that didn’t need proving |
| “People didn’t adopt it” | No training or communication plan existed outside the original pilot team |
A register mapping the vague explanation people give to the specific mechanical cause behind it. Use it to check a stalled pilot against the real reason before assuming it was the technology.
The scale of this shows up in research too. MIT’s Project NANDA reviewed more than 300 public AI deployments, interviewed 52 organizations, and surveyed 153 senior leaders in 2025, and found that 95% of organizations piloting generative AI were seeing zero measurable return on the P&L.[2] The report’s own explanation lines up with the pattern above: the pilots that did work were narrowly scoped to a real workflow with a defined owner pushing them forward, not the ones using the most advanced tools.
A pilot owner without budget authority is running a demo, not a pilot.
Picture the same chatbot pilot from earlier, run differently from the start. Before it launches, the VP of customer experience sits down with the COO for twenty minutes. They agree on four things: the pilot exists to answer whether AI can handle a defined slice of the support queue without hurting satisfaction scores, the COO is the one who decides whether it scales, the pilot needs to cut average handling time by 20% with satisfaction holding steady, and if it clears that bar, the VP already has a name and a rough budget earmarked for licensing the wider team. That takes about the length of one meeting, before a single ticket has touched the bot.
Eight weeks later the numbers come in almost identical to the earlier version of this story. The difference is that this time there’s a pre-agreed threshold to check them against and a named person waiting to make the call, instead of a slide and a nod. The COO looks at the four things agreed at the start, sees the bar was cleared, and signs off on the budget that was already set aside. The whole conversation about whether to scale takes about ten minutes, because the hard part happened before the pilot started, not after.
Call this the Scale Threshold: four things a pilot needs settled before it launches, not after it succeeds.
| Field | Filled in for this pilot |
|---|---|
| The decision this pilot exists to inform | Whether AI can handle a defined slice of the support queue without hurting satisfaction |
| Who owns that decision | The COO, named before launch, not “leadership” in general |
| The threshold that triggers scale, kill, or rerun | 20% drop in average handling time with satisfaction scores flat or better |
| The budget and mandate assumed if it clears the bar | Licensing budget for the wider team, roughed out and set aside before results come in |
A filled-in example of the Scale Threshold framework. Copy the left column for any pilot and fill in the right column before launch, not after.
Notice what’s missing from that list: nothing about the technology itself. The Scale Threshold isn’t a technical gate, it’s an organizational one. Two pilots can use the identical tool and get identical results, and only one of them scales, because only one of them had someone on the other side of the threshold ready to act on the answer.
If you can’t fill in the four fields before launch, you’re not ready to pilot yet, you’re ready to experiment.
That distinction is worth sitting with. Experimenting is fine, and sometimes it’s exactly the right first move on something genuinely uncertain. It just isn’t a pilot in the sense that leadership usually means when they ask for one, and calling it a pilot anyway is part of how expectations get set that nothing can actually meet.
Ask “who decides whether this pilot scales” in most companies and you’ll get a pause, then something like “well, it would probably go to leadership.” That pause is the whole problem in miniature. If the decision doesn’t have a name attached before the pilot starts, it gets made by default, usually by whoever’s still paying attention when the pilot ends, which is often nobody in particular.
The right owner depends on what the pilot is actually testing, and the situation genuinely differs by role.
A VP of Operations is usually the right owner when a pilot touches a process that crosses functions, like order fulfillment or a shared services workflow, because scaling it will spend budget that isn’t any single department’s to commit alone. The risk with this role is distance: a VP of Ops wasn’t in the room for the day-to-day pilot, so they’re being asked to trust a summary rather than the underlying data, and summaries flatter the summarizer.
A department head, say a head of marketing or a head of customer service, is the right owner when the pilot stays entirely inside their own function. This is usually the fastest path to a scale decision, because there are fewer people to convince. The risk here runs the other way: a department head can say yes without ever scheduling the IT or security review that a company-wide rollout will eventually need anyway, which just moves the delay downstream instead of removing it.
An IT director is rarely the person who should make the business call on whether a pilot scales, but they usually hold a veto on security, data governance, and integration that can stop a “yes” cold months after the business decision was already made. Looping them in only after the department head or VP has decided is one of the most common ways a cleared pilot still stalls, because the approval that everyone assumed was a formality turns out not to be one.
| Role | When they’re the right decision owner | Where they usually get it wrong |
|---|---|---|
| VP of Operations | Pilot crosses functions, budget spans departments | Too far from the pilot to trust the data without a translator |
| Department head | Pilot stays inside one function | Says yes without scheduling the IT or security review it will need anyway |
| IT director | Rarely the yes/no owner, but holds veto power on security and integration | Brought in after the decision, so approval stalls anyway |
A register for matching a pilot’s scope to the role that should decide whether it scales, and the failure mode most common to each.
Scaling decisions need one name attached to them, not a committee.
The decision point everyone skips isn’t really about picking the right box on an org chart. It’s about picking someone, in writing, before the pilot starts, and getting them the pilot’s actual data when the time comes rather than a polished summary. Getting the right people some real voice earlier, rather than a courtesy update at the end, is most of what stakeholder engagement for AI adoption is actually about, and it’s cheaper to do before a pilot launches than to do as damage control after a scale decision has quietly stalled.
Say the Scale Threshold was defined, the owner was named, and the pilot cleared its bar. That’s still not the same as a company-wide rollout, and treating it as though it were is its own way to stall. A five-person pilot team’s habits don’t transfer to three hundred people on their own.
Six steps carry a cleared pilot into an actual rollout without losing what made the pilot work in the first place:
Step three deserves a beat of its own, because the instinct is usually to hand the rollout to whichever team is loudest about wanting it, and loud interest isn’t the same as workflow fit. The AI adoption checklist is a useful gate to run each candidate team through before it gets a rollout date, because it forces the same what-do-we-actually-need-in-place questions the original pilot should have answered for itself.
Measurement matters here as much as it did during the pilot, and the same trap catches rollouts that already avoided it once: counting logins. A better lens has three layers. Use asks whether people tried the thing at all. Persistence asks whether they’re still doing it a month later without being reminded. Impact asks whether the actual work is different now, and whether anyone would notice if it stopped. Rollouts that only track the first layer look successful for exactly as long as it takes the novelty to wear off.
| What AI does | What the human still owns | How it gets checked |
|---|---|---|
| Drafts the first version of training materials from the pilot team’s own notes | Deciding what actually needs teaching versus what people will pick up on their own | The new team’s local adoption owner reviews before it goes out |
| Surfaces which of the 30/60/90-day numbers moved, and by how much | Judging whether the move is real change or a one-off week | The named decision owner from the Scale Threshold, at each check-in |
| Answers routine how-do-I-do-this questions once the rollout is live | Deciding what “routine” means for this team, and who handles what isn’t | Reviewed after week one, adjusted if the wrong things are landing on the wrong desk |
A working split for rollout, not a measured finding. The middle column is where a rollout either holds or quietly reverts.
None of this requires an outside team to run it. Plenty of companies work through the checklist and the ownership table above on their own and get a rollout that holds. Where a facilitator tends to earn their place is earlier than that, in the first pass of building the four Scale Threshold fields for a pilot that hasn’t launched yet, when it’s genuinely hard to see your own project clearly enough to name its own kill criteria. That first-pass structuring is most of what we run in Future Factors’ corporate workshops, and companies tend to run their second and third pilot on their own once they’ve seen the first one done properly.
If you’re doing one thing this week, do this: take whichever AI pilot is currently running in your company without a defined end point, and get its owner, its threshold, and its budget commitment written down on one page before the next update meeting. That page is the difference between a pilot that quietly lives forever and one that actually goes somewhere.
Because the mechanism connecting “the pilot worked” to “now we scale it” usually doesn’t exist. Nobody agreed in advance on what success meant, who got to decide based on it, or what budget and mandate were waiting on the other side. Research backs up how common this is: BCG found only 26% of companies have built the capabilities to move past proof-of-concept work and generate real value from AI, meaning roughly three-quarters stay stuck somewhere between testing and transforming.
Long enough for the metric the Scale Threshold set for it to move past a normal week’s noise, which for most operational pilots lands somewhere in the six-to-twelve-week range rather than a fixed universal number. The point isn’t a set duration, it’s that the length gets decided before the pilot starts, tied to how long it actually takes the metric you care about to stabilize, instead of getting extended indefinitely because “we want more data” once results start coming in.
A successful pilot hit a target. A scalable one had a named decision owner and a pre-agreed budget waiting for that moment, plus a use case that was actually representative of harder, real work rather than the safest slice available. A pilot can succeed completely on its own terms and still have nowhere to go, because success and permission to expand were never treated as the same question.
Whoever owns the budget and mandate for whatever the pilot would become if it scaled, named before the pilot launches rather than after. For a pilot that stays inside one function that’s usually the department head; for anything crossing functions it’s typically a VP of Operations or equivalent; IT or security leadership rarely make the yes-or-no call but can stall it for months if they’re only brought in after the decision is made.
The person running it never had the budget or mandate to take it further, and the person who did have that authority was never really part of the pilot. Good results with no decision-maker attached just become an anecdote passed around in meetings, cited approvingly and acted on by nobody.
The three statistics in this piece were checked against their original publishers on 31 August 2026. The 26% figure comes from BCG’s own October 2024 report, based on a survey of 1,000 CxOs and senior executives across 59 countries; it’s a self-fielded and self-analyzed survey, which is a normal limitation of vendor-published research rather than a flag specific to this number. The 95% figure comes from MIT Project NANDA’s ‘The GenAI Divide: State of AI in Business 2025,’ built from a review of more than 300 public AI deployments, 52 structured interviews and a 153-person survey; the report has circulated mainly as a PDF rather than being hosted on an official MIT domain, and this piece cites the report’s own content directly rather than a summary of it. The 30% figure from Gartner is a prediction made in July 2024 about outcomes by the end of 2025, not a measured result, and it’s presented that way throughout rather than as something already observed. One widely repeated claim was deliberately left out: the figure that ‘80% of AI projects fail,’ often attributed to a 2024 RAND Corporation study. RAND’s actual report is based on 65 interviews with data scientists and explicitly frames 80% as an outside estimate it is characterizing, not a number RAND itself measured, so it doesn’t meet the bar for a sourced statistic here. The Scale Threshold framework, the role-ownership guidance, and the six-step rollout path are Future Factors’ own thinking, developed from advising on AI rollouts, and are presented as a practical framework rather than as research findings.