Rollouts rarely fail on the tool. They fail on the four or five things that were everybody's job and therefore nobody's.
Six categories, and you need all six before day one: define the problem and pick a pilot, agree governance and data guardrails, select the tool, plan enablement, plan communication, and decide how you’ll measure. The item most often missing isn’t on any of those lists, it’s the owner. A checklist item with nobody’s name against it doesn’t get done and doesn’t get noticed until something breaks. On measurement, stop at usage numbers and you’ve measured access rather than adoption. The three things worth tracking are whether people use it, whether they still use it a month later without reminders, and whether the work itself has actually changed.
The version of this I’ve watched go wrong most often looks completely competent from the outside. Licences bought, a launch email, a decent one-hour session with a good trainer, a Teams channel for questions. Six weeks later the channel has four messages in it, three of them from the person who set it up.
Nothing on that list was wrong. The problem is what wasn’t on it, and it’s usually the same handful of things: nobody decided what people were allowed to put into the tool, nobody picked which specific piece of work it was for, and nobody’s name was against the question of whether any of it was working.
Those aren’t hard tasks. They’re just tasks that belong to nobody by default, which means they belong to nobody at all.
A checklist item with no name against it is a wish.
There’s a version of this problem at the top too, and it’s more common than most leadership teams would guess. In Microsoft’s 2026 survey of 20,000 AI-using knowledge workers, only around a quarter, 26%, said their leadership was clearly and consistently aligned on AI.[1] Three quarters of people are being asked to change how they work by a leadership group that hasn’t visibly agreed with itself yet.
Which is why the first half of this checklist is decisions rather than actions. If the decisions haven’t been made, the actions produce activity and not much else.
Here it is in one place. Copy it, put real names in the owner column, and be honest about the ones you can’t fill in yet, because those are the ones that will surface in month three.
| Category | The item | Typical owner |
|---|---|---|
| 1. Problem and pilot | One named business problem, written as a sentence with a number in it | Sponsor |
| One pilot group and one recurring workflow, chosen and written down | Sponsor | |
| What “this worked” looks like, agreed before anything starts | Sponsor and the pilot team lead | |
| 2. Governance and data | What can and cannot be put into the tool, in plain language people will read | Legal or compliance, with IT |
| Where outputs may and may not go: customer-facing, contractual, HR decisions | Legal or compliance | |
| Disclosure position: when someone has to say AI was involved | Legal or compliance | |
| Who to ask when it’s ambiguous, by name, with a response time | Named individual | |
| 3. Tool selection | Does it reach the material the workflow actually needs? | IT with the pilot team lead |
| Admin controls: what can be turned off, what’s on by default, who sees usage data | IT | |
| Procurement, security review, and the renewal date in someone’s calendar | Procurement | |
| 4. Enablement | Day-one pack: the workflow, three worked prompts, the quality bar, where to ask | L&D or the team lead |
| Reinforcement: what happens in weeks two to six, in the diary now | Line managers | |
| What managers specifically are expected to do differently | Sponsor | |
| 5. Communication | The job security question, answered directly, before anyone asks it | Sponsor |
| Who hears about it, when, and who hears it from their own manager first | Internal comms | |
| 6. Measurement | Use, persistence and impact defined as three separate questions | Sponsor |
| A review date in the calendar with the decision it feeds | Sponsor | |
| The stop condition: what would make you end this rather than extend it | Sponsor |
Our working checklist for a first AI rollout. The owner column is the part that changes outcomes; the items themselves are mostly obvious once written down.
Twenty items looks like a lot until you notice that roughly half of them are a single decision recorded in a sentence. The three that consistently take real work are the governance language, the day-one pack, and the manager expectations.
The rest of this article goes deeper on four of the six categories. For tool selection specifically, we’ve already got a longer piece on the questions worth asking before you buy: the AI tool evaluation checklist was written for marketing teams but the eight questions transfer to any function.
The failure here isn’t skipping the step, it’s doing it at the wrong altitude. “Improve productivity across the business” is a problem statement in the same way that “be healthier” is a fitness plan. What you want instead is something small enough that you could be wrong about it in eight weeks and everyone would know.
| The version that gets approved | The version that works | |
|---|---|---|
| Problem | “Our teams are spending too much time on manual work and we want to use AI to improve efficiency.” | “Our four bid writers spend roughly two days per RFP pulling answers out of past submissions, and we turn down about one RFP a month because we can’t staff it.” |
| Pilot group | “Early adopters across the business.” | The four bid writers. All of them, not volunteers. |
| Workflow | Not specified. | First-pass answers to the standard sections, drafted from the last three submissions, before a human rewrites. |
| What good looks like | “Increased adoption and positive feedback.” | “We bid on one more RFP a month by Christmas, and the writers say the first pass is worth editing rather than binning.” |
A worked example of the difference in altitude. The right-hand column can be proven wrong, which is what makes it useful.
Two things about the pilot group are worth arguing for even when they’re unpopular, and both tend to get negotiated away in the first planning meeting.
Take the whole team, not volunteers. A volunteer group tells you what happens when motivated people use a tool, which you could have guessed. It tells you nothing about your actual rollout, where a third of people are indifferent and two are quietly hostile. The indifferent third is the population you need data on.
Pick work that repeats. A pilot on a one-off project can’t distinguish between a tool that works and a week that went well. If the task doesn’t come round again within the pilot window, you’ll finish with an anecdote rather than a finding, and anecdotes are what get rollouts extended forever.
Who you pick matters more than the size. A finance team will hit a data-access wall in week one and you’ll learn about your governance gaps fast. A marketing team will produce visible output quickly and teach you about quality standards. An HR team will surface the confidentiality question before anything else, which is genuinely useful to learn early but will make the pilot look slow. None of those is the wrong choice; just know which lesson you’ve signed up for.
Governance has a reputation as the thing that slows rollouts down, and the version that does is the twelve-page policy nobody finishes. The version that speeds things up is four answers people can remember.
The reason to do it first is not caution. It’s that people are already using AI at work, and in the absence of a rule they’re making one up individually, which is a worse outcome than any policy you’d have written.
Decide what people are allowed to put in before you decide what tool they get.
The four questions, and what a usable answer looks like rather than a legal one:
Then make the split explicit for the workflow you’ve picked, so that “a human reviews it” doesn’t quietly become “somebody probably looked.” Writing it down takes ten minutes and it’s the thing you’ll want to be able to point at if anyone ever asks how a particular output got approved.
| What AI does | What the human still owns | How it gets checked |
|---|---|---|
| Drafts first-pass RFP answers from the last three submissions | Whether the answer is still true, and every commercial commitment in it | Bid writer initials each section against the source submission. No initials, it doesn’t go in the pack. |
| Summarises a policy document into a plain-language brief | Whether the simplification changed the meaning | The policy owner reads the brief against the original once, at first use, then quarterly. |
| Themes candidate feedback from an interview process | Any decision about an individual candidate, entirely | Hiring manager sees the raw notes, not only the themes. This one is a hard rule, not a preference. |
Fill in a row for each workflow in the pilot. If a row is hard to write, that workflow isn’t ready.
The standard shape is a launch session followed by silence, and the silence is where it fails. Not because the session was bad, but because the first time someone actually tries the thing on their own work is usually a Tuesday about nine days later, and there’s nobody around.
It helps to be precise about four words that get used interchangeably and mean quite different things, because the confusion between them is what produces the silence.
Training builds capability: after it, someone can do the thing in principle. Workflow design gives that capability somewhere to land, which is a real recurring task rather than a general encouragement to use AI. Reinforcement turns it into behaviour, through repetition and a manager who actually checks. Adoption is what you observe when all three have happened. It isn’t a lever you can pull on its own, which is why “driving adoption” as a standalone activity tends to produce posters. Most rollouts do the first, skip the second and third entirely, and then wonder why the fourth didn’t arrive.
Put the support where the first real attempt happens, not where the training calendar has room.
What actually goes in a day-one pack, and it fits on two sides of paper:
Here’s one filled in, so the level of specificity is visible rather than described. This is the bid-writing pilot from earlier.
| Section | What it actually says |
|---|---|
| The workflow | First-pass answers to the standard RFP sections. Input: the three most recent submissions plus this RFP’s question list. Output: a draft section you edit, not one you send. |
| Three prompts | Already run on our own last three submissions by Sam, who fixed two of them before this went out. Each names the section it’s for. |
| The quality bar | Good: every claim traceable to a specific past submission, hedging language preserved. Reject: any number, date or commitment you can’t find in the source. |
| Governance, on the same page | In: past submissions, the RFP itself, published product docs. Out: signed contracts, anything from the CRM. Nothing goes to the client without a named reviewer. Ask Priya in the bid channel, same day. |
| Where to ask | Priya, #bid-team channel, answers within the working day. Not a mailbox. |
A worked example of a day-one pack rather than a blank template. Every row names a person, a source or a limit.
And what managers are meant to do differently, because this is the item most likely to be assumed rather than stated. Something like: use it yourself visibly, ask about it in your next one-to-one rather than waiting for it to come up, and say out loud what a good output looks like at least once. That’s it. Microsoft’s own survey found only 13% of AI users say they’re rewarded for reinventing how they work with AI even when the results don’t land,[1] which is a fair proxy for how much cover most people feel they have to try something and have it not work.
On whether you need outside help: teams genuinely can start this internally, and if you’ve got someone with the time and the credibility to build the day-one pack, start. What structured training tends to add is speed and consistency. Someone who has watched a lot of these knows which pilot choices go wrong, and a team trained together ends up sharing a quality bar rather than having one confident person and eight who are guessing. That shared standard is the hard part to build from the inside. If it’s useful, our corporate workshops are built around a real recurring workflow per team rather than a tour of the tool, and we’ve written up why so many of the generic versions fail in why most corporate AI training fails.
Usage dashboards are seductive because they exist. Your admin console will hand you weekly active users and prompt counts without anyone having to think, and those numbers will go up during a rollout regardless of whether anything real is happening.
What they measure is access. Adoption is a different question, and it has three parts that need asking separately.
| Stage | The question | How you check it | What a bad answer looks like |
|---|---|---|---|
| Use | Are people doing it at all? | Ask the pilot group to say, in a sentence, what they used it for last week. Not a survey score, a sentence. | Everyone describes something different, which means there’s no workflow, just a tool. |
| Persistence | Are they still doing it a month later, without reminders? | Stop reminding them for three weeks and see what happens. Uncomfortable, and the only way to know. | It stops within a fortnight of the nudges stopping. |
| Impact | Has the work itself changed? | Compare against the number in your problem statement. The bid writers’ one: are we bidding on more RFPs? | You can’t answer it, because nobody wrote the number down in week one. |
Our standard measurement lens. Usage numbers belong in the first row only, as the easiest and weakest signal.
You haven’t measured adoption until you’ve stopped reminding people and looked again.
The persistence test is the one people resist, and I understand why: deliberately withdrawing support from a rollout you’re invested in feels like sabotage. It’s the only honest reading you’ll get, though. A workflow that survives three weeks of nobody mentioning it is real. One that doesn’t was being carried by the reminders, and it would have stopped in month four anyway, just further from anyone’s attention.
Write down the stop condition too, in week one, while you can still be objective about it. Something like: if the bid writers aren’t editing rather than rewriting by week eight, we stop and look at why rather than extending. Rollouts almost never get formally ended, they get quietly extended, and an extension with no new evidence behind it is how a pilot becomes a permanent fixture nobody evaluates. If you want the longer version of this argument, we’ve written it up in measuring your AI ROI without a data science degree.
The plan in most kickoff decks runs about six weeks and assumes governance happens in parallel with everything else. It doesn’t, because governance depends on people whose calendars you don’t control.
Here’s a shape that survives contact with a real organisation of, say, 150 to 500 people. It’s one team’s version rather than a template, and the useful part is the ordering rather than the exact weeks.
Problem statement with a number in it. Pilot group and workflow chosen. What “this worked” means, written down and agreed with the sponsor. Mostly conversations, almost no activity, and the most valuable three weeks in the plan.
Governance, running in parallel and starting early because it will take longer than you think. Four answers, plain language, plus the named person to ask. This is the item that slips.
Tool selection and security review against the workflow you already defined. Doing it in this order rather than the reverse is what stops you buying something that can’t reach your material.
Day-one pack built and tested on your real files by one person before anyone else sees it. Manager expectations agreed and said out loud.
Launch. Communication goes out with the job security question answered directly, and people hear it from their own manager first.
Reinforcement. Short check-ins, quality bar discussed in one-to-ones, prompts refined by the people using them. Week 10 is when it feels like it isn’t working. It usually is.
Stop reminding anyone. This is the persistence test and it needs three full weeks to mean anything.
Review against the number from week one, and make an actual decision: extend, change, or stop. The stop option has to be genuinely available or the review is a formality.
One team’s sixteen weeks, illustrating the ordering rather than prescribing dates. A smaller team can compress the middle; nobody can compress governance.
A small team of ten or fifteen can run this in about ten weeks, mostly by shortening reinforcement, and a large enterprise will spend longer on weeks two to six than the whole rest of the timeline. The sequence holds in both cases. What doesn’t survive compression is the persistence window, because three weeks is roughly the minimum before you can tell the difference between a habit and a recent memory.
If you’re earlier than this and still working out what the organisation is capable of absorbing, our piece on taking an organisation from AI awareness to AI fluency covers the stage before the pilot, and what AI enablement actually involves goes deeper on the enablement category specifically.
The move for this week is smaller than any of that. Take the checklist, put it in a document, and fill in only the owner column. Leave the items blank. The gaps will tell you where your rollout is going to stall, and you’ll know by lunchtime.
Six categories: a named business problem and a chosen pilot workflow, governance and data guardrails, tool selection, enablement, communication, and measurement. What matters more than the items is that each one has a named individual against it, because the parts of a rollout that fail are usually the ones that were everybody’s responsibility. Roughly half the items are a single decision recorded in a sentence rather than a project. The three that take real work are writing governance in plain language people will actually read, building a day-one pack tested on your own material, and stating what managers are expected to do differently, which is the item most often assumed instead of said.
For a mid-size organisation of 150 to 500 people running one pilot team, about sixteen weeks from first conversation to a real decision. Roughly three weeks defining the problem and pilot, governance running in parallel from week two because it depends on calendars you don’t control, tool selection in weeks four to six, the day-one pack in week seven, launch in week eight, four weeks of reinforcement, then three weeks with no reminders at all as a persistence test, then review. A team of ten or fifteen can compress the middle to around ten weeks. The persistence window is the part that can’t be shortened, because under three weeks you can’t distinguish a habit from a recent memory.
One executive sponsor owns the checklist as a whole and personally owns the problem statement, the pilot choice, the definition of success, the measurement decisions and the stop condition. Individual items sit elsewhere: governance and disclosure with legal or compliance, tool reach and admin controls with IT, procurement and the renewal date with procurement, the day-one pack with L&D or the team lead, and reinforcement with line managers rather than with the project. The distribution matters less than the rule that every line has a person’s name on it, not a function’s.
In our experience running these, it’s what happens between the launch session and the first real attempt. Someone sits through a good hour of training, then genuinely tries it on their own work about nine days later, on a Tuesday, with nobody around and no idea whether what came back is good. That gap is where rollouts quietly end. The related miss is that nobody tells managers what they’re supposed to do differently, so they assume the answer is nothing. Both are cheap to fix and neither appears on a typical project plan, because they’re behaviours rather than deliverables.
The same six categories, but the effort splits differently. A team of twelve can answer the governance questions in a single conversation and should still write the answers down, because the value is that everyone can recall them rather than that they’re formal. A large enterprise will spend more time on governance and procurement than on everything else combined. What genuinely doesn’t scale down is the measurement discipline: small teams tend to skip writing the number down in week one because everyone can see what’s happening, and then can’t answer the impact question in week sixteen. Write the number down regardless of size.
The checklist here comes from our own practice running AI rollouts and corporate workshops, and it’s presented as that rather than as a research-derived model. The two external figures both come from the Microsoft Work Trend Index 2026 annual report itself, read on 26 August 2026, rather than from coverage of it. One widely-circulated framing of that report was deliberately left out: several summaries render its findings as “manager behaviour, not training, drives AI adoption.” The report doesn’t say that. It says organisational factors correlate more strongly with reported AI impact than individual ones, and it explicitly describes those as statistical associations rather than causes, with every variable self-reported by the same person at the same moment. The timeline is one team’s worked shape rather than a benchmark, and it’s labelled that way, because we don’t have a defensible average rollout duration and neither does anyone else quoting one.