Your AI training was accurate the day you built it. That is the problem. Here is how to design a program that survives the tools changing underneath it.
Skill decay is real and it is worst for exactly the kind of cognitive, judgement-heavy skills AI training teaches: a meta-analysis of 53 studies found skill loss reaching an effect size of -1.4 after a year without practice[1]. Layer on top of that the fact that OpenAI published 78 dated ChatGPT release-note entries in the first seven and a half months of 2026 alone[4], and you have a curriculum with a shelf life measured in weeks. This guide covers why programs decay, the three causes worth fixing, how to design content that tolerates drift, the review cadence that works without a full rebuild, why you need one named owner, what to measure instead of attendance, and a quarterly refresh checklist you can run on Monday.
A learning director emailed me in June about a Copilot curriculum her team had built in February. Four modules, screen recordings, a workbook, the lot. She wanted to know how much of it she needed to redo. The honest answer was most of the screen recordings, because the interface in them no longer matched what her people saw when they opened the app.
She had done nothing wrong. She built a good program, and then the ground moved.
Here is the thing that makes AI training genuinely different from every other subject an L&D team handles. Compliance training goes out of date when a regulation changes, which is maybe once every couple of years and comes with warning. Leadership training barely goes out of date at all. AI training goes out of date on a schedule closer to a news cycle.
You can check this yourself rather than taking my word for it. OpenAI keeps a public changelog for ChatGPT, and between 1 January and 14 August 2026 it carried 78 separate dated entries[4]. That’s roughly one every three days. Anthropic’s release notes show seven new flagship Claude models shipped in under six months of 2026, and a January entry recording that Opus 4 and 4.1 were removed from the model selector entirely[5]. Microsoft ships generally available Copilot feature updates on a two-week rhythm, with each batch headed “Updates released between” one date and another a fortnight later[6].
Now think about what that means for a slide that says “click the model selector and choose Opus 4.1.”
And that’s only half the decay. The other half would happen even if the tools froze solid tomorrow, because a training event on its own has always faded. A meta-analysis of 53 studies, covering 189 independent data points, found substantial skill loss with non-use, growing from an effect size of roughly zero immediately after training to -1.4 after more than 365 days without practice[1]. The same analysis found that cognitive tasks were more susceptible to decay than physical ones. AI skills are cognitive tasks. You are teaching the least durable category of thing there is, about the fastest-moving tool set anyone has ever put in front of an office worker.
When we audit a stalled AI program at Future Factors, the diagnosis is nearly always one of three things, and usually all three at once.
The content described a product that has since changed. Screenshots are the obvious casualty, but the expensive version is subtler: a workflow you taught is now the wrong workflow, because the tool grew a better way to do it and your people are still doing it the slow way you showed them in March.
The program was an event. People attended, felt good about it, went back to their inbox. Nothing brought them back to the material.
This is the one the research is bluntest about. Dunlosky and colleagues reviewed ten common learning techniques for Psychological Science in the Public Interest and rated only two as high utility: practice testing and distributed practice. Rereading and highlighting, the two things learners actually default to, landed in the low-utility group[2]. Watching a recorded demo is the corporate cousin of rereading. It feels like learning and transfers almost nothing.
Spacing is equally well established. A synthesis of 839 assessments across 317 experiments found spaced presentations produced markedly better final-test performance than massed ones, with only 12 of 271 comparisons showing no effect or a negative one[3]. The same work found the optimal gap between sessions grows as the retention target grows: if you want people to still have this in six months, your follow-ups need to be weeks apart, not hours.
A single launch day is the worst possible schedule, dressed up as the most efficient one.
Nobody’s calendar has a recurring block that says “check the AI curriculum.” The program was a project with an end date rather than a system with a caretaker, so the moment the launch was done, so was the maintenance.
Microsoft’s 2026 Work Trend Index puts a number on how much the organisational side matters relative to the individual side. Its analysis found that organisational factors like culture, manager support and talent practices account for more than twice the reported AI impact of individual factors like mindset and behaviour, 67% versus 32%[11]. Microsoft is careful to note these are statistical associations rather than causal effects, and I’d repeat that caution. But the direction is hard to argue with, and it lines up with what the same report calls “blocked agency”: 10% of AI users who “have built strong skills but lack the systems to apply them”[11].
That last group is the saddest outcome of a good training program with no maintenance behind it. You taught people well. Then you let the system around them stay exactly as it was.
You can’t stop the tools changing. You can decide how much of your program breaks when they do.
The single most useful design move is to separate your content into two layers and treat them completely differently.
The durable layer is the stuff that will still be true in two years: how to write a prompt that specifies a goal, a source and an expected format; how to check an AI output before you send it; when a task genuinely shouldn’t go to AI at all; what your organisation’s data rules are. This is 70% of the value and it barely moves.
The volatile layer is everything tied to a specific product surface: menu paths, button names, model names, which app has which feature this month, pricing. This is the part that rots.
Most programs mix the two into a single artefact, so a button rename forces you to rebuild a module about critical thinking. Keep them physically separate. Durable material goes in the core deck or the workbook. Volatile material goes in a short, dated appendix, a one-page “current as of” cheat sheet, or a linked internal page. When Microsoft renames Agent Mode to “Edit with Copilot,” which it did earlier this year, you update one page instead of four modules.
Date everything visibly. Put “Current as of August 2026” on any screen-specific material, in a place the learner sees. This does two useful things: it tells people to trust their own eyes over your screenshot when the two disagree, and it makes stale content obvious to you at a glance instead of hiding in a folder.
Teach the pattern before the click path. “Point Copilot at a specific file rather than asking in the abstract” survives a redesign. “Type slash, then choose from the dropdown” does not. When you must teach a click path, teach it as an example of the pattern rather than as the lesson itself.
Use the vendor’s own words for anything you’ll have to re-check. If a capability claim in your material links to Microsoft’s or OpenAI’s documentation page, your quarterly review becomes a link check rather than a research project. This is a small thing that saves an enormous amount of time in month nine.
None of this is exciting. It’s filing, essentially. But the difference between a program that takes two days a quarter to maintain and one that takes two weeks is almost entirely down to whether someone made these choices at the start.
If you’re building the underlying material rather than maintaining it, our guide to creating training materials your team will actually use covers the production side in more detail.
The instinct when you accept that AI content decays is to review everything constantly. Don’t. A monthly full review will get abandoned by month three, and an annual one is useless. What works is a tiered cadence where different things get checked at different rhythms, plus a small set of triggers that can interrupt the schedule.
Future Factors’ recommended maintenance loop. This is an author’s framework, not survey data.
Step 4 is the one people skip, and it’s the one that does double duty. A two-minute update on what changed this quarter is both a content correction and a spaced-practice touchpoint, which is precisely the distributed-practice pattern the research supports[3]. You get maintenance and reinforcement out of the same fifteen minutes of work.
Monthly, 30 minutes: check the changelogs of the two or three tools you actually teach. Not all AI news. Just yours. Write down anything that breaks a step in your material. Most months this list is empty or has one item.
Quarterly, half a day: run the full loop above. Update screenshots, verify every vendor link still resolves, refresh the “current as of” date, and send the re-seed item.
Annually, two days: question the durable layer. Are the roles still right? Are the example tasks still the tasks people do? This is where you’d notice that the finance team’s monthly close process changed and half your examples are about a report nobody runs anymore.
Some things shouldn’t wait for the quarter. Pull the review forward when any of these happen:
That fourth one is worth taking seriously as an early warning system. In my experience people report a broken step roughly twice before they quietly stop trusting the material altogether, and you never hear about it again.
Look, I’ll be blunt about this one, because it’s where most of the good intentions in this article go to die.
An AI steering group with eleven members and a quarterly meeting will not keep your curriculum current. A named person with two days a quarter allocated to it will. The difference isn’t seniority or budget. It’s that a committee has no calendar.
The ownership problem shows up in the data. LinkedIn’s 2025 Workplace Learning Report found that only 36% of organisations qualify as “career development champions” with robust programs that yield business results, while 33% have no initiatives or are just getting started[9]. The same report found half of respondents citing that managers lack proper support as a top barrier, and that only 15% of employees said their manager had helped them build a career plan in the past six months, down five percentage points year on year[9].
Meanwhile the demand keeps climbing. The Conference Board reported in July 2026 that while 55% of workers regularly use AI, only a third had taken part in employer-provided AI training in the previous six months, and 28% said their employer provides no AI training at all[15]. Gallup’s May 2026 survey found 52% of US workers now use AI in their role, with 30% using it frequently[16]. People are using these tools whether or not anyone is teaching them.
Be specific about this or it evaporates. The owner is accountable for four things and nothing else:
Notice what’s not on that list: designing the curriculum, running every session, choosing the tools. The owner can delegate all of it. What they can’t delegate is noticing when it’s rotting.
On seniority: this does not need to be a director. It needs to be someone whose manager knows this is part of their job, and who has the standing to email a vendor’s admin for a licence question without escalating. A capable L&D specialist with two protected days a quarter beats a VP with none.
Prosci’s research across more than 2,600 change practitioners found 88% of projects with excellent change management met or exceeded objectives, against 13% with poor change management, which they summarise as roughly seven times more likely to succeed[10]. There’s a second finding in that work worth quoting at anyone who says a review cadence will slow things down: projects with excellent change management were nearly five times more likely to be on or ahead of schedule[10].
Attendance tells you people came. Satisfaction scores tell you they didn’t hate it. Neither tells you anything about whether the work changed.
This isn’t a new complaint. Donald Kirkpatrick published his four levels of training evaluation in a series of articles starting in 1959, precisely to push the field past reaction and into learning, behaviour and results[13]. Sixty-seven years later most AI training reporting is still a completion rate and a smile sheet.
The evidence on why that matters is reasonably strong, and I want to state it carefully rather than overselling it. A 1997 meta-analysis of training criteria found the average correlation between trainee reactions of any type and immediate learning was .08, with utility-focused reactions faring better at .26. Later work by Sitzmann and colleagues in 2008 pushed back on the strongest version of that claim, concluding that reactions do have a predictive relationship with learning outcomes but not one strong enough to use reactions as an indicator of learning. So the defensible statement is the narrow one: your satisfaction scores are not a measure of whether anyone learned anything, and they should never be the headline number.
| Level | The lazy version | A better AI-specific measure | Where to get it |
|---|---|---|---|
| Reaction | Satisfaction score | “Will you use this next week, and on what?” Named task required. | End-of-session form |
| Learning | Quiz completion | A short retrieval check two weeks later, not on the day | Spaced follow-up |
| Behaviour | Licence count | Active users as a share of enabled users, and repeat use over 30 days | Vendor admin centre |
| Results | Hours “saved” | One named process that now runs differently, with a before and after | The team that owns the process |
Kirkpatrick’s four levels[13] mapped to AI-specific measures. The right-hand columns are the author’s recommendations, not published benchmarks.
Two practical notes on that table.
The two-week delay on the learning check is deliberate and it’s the cheapest upgrade in this whole article. Testing on the day measures short-term recall. Testing a fortnight later measures something closer to what people will actually still have, and the act of retrieval itself strengthens retention, which is the practice-testing effect Dunlosky’s review rated as high utility[2]. You get a measurement and an intervention from the same email.
On the results row: resist the urge to report aggregate hours saved. Those numbers are almost always modelled rather than measured, and a sceptical finance director will take them apart in about forty seconds. One process, genuinely changed, with a before and after that the process owner will sign their name to, is worth more than a spreadsheet of estimates. It’s also much harder to produce, which is exactly why it’s credible.
Here’s the actual checklist. Half a day, once a quarter, in this order. I’ve put it in the order that fails fastest, so if you run out of time you’ve done the items that matter most.
Run that four times and you have a program that’s a year old and still accurate, which puts you ahead of most organisations I see.
One last thing, and it’s the thing I’d want you to take away if you skim the rest. The reason a maintained program beats a brilliant one-off isn’t that the content is better. It’s that maintenance is the only mechanism you have for reinforcement, ownership and currency all at once, and those are the three variables that decide whether a licence turns into a habit. A program that stays alive is doing pedagogical work every quarter, whether or not anyone calls it that.
If you’re earlier in the journey and still working out what your organisation actually needs, the companion piece to this one on taking your organization from AI awareness to AI fluency covers the maturity side: where you are now, and what has to change at each stage. If you’re building the first version rather than maintaining an existing one, start with our step-by-step playbook for training your team on AI, and the nine-step upskilling walkthrough covers the same ground faster. For the specific case of a Microsoft rollout, our piece on why most Copilot rollouts fail goes deeper on the adoption side.
Two things happen at once. The tools change (OpenAI published 78 dated ChatGPT release-note entries in the first seven and a half months of 2026[4], and Microsoft ships Copilot updates fortnightly[6]), so your screenshots and click paths stop matching reality. And unpractised skills decay: a meta-analysis of 53 studies found skill loss reaching an effect size of -1.4 after a year without use, with cognitive skills more vulnerable than physical ones[1]. A single launch event addresses neither.
Use tiers rather than one blanket schedule. Spend 30 minutes monthly scanning the changelogs of the two or three tools you actually teach. Spend half a day quarterly running a full refresh: screenshots, vendor links, dates, and one reinforcement item pushed to past learners. Spend two days annually questioning the durable material itself, meaning roles, example tasks and workflows. Override the calendar when a tool ships a major version, your organisation adds or switches vendors, or two learners report a step that no longer matches.
One named person, not a committee. Committees have no calendar, so nothing recurring ever happens. The owner needs roughly two protected days per quarter and accountability for exactly four things: currency (content matches the tools as of a stated date), cadence (the reviews actually happen), reinforcement (something reaches learners between sessions), and one measure reported to a sponsor. Seniority matters less than protected time. A capable L&D specialist with two days beats a VP with none.
It depends on which layer changed. Split your material into a durable layer (prompt structure, verification habits, data rules, when not to use AI) and a volatile layer (menu paths, model names, button labels, pricing, screenshots). A rename or interface change should only ever touch the volatile layer, which is a patch of an hour or two. You rebuild when the durable layer is wrong: the roles you designed for changed, or the example tasks are no longer tasks anyone does. That is usually an annual question, not a quarterly one.
Not with attendance or satisfaction scores. A 1997 meta-analysis found trainee reactions correlate about .08 with immediate learning, and later work concluded reactions are not strong enough to use as an indicator of learning. Measure instead at Kirkpatrick’s higher levels[13]: a short retrieval check two weeks after the session rather than on the day; active users as a share of enabled users plus repeat use over 30 days from your vendor’s admin centre; and one named business process that now runs differently, with a before and after the process owner will vouch for.
This guide is based on Future Factors’ work rebuilding stalled AI training programs for corporate teams, combined with primary research on skill decay, spaced practice and change management. Every statistic here was traced to its original source: where a widely-repeated figure could only be found on blogs citing other blogs (notably the “AI skills half-life” claim), it was left out and the omission stated in the text. Vendor release cadences were counted directly from OpenAI’s, Anthropic’s and Microsoft’s own public changelogs on 20 August 2026.