Most L&D teams are sitting on more raw material than they'll ever get through. That, rather than the blank page, is the job worth handing over.
The decision rule for Copilot in L&D: point it at material that already exists inside your organisation, and use it to read, compare, condense, restructure and check. Don’t use it to invent a curriculum from a blank page. This article gives six copy-ready prompts covering course outlines from existing documents, turning a subject-matter-expert interview into training content, facilitator guides, application-level assessment questions, feedback theming, and auditing an old course for what’s gone stale. It also sets out clearly what still needs an instructional designer: learning objectives, sequencing decisions, accessibility, and the judgment about what the business actually needs people to do differently.
A learning manager I was talking to had ninety minutes of recorded interview with the company’s best field engineer, the one everybody rings when something goes wrong. She’d been sitting on it for five weeks. Not because the content was thin, but because turning ninety minutes of a very knowledgeable person talking into something teachable is genuinely slow work, and it kept losing to whatever was on fire that week.
That transcript is the shape of job Copilot is actually good at. The raw material exists, it’s text, it’s already inside the organisation, and the work is reading, structuring, and condensing rather than inventing.
So the decision rule I’d give any L&D team is narrower than the marketing around these tools suggests, and it’s worth applying before you open anything:
Point it at what you already have. Generating from nothing is where it’s weakest and where you’re most needed.
That rule does more work than a feature list would, because it tells you what not to do. “Design me a leadership programme” produces something that looks like a leadership programme and isn’t anchored to anything your business needs. “Here are the eleven exit interviews from the last two quarters, tell me which management behaviours come up more than twice” produces something you can act on by Thursday.
Concretely, the L&D workload splits into three tiers on this:
| Tier | Kind of work | Examples from a normal L&D week |
|---|---|---|
| Hand over | Reading, comparing, condensing, restructuring material that already exists | SME transcript to draft content. Old course audited against a new policy. 400 feedback responses themed. Two versions of a deck compared for what changed. |
| Hand over carefully | First drafts where you’ll need to rewrite substantially and you know what good looks like | Facilitator guide from an existing deck. Assessment questions. A summary email to a sponsor. Anything a learner will see unedited: nothing. |
| Keep | Decisions about what people must be able to do differently, and for whom | Learning objectives. Sequencing. What to cut. Whether training is even the right intervention. Accessibility of the finished materials. |
Our recommended split for L&D work, based on where these tools are reliable rather than on what they’ll attempt.
One product note, because it affects which of these is practical. Microsoft ships two agents inside Copilot that matter here. Researcher handles deep multi-step questions across your files and, if your admin allows it, the web, and it cites its sources; Microsoft caps it at 25 queries per user per month and its own documentation says it can’t process images in input documents.[1] Analyst is the one to reach for when the material is a spreadsheet or a CSV, because it calculates statistics, finds outliers and returns charts.[2] Microsoft’s own guidance is that Analyst, not Researcher, is the better fit for Excel work.[1]
These are written to be pasted and edited, not admired. Each one names the material it expects, so if you don’t have that material the prompt isn’t for you yet.
Three habits make all six of them work noticeably better, and they take under a minute between them:
| You are helping me build a training outline from a document that already exists. Read the attached [POLICY / PROCESS DOC / PLAYBOOK]. First, list every distinct thing a [ROLE, e.g. new field engineer] would need to be able to DO after reading it. Actions, not topics. If something in the document is background rather than an action, put it in a separate list called “context only”. Then group the actions into no more than five modules, ordered so that each module only depends on ones before it. For each module give me: the module name, the actions it covers, and one sentence on how I would know a learner can do them. Flag anything in the document that is ambiguous enough that you had to guess what the reader is meant to do. I want that list. |
Prompt 1. The “context only” list and the ambiguity flag are the two parts people delete, and they’re the two parts that make the output useful.
How to adapt it: swap the role, and if your material is several documents rather than one, add “where the documents disagree, tell me where rather than picking one.” That disagreement list is often more valuable than the outline.
| Attached is a transcript of an interview with [ROLE], who is one of the most experienced people we have at [TASK]. Pull out three things separately: 1. The steps they describe, in the order they do them. 2. The judgment calls: the places where they say something like “it depends” or “you get a feel for it”. Quote them directly. Do not smooth these into rules. 3. The mistakes they mention other people making. Then tell me which of the judgment calls in list 2 you think cannot be taught from the transcript alone and would need a worked example or a conversation with them. Use only what is in the transcript. If a step seems to be missing, say so rather than filling the gap. |
Prompt 2. Separating steps from judgment calls is the whole point; the judgment calls are the expertise, and they’re what a generic model would otherwise flatten.
How to adapt it: if you’re working from a recording rather than text, get the transcript first. And keep the “use only what is in the transcript” line in every version. It’s the single instruction that most reduces confident invention.
The gap between an outline and something you could actually run is where most L&D time disappears, and it’s mostly assembly rather than thinking. Two prompts for that, and then a worked example of what comes back so you can calibrate what “good enough to edit” looks like.
| Attached is the deck for [SESSION NAME], which runs [DURATION] with roughly [N] participants. Draft a facilitator guide. For each slide give me: the point being made in one sentence, a suggested question to ask the room, and the most likely place a participant gets confused or pushes back. At the end, give me a timing plan that adds up to [DURATION] including a break, and tell me which two slides I should cut first if I’m running fifteen minutes over. Write the questions in plain spoken English, the way someone would actually say them out loud, not as bullet points. |
Prompt 3. The “which two slides to cut” instruction is the one facilitators tell me they use most, because it forces a priority decision you can disagree with.
| From the attached module, write eight assessment questions for [ROLE]. Rules: no question may be answerable by someone who has memorised the wording without understanding it. Every question must describe a specific situation the person would actually face and ask what they would do. For each question give me the answer you consider correct, and one plausible wrong answer that someone who half-understood the material would choose. Tell me what misunderstanding that wrong answer reveals. Flag any question where the module doesn’t actually give enough information to answer it. |
Prompt 4. The distractor-plus-diagnosis structure is what turns an assessment into a source of information about your own materials.
Here’s roughly what came back from Prompt 4 on a compliance module, so you can see the standard rather than imagine it. Question: “A supplier sends you an updated certificate by email two days after the deadline. You’ve already submitted the quarterly return. What do you do?” Proposed correct answer: log it, submit an amendment, note the late receipt. Plausible wrong answer: file it and include it in next quarter’s return. Diagnosis: the learner has understood that the certificate matters and missed that the return is a point-in-time declaration.
That diagnosis line is the bit worth reading. It told the learning manager something about her own module, which was that the point-in-time nature of the return was mentioned once, in a sentence, on slide fourteen. She’d have defended that slide as clear if you’d asked her the day before.
Open-text feedback is the thing most L&D teams collect diligently and read once. Four hundred responses to “what would have made this session more useful?” is genuinely too much to hold in your head, so what usually happens is someone skims fifty, forms an impression, and reports the impression.
This is the workflow where Analyst rather than Researcher is the right tool, because the material is tabular. Analyst is built for exactly this: attach the file, ask a question in plain language, get back statistics, trends and outliers as a readable report.[2]
| Attached is [N] open-text responses to the question “[QUESTION]” from [SESSION / PROGRAMME]. Group them into themes. For each theme give me: a name, how many responses fall into it, and three verbatim quotes, including one that is critical if there is one. Then give me a second list called “small but specific”: responses that don’t fit any theme but name a concrete problem. Do not discard these into an “other” bucket. Finally, for each theme tell me whether acting on it would change the content, the delivery, or the audience we invited. If it’s none of those three, say so. |
Prompt 5. The “small but specific” list exists because the single most actionable comment in a feedback set is usually a one-off, and theming is designed to bury exactly that.
The last instruction is the one that changes what you do on Monday. A theme called “wanted more examples” is unactionable. The same theme sorted into content, delivery, or audience turns into either “add two worked examples to module three,” “stop lecturing for the first twenty minutes,” or “we invited the wrong people.”
Two honest warnings on this one. It will over-weight the articulate. People who write three sentences get themed; people who write “fine” don’t. And it has no idea who was in the room, so a theme from eleven people can look identical to a theme from eleven managers, which is a completely different signal. Read twenty raw responses yourself before you look at the themes, every time. It takes four minutes and it recalibrates you.
There’s a particular irony in an L&D team having a Copilot licence nobody uses, which is that the failure mode is the exact one they’re employed to solve for everyone else.
Our framework for this is Tool x Workflows x Behavior, and we’ve written it up properly as its own piece rather than re-explaining it here: what makes an AI-powered professional different from everyone else using AI. The short version is that a capable tool with no recurring workflow to sit in, and no behaviour repeated often enough to survive a busy month, produces roughly nothing, however good the tool is.
What’s specific to L&D is which variable usually binds. It isn’t Tool, because you have the licence. It’s rarely Behavior either, because L&D people are unusually good at sticking to processes. It’s almost always Workflows, and the reason is structural: L&D work is project-shaped rather than cycle-shaped. A new programme, then a compliance refresh, then an onboarding redesign. Very little of it repeats in the same form six weeks later, which is precisely the condition a workflow needs.
So the useful move is to look for the parts of your year that are cyclical and start there. Post-session feedback every time you run something. The quarterly content review. The onboarding cohort that arrives every month. Those repeat in the same shape, which is why the prompt you refine on the first one still works on the fourth.
This is the section I’d most want an L&D leader to read, because the tools are now good enough at producing course-shaped objects that it’s genuinely easy to mistake the object for the outcome.
Four things stay with a person, and none of them are about writing quality.
Deciding what someone must be able to do differently. A model can read your policy and infer plausible actions from it. It cannot know that the actual problem is that regional managers approve things they shouldn’t because the escalation path is embarrassing, and no amount of policy training will touch that. Choosing the objective is a judgment about the business, and it usually requires information that isn’t written down anywhere.
Sequencing for how people actually learn. A logical order and a learnable order are different things. Difficulty needs to build, practice needs to come before the stakes rise, and some things have to be taught twice in different forms. Models produce logical orders reliably and learnable ones by accident.
Whether training is even the right intervention. Ask any model to design training for a problem and it will design training for the problem. It has no mechanism for coming back and saying the issue is a broken handoff between two teams and a course would waste everyone’s morning. Getting that call wrong is expensive in a specific way: you spend six weeks building something competent that couldn’t have worked, and the original problem is still there afterwards.
Accessibility, in the real sense. Not just alt text and contrast, though those matter and they’re checkable. Whether the reading level suits the audience, whether the examples assume knowledge only half the room has, whether someone joining remotely can actually participate in the exercise you designed for a room.
A model can sequence a course. It can’t decide what the learner has to be able to do on Monday.
Most people already sense this, which is encouraging. In Microsoft’s 2026 survey of AI users, 86% said they treat AI output as a starting point rather than a final answer and that they stay responsible for the thinking.[3] The gap between saying that and doing it under deadline is where the checking below earns its place.
| What Copilot does | What the instructional designer still owns | How it gets checked |
|---|---|---|
| Extracts candidate actions and judgment calls from an SME transcript | Which of those actions are actually the performance gap, and which are already fine | The SME reads the extracted list before anything is built. Fifteen minutes, and they always correct something. |
| Proposes a module order and drafts assessment questions | Sequencing for difficulty and practice, and whether the assessment tests the objective or the wording | You answer three of the questions yourself without the module open. If you can pass by recall, they get rewritten. |
| Themes 400 feedback responses and flags outliers | Which themes are signal, who was in the room, and what changes as a result | Read 20 raw responses first, then compare. Every time, not just when the themes look odd. |
| Diffs an old course against current documentation | Whether a difference matters enough to reopen the course at all | Every row marked “now wrong” is confirmed against the source document by a person before any rewrite. |
Fill in your own version of this before the first build, not after something reaches a learner with an error in it.
Six prompts in a document nobody opens is the normal outcome here, and it’s worth designing against deliberately rather than hoping enthusiasm holds. The sixth prompt is the one that makes the difference, because it attaches to something already in your calendar rather than needing you to remember it.
| Attached: our [COURSE NAME] materials from [YEAR], and the current [POLICY / PROCESS / SYSTEM] documentation. Go through the course section by section and tell me, in a table: the section, what it currently says, what the current documentation says, and whether the difference is (a) now wrong, (b) still true but incomplete, or (c) unchanged. Only include rows where something differs. Quote both sources directly for anything you mark as “now wrong”. Do not suggest rewrites yet. I want the differences first. |
Prompt 6. Holding it back from rewriting is deliberate. Ask for the diff and the fix at once and you get a rewrite you can’t audit.
Ask for the difference before you ask for the fix. A rewrite you can’t audit is a rewrite you have to redo.
That one belongs to the quarterly review, which is the point. Each of the six attaches to a moment that already exists, so nothing depends on anyone remembering. Not a new ritual, an existing one with a step added.
| Moment already in your year | Prompt | Who runs it, and when |
|---|---|---|
| Every time a session finishes | Prompt 5, feedback theming | Whoever ran the session, within 48 hours while the room is still in their head |
| Quarterly content review | Prompt 6, the staleness diff | Content owner, first week of the quarter, one course at a time |
| Any new SME conversation | Prompt 2, transcript to teachable content | Same day as the interview, before your memory of it fades and starts filling gaps |
| Before a session runs for the second time | Prompt 3, facilitator guide | Facilitator, the week before, because the first run always changes it |
The pairing matters more than the prompt quality. A good prompt with no moment attached is a document.
Then a quality bar, so that “we use Copilot for this now” doesn’t quietly become “we ship whatever it produced.” Four checks, and they take under ten minutes on a normal build:
When a prompt has survived a few real runs and the quality bar consistently passes, it’s earned a place in a shared library, and the record is worth keeping properly. Here’s what one entry looks like filled in, because a stored prompt with no context around it gets ignored by everyone except the person who wrote it.
| Field | Entry |
|---|---|
| Name | SME transcript to teachable content (Prompt 2) |
| Use it when | You have a transcript of an experienced person explaining how they do something, and you need steps and judgment calls separated. |
| Don’t use it when | You only have notes rather than a transcript. It fills gaps in notes and you won’t spot where. |
| What good output looks like | The judgment-call list is longer than you expected and contains direct quotes with hedging language still in them. |
| Known failure | It smooths “it depends” into a rule if you drop the “do not smooth these” line. Keep that line. |
| Owner and last run | Priya, 14 August 2026, on the field engineer transcript. |
A worked example of a library entry, not a blank template. The “don’t use it when” and “known failure” rows are the ones that stop a colleague wasting an afternoon.
On knowing whether any of this worked, prompt counts and licence usage reports won’t tell you. What we look at, in this order: are people actually running these (use), are they still running them two months later without anyone asking (persistence), and has anything about the finished materials or the turnaround time genuinely changed (impact). Only the third one is the point, and it’s the only one your L&D metrics probably already capture.
Plenty of teams get through this on their own with a shared prompt document and someone stubborn enough to maintain it, and there’s a real argument for starting that way. What structured training tends to add is speed and consistency: the whole team ends up with the same quality bar rather than one person with good habits, which is the part that’s genuinely hard to build from the inside. If you want a hand with that, our corporate workshops run this on your own material rather than a generic case study. Either way, the first move is the same. Take the transcript you’ve been meaning to get to, and run Prompt 2 on it this afternoon.
The reliable uses all involve material that already exists inside your organisation: turning a subject-matter-expert interview transcript into structured content, drafting a facilitator guide from a deck you already run, diffing an old course against current policy documentation, writing application-level assessment questions from an existing module, and theming open-text feedback. The pattern is reading, comparing, condensing and checking rather than inventing. Where it’s weakest is designing a programme from a blank page, because that requires deciding what the business needs people to do differently, and no model has access to that. Attach the material first and ask it to describe what it can see before asking it to do anything.
Attach the source document and ask for actions rather than topics: “Read the attached policy. List every distinct thing a [role] would need to be able to DO after reading it. Actions, not topics. Put anything that is background rather than an action in a separate list called context only. Then group the actions into no more than five modules, ordered so each module only depends on ones before it, and for each give me the module name, the actions it covers, and one sentence on how I’d know a learner can do them. Flag anything ambiguous enough that you had to guess.” The ambiguity flag is the part people delete and the part that most improves the source document.
Yes, and for tabular feedback the Analyst agent is the right one rather than Researcher. Microsoft’s own guidance is that Analyst is better suited to spreadsheet work; it calculates statistics, identifies trends and outliers, and returns a readable report with charts. Two limits are worth knowing before you rely on it. It over-weights articulate respondents, because someone who writes three sentences gets themed and someone who writes “fine” doesn’t. And it has no idea who was in the room, so a theme raised by eleven people looks identical to one raised by eleven managers. Read twenty raw responses yourself before looking at the themes.
No, and the reason isn’t writing quality. Four things stay with a person: deciding what someone must be able to do differently (a judgment about the business that usually isn’t written down anywhere), sequencing for how people actually learn rather than in logical order, judging whether training is the right intervention at all rather than a broken process, and accessibility in the real sense of reading level, assumed knowledge and whether remote participants can join in. Models are now good enough at producing course-shaped objects that it’s easy to mistake the object for the outcome, which is exactly why the objective decision matters more than it used to.
The practical difference is reach rather than capability. Copilot sits across the material in your Microsoft environment, so it can read the exit interviews in Outlook, the policy in SharePoint and the deck in OneDrive in the same conversation. Your LMS’s built-in AI usually only sees what’s already inside the LMS, which is the finished courses rather than the raw material they came from. That makes Copilot the better fit for the upstream work of getting from source material to a draft, and your platform’s own features often better for what happens after publication, like generating variations or nudging learners. They aren’t really competing for the same job.
Every Copilot capability described here was checked against Microsoft’s own current support and Learn documentation on 26 August 2026, not against press coverage or roundup blogs, because agent availability and limits have moved repeatedly this year. Three things are deliberately not stated because they couldn’t be confirmed from Microsoft’s own pages: any licence price, whether the Researcher agent’s 25-query monthly cap is counted separately from other agents, and whether Copilot can read images inside an uploaded deck. On that last one, Microsoft’s Researcher FAQ says plainly that it cannot process images as input, so this article assumes text only. The article contains no interface screenshots, and that’s a deliberate choice rather than an oversight: it gives prompts and a decision rule rather than navigation steps, so there’s nothing to show that a screenshot would clarify. The example outputs described are illustrative of the shape and standard to expect, not transcripts from a named client.