Microsoft thinks managing AI agents is going to become part of a lot of our jobs. When I first heard the phrase, my mind went somewhere far more dramatic than what this looks like on a normal Tuesday.
You don’t need a small army of AI workers running your marketing department to get started. Plenty of teams will end up running several agents, and some will genuinely need to, but that isn’t where you begin. Most teams should start with one agent doing one annoying job everyone keeps putting off. Becoming an agent boss has less to do with building impressive systems and much more to do with something you already have to learn at work: how to delegate well. The difference is that some of what you delegate now goes to software, and software needs more managing than people expect. It needs context, boundaries, someone responsible for checking the work, a definition of good, and a point where you decide whether it deserves to keep the job.
The phrase comes from Microsoft, who describe an agent boss as someone who builds, delegates to, and manages AI agents to increase their impact, and expect more of us to end up directing small groups of agents with skills like research, analysis and coordination[1]. That can sound futuristic until you translate it into a normal workday.
Say you manage marketing. Today someone on your team spends 45 minutes every Monday checking competitors. Tomorrow an agent does the first scan and hands you a short summary of what changed. You still decide whether those changes matter, whether your team should respond, and what the response should be. The judgment is still yours. What’s moved is the 45 minutes.
That distinction matters, because the first question people usually ask about agents is “how much work can I hand over?” I think there’s a better one: which parts of this work actually require my judgment, and which parts are stopping me from spending more time on that judgment? That second question leads to much better use cases, and it’s the question this whole article is built around.
You do not need a small army of AI workers running your department to get started. Some teams will genuinely grow into that, and there’s nothing wrong with it, but it isn’t step one. Most teams should start with one agent doing one job everyone keeps putting off. Maybe somebody checks six competitor websites every Monday. Maybe customer feedback piles up because nobody has time to read hundreds of comments. Maybe every new campaign begins with someone digging through old briefs trying to remember what was already tried six months ago.
This is the part people skip. Someone hears about agents and immediately asks “what agent should we build?” I’d start one step earlier and ask what work we’re already doing that takes more human effort than the work deserves.
Start with the work, not the agent.
Look at your week. Which tasks make you think, I cannot believe we’re doing this again? Maybe someone repeatedly gathers the same information. Maybe someone spends hours sorting feedback before a pattern becomes visible. Maybe a team member rebuilds the same report every Friday, or research keeps sliding down the list because everyone has something more urgent. Those are better starting points than asking AI to take over your most strategic work.
Three or more yeses and it’s worth a closer look.
| Marketing task | Agent fit | Why |
|---|---|---|
| Monitor competitor pricing | Strong | Repetitive research with a clear output |
| Summarise customer feedback | Strong | Large amounts of information to sort |
| Analyse recurring campaign performance | Strong | Repeatable analysis |
| Prepare a first campaign brief | Medium | AI can prepare, a human should decide |
| Schedule approved content | Medium | The agent takes an external action |
| Develop final brand positioning | Weak | Needs significant judgment and context |
| Respond autonomously to a PR issue | Poor | High consequence and high judgment |
A starting shortlist for marketing teams. Author’s framework, not survey data.
There’s a pattern in that table worth naming. The best first agents usually aren’t glamorous. A competitor watcher, a campaign historian, a customer-feedback analyst, a meeting-prep agent. Those jobs matter, but they don’t need your best strategic thinking spent collecting the raw material.
The rule I’d hold onto: automate preparation before judgment. If someone spends an hour gathering information before spending ten minutes making a decision, I want AI helping with the hour. That gets you a useful starting point without asking AI to make calls the team should still own.
Once you know which job you’re handing over, resist the urge to start building. Write the job description first, and I mean a real one. “Help me with competitor research” isn’t enough. Neither is “be my marketing research agent.”
Before an agent does any work I want to know what business process it’s improving, what information it can reach, what a good result looks like, who owns it, and how we’ll know whether any of this has value.
| Field | Competitor Watch, owned by the demand gen lead, reviewed quarterly |
|---|---|
| Business process the agent improves | Competitive research. We were checking six competitor sites by hand, badly, roughly once a month. |
| What good looks like | A Monday note listing only what changed since last week, with a link to each change. |
| What the agent can access | Public pricing pages, the Meta and Google ad libraries, our press-mention feed. Nothing internal. |
| What the agent must never do | Email anyone. Post anything. Touch the CRM. Draft public copy. |
| Who reviews it, and when | Demand gen lead, first Monday of each quarter. Checks the sources still exist and the note still helps. |
| How we know it worked | It surfaced a competitor pricing change early enough that we changed a campaign because of it. |
A worked example you can copy. The two rows teams leave blank are what it must never do, and who reviews it.
Two rows deserve more attention than the rest, because they’re the two teams leave blank. What the agent must never do is where you decide calmly, in advance, that this thing doesn’t email anyone and doesn’t touch the CRM. Much easier to settle now than in the moment when it would be convenient.
Who reviews it, and when matters because an agent can keep running long after whoever built it stopped paying attention. Sources change, priorities change, instructions go stale, someone quietly grants access to a new system. Or the Monday update keeps arriving even though nobody has opened it since March. Ownership belongs in the design from the start.
I’d add one more question: what would have to happen for us to turn this off? Maybe the source data disappears, a platform changes, the team stops reading the output, or another workflow replaces it. Deciding that upfront stops you collecting agents purely because somebody once spent time building them.
With the job clear, the instructions almost write themselves. For the competitor example, something like this:
You are our Competitor Intelligence Agent. Your job is to monitor these competitors: [LIST]. Each week, identify meaningful changes in pricing, offers, positioning, campaigns, products and major announcements. Ignore routine content updates unless they signal a real change.
For every finding: explain what changed, include the source, explain why it may matter to our marketing team, and flag anything that needs human judgment. Do not make strategic decisions for us. Do not contact anyone, publish anything, or modify any business system.
Now the agent knows the job. More importantly, so do you. That sounds obvious, and yet plenty of agents get built without anyone being able to explain in one sentence which business process they’re supposed to improve.
If you cannot name the business process the agent improves, you are not ready to build it.
Here’s another part teams underestimate. You can write excellent instructions and still get generic work back, because the agent knows nothing about your business.
Think about what you’d give a marketer joining next Monday: brand guidelines, products, ICPs, personas, messaging, current priorities, past campaigns, examples of work you like and work you don’t, KPIs, terminology, constraints. Your agent needs the same foundation. I think of it as building the agent’s business brain. The instructions explain the job; the context explains how that job has to happen inside your organisation.
That second part is usually what separates a generic AI response from something your team can use. If you’ve ever read an output and thought this sounds fine, but it could have been written for any company, missing context is normally why.
Sometimes the problem isn’t the model. Sometimes the agent has simply been given a terrible first day at work.
This is where managing agents starts to feel different from prompting. You’re no longer only asking for an answer, you’re deciding how much authority something gets. Three levels are enough.
The agent gathers, analyses and reports. A human decides what happens next. Summarising campaign results, monitoring competitor changes, analysing feedback, finding patterns in sales-call notes. For most teams this is the sensible place to start, because the agent can be genuinely useful without taking a single external action.
The agent prepares something that may eventually leave the building, but a person approves it first. A campaign brief, ad variations, an email, a content calendar, a first version of a report. The agent handles the first pass; a person decides whether it moves forward.
Now it updates a system, sends something, schedules something, changes a record or triggers another workflow. This is where I’d slow down and think properly about controls.
Match oversight to consequence, not speed.
The closer an agent gets to customers, money, sensitive information, public communication or anything irreversible, the more human oversight I want. Fast does not mean low risk. An internal research summary can be wrong and cost you an hour. A public response can be wrong and cost you a relationship. Same technology, different consequence, so the supervision should differ too.
A lot of agent conversations start with “look what it can do.” I’m more interested in what happens if it gets this wrong. That question gives you a far better answer about how much rope to hand over.
Here’s a habit I wish more people built. When AI gives you something that looks good, don’t assume it’s finished. Turn the conversation around and ask the agent to inspect its own answer.
I use versions of these because a weak AI answer looks a great deal like a strong one. Both arrive neatly structured, both sound confident, both come with headings and clean recommendations. Our brains associate presentation quality with thinking quality, and those are different things. Asking the agent to interrogate its own output gives you one more chance to catch a shaky assumption before it travels any further.
There’s one piece of research anyone managing AI should know about. Researchers ran an experiment with 758 consultants at Boston Consulting Group. Across 18 knowledge tasks, the people using GPT-4 completed more work, faster, at better quality. Then the researchers handed everyone a task sitting just outside what the model does well, and the people using AI became 19% less likely to reach the correct answer than the people working without it[3].
Same people, same tool, similar-looking work, opposite outcome. They called it the jagged technological frontier, and I like the phrase because it captures something genuinely strange about this work: you can’t see where the boundary sits. AI may be excellent at Task A while Task B, which looks almost identical to you, goes badly. Nothing flashes red. The answer still sounds confident.
Which means “the agent has been good so far” isn’t a quality-control process. I’d pay extra attention when the task is new, the stakes are high, information is incomplete, the agent is clearly making assumptions, or the answer contradicts something you already know. And every so often, spot-check something you’d normally wave through. You want the checking built into the workflow, rather than added after the first bad week.
Once an agent is working, review it. And I wouldn’t make hours saved your only measure. Microsoft’s own guidance on this warns against treating usage as value and leaning too hard on theoretical time savings[4], which strikes me as right. If someone tells you an agent saves four hours a week, where did that number come from? Usually someone estimated it once and everybody kept repeating it.
I’d rather ask whether the agent changed something that mattered. Five questions, scored one to five, takes about as long as making a coffee.
| Question | Score |
|---|---|
| Is the output still accurate? | /5 |
| Does someone actually use the output? | /5 |
| Does the agent need minimal correction? | /5 |
| Has it changed a decision or an action? | /5 |
| Would we notice if we turned it off? | /5 |
| 20 to 25 keep it · 15 to 19 fix the workflow · below 15 ask whether it should still exist | |
Run this per agent, quarterly. Five minutes, and it settles the argument about whether something stays.
The scoring matters less than the conversation it forces. Most teams have never sat down and asked whether an agent is still pulling its weight, so the first time you run this you will usually find one that quietly stopped being useful in March.
That last question matters more than it sounds, because teams get strangely attached. Someone spent two weeks setting it up, so nobody wants to be the person who switches it off. But agents should earn their jobs too. If nobody reads the output, improve it or retire it. If someone rewrites the whole thing every time, find out why. If it keeps producing work and nobody makes a different decision because of it, ask what the process is actually creating.
An agent running is not the same as an agent creating value.
This is where things accelerate. Your first agent works, and suddenly you see possible agents everywhere. Before long four people have built four slightly different versions of the same thing. You need a little structure before that happens, and a shared document is enough.
| Agent | Job | Owner | Authority | Last reviewed |
|---|---|---|---|---|
| Competitor Watch | Monitor competitor changes | Demand gen | Research only | Aug 2026 |
| Voice of Customer | Analyse feedback | Product marketing | Research only | Aug 2026 |
| Campaign Brief | Prepare first drafts | Campaign lead | Draft only | Aug 2026 |
Three columns of discipline that stop four people building four versions of the same thing.
You don’t have to call it a centre of excellence if that phrase makes everyone tired. It’s a list. But knowing what exists, who owns it, what it can reach and when someone last looked at it becomes much more important once agent use spreads past one team.
If this all sounds useful and you’re still wondering where to start, here’s the sequence I’d use.
Monday, find the job. Write down three recurring tasks your team dislikes. Pick one, and don’t pick the most ambitious. Pick the one where you’ll learn something quickly.
Tuesday, write the job description. The business process, what good looks like, what it can access, what it must never do, who owns it, when it gets reviewed, and how you’ll know it worked.
Wednesday, build the business brain. Give it the context, examples of strong work, what your team cares about and what it ignores.
Thursday, run it manually. Don’t automate yet. Run the process alongside the agent and watch where it breaks. Maybe the instructions are vague, a source is missing, the output is bloated, or it keeps making the same wrong assumption. Good. You’re learning how the job actually needs to work.
Friday, review week one. What worked, what was wrong, what assumptions it made, what context would have helped. Then update it.
In week two, run it again and look for repeated problems. One strange answer is noise; the same failure three times is a design problem. By week four you only need to answer one question: has this agent earned a permanent job?
There’s going to be a lot of pressure to build agents simply because the technology makes them easy to create, and I’d resist it. You don’t win because you have twenty-seven agents. You win because the right work gets done better. Sometimes an agent handles an entire first pass, sometimes it researches and stops, sometimes it drafts and waits, and sometimes you look at a task and decide a person should own the whole thing. That’s part of the job too.
The skill isn’t building the most agents, it’s knowing what to delegate, what context to provide, where to put the boundaries, when to trust, when to check, when to override and when to switch the thing off.
So here’s the challenge. Don’t build five agents this week. Find one annoying recurring job and give a single agent one job, one owner, clear inputs, a defined output, clear boundaries, a review date and one measure of success. Run it for a few weeks, pay attention to what breaks and where you still need human judgment, then decide whether it stays.
That’s a decent first step. If you want the wider version of this, our guide on using AI for competitor research is a good first agent to copy, and building an AI prompt library for your team covers the shared-standards problem before it turns into one. Microsoft’s own write-up on the idea, featuring Copilot Studio’s Jack Rowbotham, is worth a read if you want their framing[2].
It’s Microsoft’s term for someone who builds, delegates to and manages AI agents rather than only prompting a chatbot. In day-to-day work the idea is simpler than the phrase suggests: you’re learning to manage a mix of human and software work. That means assigning the right jobs, providing enough context, setting boundaries, checking quality, and deciding which parts of the process still need a person’s judgment. Most of it is ordinary delegation, applied to something that doesn’t get bored and doesn’t tell you when it’s out of its depth.
Something repetitive and mildly annoying. Competitor monitoring, customer-feedback analysis, recurring research, campaign history, reporting and meeting preparation are all good places to look. Avoid starting with your most strategic or sensitive work: you’ll learn far more from a small, low-risk workflow, and the cost of getting it wrong is an hour rather than a relationship. A useful filter is whether AI can prepare the work while a human still makes the decision at the end.
No. Most current platforms let non-technical people create agents and recurring AI workflows without writing anything. The harder part is designing the work properly: what job does this agent have, what information does it need, how much authority should it get, and what happens when it gets something wrong. Those questions matter far more than whether you can code, and they’re the ones marketing teams are already equipped to answer.
Base it on consequence rather than on how fast the output arrived. An internal research summary needs less oversight than anything customer-facing. Anything touching sensitive information, money, public communication, strategy or an external action deserves real human involvement before it goes anywhere. The common mistake is treating everything as low-stakes because the agent produced it in four seconds, and speed tells you nothing at all about risk.
Look past usage counts and estimated time saved, both of which are easy to produce and hard to verify. Ask whether the agent improved a decision, caught something the team would have missed, raised the quality of the work, or made a process faster without creating a pile of correction work. And ask the question teams forget: would anyone notice if we turned it off? If the honest answer is no, you have your answer, and retiring it is a better outcome than leaving something running that people half-trust.