Every company seems to be building one of these now. Most of them are built wrong, and I keep watching the same mistakes repeat.
An AI Center of Excellence works when it has a named owner, a real mandate to approve or block AI use cases, and metrics tied to business outcomes, not just tool logins. It fails when it’s a task force nobody remembers to disband, a Slack channel dressed up as governance, or a compliance layer with veto power and no forward motion. McKinsey’s most recent State of AI survey found 88% of organizations now use AI in at least one business function, but only 39% report any enterprise-wide EBIT impact from it, and most of those say it’s under 5%. That gap between using AI everywhere and it actually moving the numbers is exactly what a well-structured CoE is supposed to close. This guide covers the three operating models (centralized, federated, hub-and-spoke), who needs to be in the room, the three KPIs worth tracking from day one, why governance keeps failing at the exact same rate Deloitte just measured, and a rough 90-day plan if this landed on your desk.
Two things happened in the back half of 2025 that explain why “AI Center of Excellence” turned into one of the more searched phrases in corporate IT and HR departments this year. Gartner published a report in November titled Transform Your AI COE Into a Strategic Value Enabler, and its own summary is blunt about the problem: plenty of enterprises already launched an AI CoE to build initial technical expertise and pilot use cases, but as adoption grows, few have evolved that CoE’s mandate to match where the company actually is now[1]. A few weeks before that, Deloitte announced the launch of its own global AI Infrastructure Center of Excellence, which tells you something when the firm selling AI advice is standing up a formal structure instead of just writing about one[4].
I sat in on a CoE kickoff call for a client earlier this year, and about ninety minutes in, nobody in the room could actually tell me who had the authority to say no to a use case once the pilot phase ended. That’s not a knock on the people on that call. It’s the default state of a structure that gets built because a board member asked for one, not because anyone worked out what job it’s supposed to do.
Here’s the gap that makes this worth fixing. McKinsey’s most recent State of AI survey found 88% of organizations now use AI in at least one business function, up from 78% a year earlier. But only 39% report any enterprise-wide EBIT impact from it, and most of those say the impact is under 5%[2]. We covered a version of this same gap when 48% of executives told researchers AI had been a disappointment (see our breakdown of what the other 52% are doing differently), and the pattern is consistent: usage is nearly universal, value capture is not.
A Center of Excellence exists to close that gap. Skip the structure, and adoption plateaus exactly where Gartner describes it: a lot of pilots running in parallel, and nobody with the standing to turn any of them into something bigger.
An AI Center of Excellence, CoE for short, is a small, named group inside your company with the authority to standardize how AI actually gets used: which tools are approved, how a new use case gets vetted before it ships, who signs off when something touches customer data, and how you measure whether any of it worked. It sits in an awkward middle spot between a task force and a permanent department, and that in-between status is exactly why so many of these fail before they get started.
A task force disbands once its charter runs out. A department gets a permanent seat on the org chart and reports up through someone with real authority. A CoE starts out with neither of those things, which is normal, not a design flaw, but it can’t stay that way forever. Sitting in that in-between spot too long is what turns it into a group that just meets. It produces slide decks, it runs a few workshops, and six months later nobody outside the group can name a single decision it actually made.
If you want the plain-English fundamentals of what enablement means before this piece, we already covered that ground in our AI enablement explainer. This piece assumes you have the basics and picks up where that one leaves off: the org-design part. And the org-design part comes down to one word, mandate. A CoE with a real mandate can say yes to a use case and make it happen, or say no and have that no actually stick when a department head pushes back. Without that authority sitting somewhere specific, what you’ve built is a working group with a nicer name on the org chart.
One team owns every AI decision company-wide: which tools get approved, which use cases get built, who gets access. This is the fastest model to stand up and the easiest to govern, because there’s exactly one place decisions get made. The tradeoff shows up fast at any real scale: every team that wants to do anything with AI has to wait on one small group, and that group becomes a bottleneck the moment more than a handful of use cases are in flight at once.
Each business unit builds and governs its own AI use, with a thin central layer offering guidance rather than approval. This moves fast and lets teams closest to the actual work make the calls. It also means five departments can quietly buy five different tools that do the same thing, with five different data-handling standards, and nobody notices until a security review surfaces it. Federated works when your culture already has strong cross-team communication. It’s a mess when it doesn’t.
A central hub sets the standards, the approved tool list, the governance rules, and the metrics. Business units, the spokes, own execution of their own use cases inside those guardrails. This is the model most mid-size companies actually end up with, whether they planned it or backed into it, because it keeps one place accountable for standards without making every single decision a bottleneck. If you’re starting from nothing, I’d start here rather than trying to centralize everything on day one and loosening it later. It’s a lot easier to hand out more autonomy than to claw it back once teams have gotten used to going around you.
The people who need a seat at the CoE table are fewer than most kickoff decks suggest, and the ones who matter most are rarely the ones who get invited first.
Someone senior enough that when a department head pushes back on a CoE decision, the CoE’s answer holds. Without this person, every disagreement becomes a negotiation the CoE loses, because nobody below the C-suite can out-rank a VP who wants an exception. This is the single role that determines whether the rest of the structure means anything.
One person runs this day to day, and their job title should say so. “AI lead” or “head of AI enablement” works. A committee chair who rotates quarterly does not, because nobody builds institutional memory or takes real accountability for a rotating seat.
Pull people from the business units where AI is actually getting used, not just from IT. A CoE staffed entirely by technologists tends to approve technically sound use cases that nobody in the actual department wants, and reject practical ones for reasons that only matter on a whiteboard.
Bringing risk and compliance in after a use case is already built is how good ideas die in review. Bring them in when the use case is still a proposal, and they become collaborators instead of gatekeepers.
This is the piece that turns a CoE from a policy document into something people actually feel day to day. Citi built exactly this: a network of roughly 4,000 AI Champions and Accelerators spread across a workforce of 182,000, and reached over 70% adoption of firm-approved tools as a result[6]. What stands out to me about Citi’s version is what it didn’t include. No promotions attached to the role, no pay bump, just training access, internal badges, and standing with your own team. Peer-led adoption spread faster than a top-down mandate would have, according to the reporting on it, and honestly that tracks with what I see in the workshops I run. People trust a colleague who’s already tried the thing over a slide from corporate every single time.
A CoE without metrics is a group of well-meaning people having meetings. The metrics are what turn opinions about whether it’s working into something you can actually defend in front of a skeptical VP.
CIO.com, in a piece sponsored by CrowdStrike, laid out three KPIs that hold up well beyond the security context they were written for: time from idea to production deployment, employee adoption rates of approved tools, and security incidents caught through prevention[5]. The insight that matters most here isn’t any single number, it’s that these three have to move together. Deployment speed going up while incidents also go up means your controls have gaps. High adoption with slow deployment means you’ve built a bottleneck. Low incidents paired with low adoption usually means you’ve blocked so much that nobody’s using anything, which isn’t safety, it’s just a different kind of failure.
Layer a fourth number on top of those three, because otherwise you’re only measuring activity, not value: some version of business impact tied to the specific use cases your CoE approved. This is where that 39% EBIT figure from McKinsey becomes useful as a gut check rather than just a scary headline[2]. If your CoE can’t point to which of its approved use cases moved a real number (hours saved, error rate down, cycle time shorter), you’re measuring whether people logged in, not whether the thing worked.
Start with a handful of metrics you can actually explain out loud to a skeptical CFO. A dashboard with forty tracked numbers and no clear story is worse than three numbers everyone in the room understands and can argue about.
Deloitte’s 2026 State of AI in the Enterprise report found that only one in five companies has a mature model for governance of autonomous AI agents, even as agentic AI usage is set to rise sharply over the next two years[3]. That gap between how fast the technology is moving and how ready the guardrails are is precisely the job description of a CoE’s governance function, and it’s also exactly where most of them get the balance wrong.
Deloitte’s own framing on this is worth sitting with: organizations where senior leadership actively shapes AI governance see meaningfully more business value than those who hand the whole job to a technical team and walk away[3]. Effective governance gets built into how people are actually evaluated at review time, not bolted on as a separate approval gate that exists parallel to how work really gets done.
The honest tension here, and I’ll admit this took me longer to land on than I expected, is that governance built too loosely lets shadow AI run wild, and governance built too tightly turns your CoE into the group everyone routes around. I don’t think there’s a formula that solves this cleanly. What I’ve seen work is a tiered approval process: low-risk, internal-only use cases (drafting, summarizing, brainstorming) get a fast, almost automatic green light, while anything touching customer data, financial decisions, or external-facing output gets the fuller review. Treat every request the same way and you’ll either approve something risky out of review fatigue, or you’ll bottleneck the boring, obviously-fine stuff behind the same process built for the risky stuff.
For a broader look at where AI adoption is actually landing across industries right now, Microsoft’s diffusion research is worth a read alongside this (see what Microsoft’s AI Diffusion Report reveals about global adoption), since governance maturity tends to track pretty closely with how far along a company already is on adoption overall.
A committee with a rotating chair produces committee output: cautious, unaccountable, and slow. Pick one person, give them a real title, and make their name the answer to “who do I ask.”
A CoE staffed entirely by data scientists and engineers builds technically elegant solutions nobody downstream asked for. The functional reps aren’t optional extras, they’re the reason anything gets adopted instead of just built.
The flashy autonomous-agent demo gets the board’s attention. The unglamorous, repetitive task that eats forty hours a month across a department is usually where the actual return sits. A CoE that only greenlights the exciting stuff will have a great slide deck and no measurable impact.
If every proposal takes six weeks of review regardless of risk level, people stop proposing things through the CoE and just find a way around it. That’s exactly how shadow AI takes root, quietly, inside the very structure meant to prevent it.
This is the exact failure Gartner called out in the report that opened this piece: plenty of companies stood up a CoE to run early pilots, and then never came back to update its job description once those pilots turned into real, everyday usage[1]. The structure that was right for approving three cautious pilots a year ago is usually the wrong structure for the fifty requests coming in now. Set a recurring check, every two quarters is reasonable, to ask plainly whether the CoE’s mandate still matches how much AI the company is actually running. If nobody’s asked that question in a year, that’s the answer.
Audit what already exists. Map every AI pilot, every tool someone quietly expensed, every shadow AI habit nobody officially approved. Start the executive sponsor conversation now, don’t wait for month two. Loop in someone from legal or compliance from the start. Sketch a one-page draft governance framework (what’s fast-track, what needs full review), good enough to test, not perfect.
Confirm your executive sponsor and pick your operating model (hub-and-spoke, most likely). Run one or two governed use cases through the draft framework from month one, with a metric attached from day one, not bolted on after. Whatever breaks in that first real use case tells you more than another week of planning would.
Formalize the governance process based on what actually broke or held up in the use cases you just ran, not on what you guessed on day one. Recruit your first wave of champions. Bring a board-ready update built entirely on the real numbers from the last sixty days.
If this landed on your desk with a vague mandate and a deadline that feels unreasonable, that’s normal, and I won’t pretend the first month feels organized. It won’t. You’ll spend more time than you expect just figuring out what AI is already being used, quietly, without anyone’s blessing, before you can govern anything new.
Notice the order on purpose: the draft governance framework comes before the first live use case, and the formalized version comes after. Write the polished policy document before you’ve tested it against one real case, and you’ll formalize the wrong things, the parts that felt important on a whiteboard instead of the parts that actually broke when a real request came in. The executive sponsor conversation starts on day one for the same reason. That’s usually the slowest-moving piece, since it depends on someone else’s calendar and appetite, not yours, so it can’t wait until you’re otherwise ready.
By day 30, you should have an honest map: what’s actually running, who’s using what, where the real appetite for AI sits inside the company, which is often not where the org chart suggests it should be, and a sponsor conversation already underway. By day 60, your model is picked, your sponsor is confirmed, and something small and real is live with a metric attached, running through a rough governance draft rather than no governance at all.
By day 90, you want three things you can put in front of leadership without hedging: a governance process that’s actually being followed because it was shaped by a real use case, not just written down and never tested, a champions network with real people in it, even a small one, and two or three use cases with a number attached to what they saved or improved. If you personally want a faster way to apply AI to your own workload while you’re standing all this up, our guide on using AI to be a better manager is worth a look, since the person building the CoE is usually also the person with the least time to spare for it.
Nobody gets this fully right in ninety days. The goal isn’t a finished structure, it’s a structure with enough real decisions behind it that it survives the first time someone senior asks what it’s actually done.
It’s a small, named group inside a company with the actual authority to decide which AI tools get approved, how new use cases get vetted before they ship, and how success gets measured. The word that matters most is authority. Without a real mandate to say yes or no and have it stick, what you have is a working group, not a Center of Excellence.
You need the function, even if you don’t need the full-time headcount. A smaller company can run a lightweight version, one named owner spending part of their time on it, a short approved-tools list, and a simple tiered review process, rather than the multi-person structure a large bank like Citi runs. The mistake smaller companies make isn’t skipping a CoE, it’s skipping the mandate part and assuming informal habits will scale, which they don’t past a certain headcount.
Neither one alone tends to work well. A CoE that reports purely into IT ends up technically sound and business-disconnected. One that reports purely into a business unit tends to lose the cross-company standardization that’s the whole point of having one central structure. The pattern that holds up best is a business-side executive sponsor with a technical co-lead, so decisions get made with both the authority to enforce them and the technical judgment to make them sound.
A champions network is the peer layer, employees across teams who’ve gone deeper on AI and help their colleagues day to day, the way Citi’s 4,000-person Champions and Accelerators program works. A CoE is the structural layer above that: the governance, the approved-tools list, the metrics, and the mandate. A good CoE usually runs a champions network as one of its tools, but the champions network by itself has no authority to approve or block anything.
Expect two or three governed use cases live with a real metric attached inside the first 90 days if things go reasonably well. Enterprise-wide financial impact takes considerably longer. McKinsey’s data shows only 39% of organizations report any enterprise-wide EBIT impact from AI at all, and most of those say it’s under 5%, so treat the 90-day milestone as proof the structure works, not as proof of company-wide ROI. That part takes quarters, not weeks.
This piece draws on Gartner’s November 2025 research on transforming AI Centers of Excellence, McKinsey’s State of AI in 2025 survey, Deloitte’s State of AI in the Enterprise 2026 report and its own AI Infrastructure CoE launch announcement, CIO.com and CrowdStrike’s KPI framework for AI-enabled success, and reporting on Citi’s AI Champions and Accelerators program. All sources are linked below and were checked live while writing this.