A working system for quarterly OKR planning, built from what actually happens when AI meets a goal-setting process most teams already dread: real time saved on the first draft, and a specific new way to write goals that sound good and measure nothing.
Only 45% of teams say they’re effective at measuring their goals, and almost one in five rate themselves not effective at goal-setting at all, according to a Microsoft Research study of over 500 engineers (Butler, Zimmermann, and Bird, 2023). Just 47% of U.S. employees strongly agree they know what’s expected of them at work, a figure that’s climbed back to 49% in 2026 but still sits well below the 61% recorded in 2015 (Gallup). AI can close real parts of that gap: faster first drafts, better-structured key results, and check-in summaries nobody has to write by hand. It also drafts hollow key results that look rigorous and measure nothing, and honestly, it needs a human who already knows how to write a decent OKR to catch it.
Open it. I mean the real one, not the tidied-up version from the all-hands slide, still sitting in some tab you haven’t closed since last quarter. Count how many key results actually got graded above a 0.7. When I run this exercise live with a leadership team, the room usually goes quiet around the third one. Fewer than half, most of the time, and a chunk of the ones that cleared 0.7 only did because someone quietly rewrote the target in week nine so it would.
This isn’t really a you problem, even though it feels like one when you’re staring at the doc. Most teams write OKRs the same broken way: three tired people in a room in the last week of the quarter, pulling numbers that sound ambitious enough to say out loud but vague enough that nobody gets held to them precisely. Then everyone goes back to doing whatever they were already doing. The doc sits untouched until the next deadline forces someone to open it again.
AI changes a real part of that, and I think it’s worth being specific about which part. It’s genuinely good at turning a messy list of priorities into a structured first draft, and it’s decent at spotting when a key result is actually just an activity wearing an outcome’s clothes. It’ll also summarize a quarter’s worth of scattered check-ins into something a leadership team can actually read, which alone has saved me hours of prep before a quarterly review. What it can’t do is know what your team should actually be trying to achieve this quarter. That part is still entirely yours, and no prompt changes that.
The full system described in this guide, from first draft to ongoing check-ins.
Before you touch any tool, be honest with yourself about what actually happened to last quarter’s OKRs. Nobody looked at them after week one? Then a faster AI-assisted process just gets you a nicer-looking version of the same doc nobody opens again. Speed doesn’t fix an attention problem.
A Microsoft Research team studied OKR adoption across a software organization of more than 4,000 engineers: 47 interviews, a 512-person survey. I cite this one a lot in workshops because it’s peer-reviewed research, not a vendor survey with a product to sell. Tooling wasn’t the top complaint. Leadership buy-in wasn’t either, which surprised me the first time I read it. The plain act of writing the objectives and key results was the single biggest pain point: “creating and setting OKRs” came up almost twice as often as any other challenge. Only 45% of teams believed they were actually effective at measuring their goals, and nearly one in five said they weren’t effective at goal-setting at all.
That tracks with what I see when I’m sitting with a leadership team building their OKRs from scratch, which happens most quarters these days. Nobody struggles with the spreadsheet. What actually stalls the room is staring at a blank objective field, trying to compress six competing priorities into three to five statements that are both ambitious and honest about what’s achievable in twelve weeks.
Google’s own re:Work guide, which documents the OKR process the company has run internally for two decades, recommends keeping it tight: three to five objectives per level, roughly three key results under each, graded on a 0.0 to 1.0 scale where 60 to 70% is the actual target, not a disappointment. Full attainment on everything usually means the goals weren’t ambitious enough to begin with. Most teams I work with never get anywhere near that discipline on their own, and that’s the actual gap AI is useful for closing, but only if you use it to sharpen the writing instead of skipping the thinking part entirely.
The clarity problem shows up in the engagement numbers too. Gallup’s most recent workplace data puts the share of U.S. employees who strongly agree they know what’s expected of them at work at 49% as of mid-2026, an improvement from 47% the year before, but still nowhere near the 61% Gallup recorded back in 2015. A decade of investment in performance software and goal frameworks, and expectation clarity is still worse than it was ten years ago. That’s the actual baseline you’re working from. Not a hypothetical, not a worst-case scenario I made up to scare you into fixing it.
Self-rated goal-setting effectiveness among engineers surveyed at a 4,000-person organization. Source: Butler, Zimmermann, and Bird, Objectives and Key Results in Software Teams (arXiv, 2023). [5]
If your last OKR cycle ate a full week of leadership time and produced goals nobody could recite by February, a better tool isn’t the fix. Nobody pushed the objectives to be specific in the first place, and that’s a discipline problem, not a software one. AI can make the sharpening step faster once you decide to actually do it. It won’t decide for you.
Skip the blank page entirely, that’s the first thing I tell every group I train on this. Feed the model your actual inputs, not a request to invent goals from nothing: last quarter’s results, this quarter’s known constraints, the two or three things leadership has already said matter, and any customer or revenue data you have handy. A useful prompt looks like this: “Here are our last quarter’s OKR results, our team’s current priorities, and our headcount and budget constraints for this quarter. Draft three to five objectives that are ambitious but achievable in twelve weeks. Flag anything that sounds like a to-do list dressed up as a goal.”
That last instruction matters more than it looks. Google’s own guidance is blunt about the most common OKR-writing mistake: objectives that are really just a checklist of ongoing work, phrased to sound strategic. “Keep hiring,” “maintain market position,” and “continue doing X” are the classic offenders. AI is decent at catching this pattern once you ask it to, because it’s a language pattern, not a strategic judgment call.
Google’s re:Work guide describes a genuinely useful gut check: an objective should feel a little uncomfortable to commit to. Ask AI to generate two versions of each objective, a safe one and a stretch one, and sit with the difference for a minute before you pick. Most teams, left to their own devices, default to the safe version without ever noticing they did. I’ve watched this happen in real time in a workshop: a team quietly picks the safe draft, nobody flags it in the moment, and by week six the quarter feels suspiciously easy.
I still read every AI-drafted objective out loud before it goes in front of a team. A model can generate something that’s technically specific and still sounds like nobody in particular wrote it, and a stiff, generic-sounding objective doesn’t rally anyone. If you’re building this alongside a broader planning cadence, the same discipline applies to how you use AI to plan your week: AI drafts fast, but the version that ships has to sound like something a real person actually decided to commit to.
Never let AI invent your priorities from scratch. Feed it what you already know is true: last quarter’s numbers, this quarter’s constraints, the two things leadership already said out loud. Ask it to structure and sharpen that. Your judgment about what actually matters doesn’t get outsourced here.
This is where most OKRs quietly fall apart, AI-assisted or not. Google’s re:Work guide is specific about the tell: key results that use words like “consult,” “help,” “analyze,” or “participate” are describing activities, not outcomes. “Assess customer service satisfaction” is an activity. “Publish customer service satisfaction levels by March 7th” is a key result. The difference is whether an outside observer could look at the number on the last day of the quarter and know, unambiguously, whether you hit it.
A prompt that actually helps here: “For each objective, propose three key results. Each one needs a number, a deadline, and a description that a stranger could verify without asking me anything. Flag any of them that describe an activity instead of a measurable outcome, and tell me why.” Asking it to explain its own flags forces the model to justify the distinction instead of just rephrasing, which matters because rephrasing is the exact failure mode you’re trying to catch.
Quantive’s platform (formerly Gtmhub, now merged with Workboard) has a built-in feature for exactly this step, a guided draft flow inside its OKR whiteboard that proposes objectives and key results based on prompts you give it, then lets you refine them in place. Perdoo does something similar with its AI Assistant, which can draft strategic pillars, KPIs, and full OKRs, and suggest key results for an objective you’ve already written. Neither replaces the judgment call about whether a key result is actually meaningful. Both just save you the blank-page time.
If a key result reads well but you can’t picture the actual evidence you’d point to on the last day of the quarter, it’s not done. Doesn’t matter who wrote it, a person or a model. Same rule applies either way.
Harvard Business Review’s guidance on OKRs makes a point that a lot of companies quietly ignore: cascading goals works best team to team, not by forcing every individual to have a personal OKR that maps neatly to the company objective above them. Rigid, top-down cascading tends to produce exactly the kind of activity-dressed-as-outcome key results Google warns against, because individuals start reverse-engineering something measurable out of whatever they were already planning to do.
Where AI is actually useful here is checking alignment after the fact, not generating it from scratch. Paste your company-level OKRs and a team’s draft OKRs into the same prompt and ask: “Does this team’s key results meaningfully advance any of the company-level objectives? If not, say so plainly. If the connection is a stretch, say that too.” A blunt answer from a tool with no stake in the political comfort of the room is worth more here than another round of a manager softening their own assessment.
The mistake I see most often when I’m walking a leadership team through their first AI-assisted cascade: one department’s OKRs get written first, they’re specific and well-argued, and every other team quietly reshapes their own goals to echo that department’s language, whether or not it’s actually the right fit for their work. Honestly, it usually isn’t. Ask AI to check each team’s OKRs against the actual company objectives, not against each other’s drafts, to catch this before it compounds.
This is also where the case for AI-assisted OKR software over raw ChatGPT starts to matter. Lattice’s cascading alignment feature, for instance, lets you see a live tree of how every team’s OKRs connect (or don’t) to the objectives above them, which a chat window can’t show you at a glance. If you’re managing goals across more than a handful of teams, that visual matters more than any single AI-generated draft.
Full alignment on paper isn’t real alignment. Ask AI to flag the connections that are a stretch, not just rubber-stamp the ones that already look fine. Otherwise you cascade a false sense of coordination straight through the quarter, and nobody notices until the numbers don’t add up in week ten.
You don’t need dedicated OKR software to get real value from AI here. A small team can run this entire system through ChatGPT or Claude: draft objectives, generate key results, check alignment, summarize check-ins, all in one ongoing conversation thread you keep coming back to each quarter. The tradeoff is that nothing persists automatically. You’re the memory system.
Once you’re managing OKRs across more than a handful of teams, dedicated software starts earning its cost. Lattice, Perdoo, and Quantive all now build AI drafting and alignment-checking directly into the OKR-setting flow, and they keep a live, queryable record instead of a chat log you have to scroll back through. Perdoo’s AI Assistant drafts objectives, KPIs, and key results directly inside the tool your team already updates weekly. Quantive’s guided draft flow does the same inside its whiteboard view, with the advantage of pulling in live data from connected tools like Salesforce or Jira to ground a suggested key result in something real instead of a guess.
That last point isn’t a throwaway line. Microsoft built real AI capability into Viva Goals, Copilot features that could draft OKRs from planning documents and auto-summarize check-ins so a team lead didn’t have to write the update themselves. Then Microsoft retired the whole product on December 31, 2025, after quietly freezing new feature development a year earlier. If you were relying on it, you’re already migrating. I’ve had this exact conversation with more than one client this year, and it’s a useful reminder that a platform’s AI features are only as durable as the company’s commitment to the product wrapped around them. Worth weighing before you build a quarter’s worth of process on top of any single vendor.
Honestly, most teams overbuy on OKR software before they’ve proven they can run a disciplined quarterly cycle at all. Get the habit right with a free AI tool first. Pay for dedicated software once the habit is the bottleneck, not the writing.
Gallup’s research on performance management is specific about the cadence that actually works: employees who have quarterly progress conversations about their goals are 90% more likely to be engaged and 2.1 times as likely to feel the process is fair and transparent, compared with the far more common pattern of setting goals once a year and revisiting them only at review time. Fifty-six percent of employees currently review their performance goals with a manager once a year or less.
AI doesn’t decide whether a goal is on track here, that judgment still sits with a manager who knows the context. What it does well is remove the friction that makes weekly or biweekly check-ins feel like too much overhead to bother with. A prompt like “summarize this week’s check-in notes across all key results into a single paragraph per objective, flag anything that’s fallen more than one grading band behind pace” turns a fifteen-minute writing chore into something that takes two minutes to review and correct.
Copilot’s now-retired integration inside Viva Goals did exactly this: it read check-in data and status updates across an objective’s child items and generated a summary a manager could use to prep for a review meeting, instead of writing it from scratch. Perdoo’s check-in feature runs on the same principle without the AI layer doing the summarizing itself, just five minutes a week per person, aggregated automatically into a status view leadership can scan. Either approach beats what most teams actually do, which is nothing between the kickoff meeting and the quarter-end scramble. One honest caveat, though: a summarized check-in strips out tone. If someone’s status update reads “on track” but they said it through gritted teeth in the actual meeting, the AI summary won’t catch that. You still have to be in the room sometimes, or at least on the call.
If check-ins are new to your team, pair this with how you already run one-on-ones and difficult conversations: a five-minute OKR check-in fits naturally inside a regular one-on-one instead of requiring its own separate meeting nobody wants on the calendar.
Set the check-in cadence before the quarter starts, not after the first one gets skipped. A quarterly-only OKR process with no cadence in between is barely different from setting goals once a year and hoping.
Most guides to AI and OKRs skip this part, and it’s the one I’d actually worry about first. Large language models have a well-documented tendency researchers call sycophancy: a pull toward output that sounds agreeable, confident, and well-structured, whether or not it holds up under scrutiny. A 2024 technical survey on the topic notes the pattern shows up readily in structured writing tasks, not just open conversation. OKR drafting is about as structured as writing tasks get.
In practice, that looks like this: you ask AI to draft key results for “improve customer onboarding,” and it hands back something like “increase onboarding satisfaction score,” complete with a target number and a deadline, which looks exactly like a well-formed key result. Except nobody defined how satisfaction gets measured, whether the current tool even tracks it, or whether “increase” from an unknown baseline means anything at all. I’ve seen close to this exact key result, word for word, in more than one draft I’ve reviewed. It has the shape of rigor without the substance.
The fix is the same discipline Google’s own guide recommends for human-written OKRs, just applied deliberately to AI output: for every AI-drafted key result, ask “what specific evidence, existing today, would prove this number on the last day of the quarter?” If you can’t answer that in one sentence, the key result isn’t finished, no matter how confident it sounds. Push the model explicitly: “For each key result you just proposed, name the exact data source I’d pull the number from. If one doesn’t exist yet, say so instead of assuming it does.”
I run every AI-drafted OKR set through that question before it goes anywhere near a team, and it catches something almost every time. If you’re new to prompting for this kind of precision, the technique carries over directly from how AI can help you delegate work without losing visibility into whether it actually got done: specificity up front is what makes accountability possible later, whether the delegate is a person or a language model.
Before any AI-drafted OKR set goes in front of your team, run each key result through one question: what specific number, from what specific source, would prove this happened? If you can’t answer it in one sentence, send it back, no matter how polished it reads.
No, and treating it that way is the fastest way to end up with confident-sounding goals nobody can actually measure. AI is genuinely useful for structuring a first draft and catching activities disguised as outcomes. It’s also decent at checking whether team-level goals actually connect to company objectives, which is a check most teams skip. What it doesn’t know is your team’s real capacity, or which specific number would actually prove success this quarter. Those calls stay with you, and honestly, that’s the point of having a human run the process at all.
It depends on scale more than which tool is objectively better. A single team can run this whole system through ChatGPT or Claude, as an ongoing conversation thread, at no extra cost. Once you’re coordinating OKRs across multiple teams, dedicated software like Lattice, Perdoo, or Quantive starts earning its price. It keeps a persistent, queryable record and shows cascade alignment visually, something a chat log genuinely can’t do for you.
Microsoft retired Viva Goals on December 31, 2025, after quietly freezing new feature development a year earlier, citing low adoption. If your team was using its AI-assisted OKR drafting or Copilot check-in summaries, you should already be migrating. Microsoft made historical data available through API, Excel, and PowerPoint export before the shutdown, so pull yours now if you haven’t. I wouldn’t wait on this one.
For every AI-drafted key result, ask one question before you accept it: what specific number, from what specific data source, would prove this happened on the last day of the quarter? If you can’t answer that in one sentence, and neither can the model, the key result is describing an activity or a vague direction, not a measurable outcome. Push the AI to name a real data source. If one doesn’t exist yet, make it say so instead of guessing.
Google’s re:Work guide, which documents the framework the company has run internally for two decades, recommends three to five objectives per level with roughly three key results under each. More than that and teams get over-extended, effort gets diffused, nothing lands cleanly. If your AI-drafted list runs past five objectives, ask it to help you cut down to what actually matters this quarter. Not what would be nice to track. What matters.
I researched this by reading Google’s re:Work OKR guide directly, a peer-reviewed Microsoft Research study of OKR adoption across a 4,000-engineer organization (published on arXiv), and Harvard Business Review’s guidance on cascading goals. Product claims about Lattice, Perdoo, Quantive, and Microsoft Viva Goals are checked against each vendor’s own product documentation and support pages, including Microsoft’s official Viva Goals retirement notice. Workplace statistics are drawn from Gallup’s 2026 workplace research and its CHRO-focused performance management study. The point on AI sycophancy in structured writing tasks is grounded in a published 2024 technical survey, not a guess. Every stat and tool claim below is sourced and linked.