Every AI marketing post promises another way to generate something. The workflows actually worth stealing this year do the opposite: they point AI at work you already made and ask it to find what's wrong with it.
The AI marketing workflows worth trying this week are diagnostic, not generative: point AI at a site, an ad library, a review corpus, or a call transcript archive you already have, and ask it one specific question instead of asking it to write something new. Six workflows here, each sourced to a named practitioner, each rated honestly for evidence strength, each flagged for whether the real version needs code (most don’t). The counterweight section covers what breaks this pattern, including a Meta ad account permanently banned for giving an AI agent write access, and an AI ad-targeting product the FTC found was reselling email lists at a markup.
Scroll LinkedIn on any given Tuesday and you’ll find someone promising the same thing: an AI prompt that writes your ad copy, your email sequence, your content calendar. I used to stop and read these. Now I scroll past most of them, not because they’re wrong, but because they’ve stopped being interesting. Everyone’s already tried using AI to make something. The novelty wore off months ago.
What hasn’t worn off is a different pattern entirely. The AI marketing uses that made me stop and think “I didn’t know you could do that” almost never asked AI to create anything. They pointed it at something that already existed, a live website, a brand’s own ad library, a pile of call transcripts, and asked one specific question. Diagnostic, not generative. Auditing, not authoring.
That distinction isn’t new to this site. We’ve made a version of the same argument about Copilot: strongest for reading, comparing, and challenging work you already have, not generating new work from nothing, covered in our piece on the Copilot features almost nobody uses. It carries into marketing more cleanly than I expected.
Point AI at work that already exists, not a blank page. The blank page is where AI marketing gets boring and generic. The work you already made is where it gets useful.
Here’s the same task, done both ways, so the difference is concrete rather than a slogan:
| Task | Generative version | Diagnostic version |
|---|---|---|
| Landing page | “Write a new landing page for this product.” | “Here’s our page and our persona. Find the highest-impact missed opportunities and tell me what to change.” |
| Customer research | “Write a buyer persona for our ideal customer.” | “Here are real sales calls and reviews. Extract the customer’s language, then rate this idea against five statements, hard no to perfect fit.” |
| Site content | “Write 30 new product pages for our catalog.” | “Here are the 30,000 pages we have. Flag any claim about pricing or results that isn’t backed up on the page.” |
Same tool in every row. What changes is whether you’re asking it to make something new, or find what’s wrong with what you made. This article is built from real examples of the right column.
One honest note: this isn’t an argument that generative AI marketing is bad. A first draft of ad copy is still a legitimate use, and we’ve covered plenty of those in our guide to using ChatGPT for marketing. This piece is about the diagnostic workflows that don’t get written about as often, because they’re less flashy and harder to demo in thirty seconds. This year, they’re also where the real surprise is.
Say your site has thirty thousand pages built over a decade by a dozen people, and somewhere in there is a pricing claim two years out of date, or a “we’re the only provider who” that stopped being true. Nobody finds that clicking through by hand.
Aimee Peake at Workshop Digital, a Richmond, Virginia agency, hit a version of this at enterprise scale: a client with ten websites and no reliable way to tell which pages were public and which had drifted into staff-only content[1]. Her fix used a tool most marketers already own and have only used for one thing: Screaming Frog, best known as a broken-link checker, also ships an AI connector that runs an arbitrary prompt against every page it crawls, writing the answer, category plus reasoning, into a spreadsheet. Swap Peake’s classification prompt for something else, tone of voice, unsupported claims, stale pricing, and it becomes a bulk auditor instead of a link checker.
The setup is a handful of settings, not code:
One line made the whole thing usable. Peake’s first prompt kept mislabeling pages because of navigation menus and login buttons appearing identically on every page, “combination content,” her team called it. The fix was one added sentence: “Ignore universal navigation, login buttons, or links common to all pages.” Without it, the audit drowns in boilerplate false positives.
Model category: your chosen LLM connector (OpenAI or Gemini). Content type: HTML. Prompt target: Page Text. Test on 10 to 20 pages before running the full crawl.
Adapted from the audience-classification prompt and navigation fix documented by Aimee Peake[1], rewritten for an unsupported-claims audit.
Who this is for: whoever owns the website, not necessarily whoever runs campaigns. Fully non-technical, the only friction is pasting an API key into a settings box.
The honest part: Screaming Frog’s license runs about £199 a year, and Gemini’s free tier was capped near 1,500 requests a day when this was written, so a 30,000-page crawl needs a paid key or a crawl split across days. Verify both prices, they move fast. Peake’s outcome claim, “time savings measured in hundreds of hours,” has no baseline attached. The mechanism is verifiable. The ROI figure isn’t. Our anti-hallucination toolkit helps decide what actually counts as unsupported once the sheet comes back.
Andy Crestodina, co-founder and CMO of Orbit Media, makes a point that’s easy to agree with and hard to act on: because of confirmation bias, marketers often “use AI” just to validate what they already believed[2]. You ask if your headline is good. It says yes. Nothing changes.
His fix is a specific prompt structure, not “be more critical.” Give the model a full-page screenshot, a B2B persona, and an explicit list of fourteen named cognitive biases (scarcity, loss aversion, anchoring, and eleven others), and ask it to find the highest-impact missed opportunities to apply each one, with drop-in rewrites tagged by bias[2]. His example: a program starting in autumn, with no deadline mentioned anywhere on the page. A scarcity-focused pass catches that gap in seconds, a specific lever, not a vague “be more persuasive” note.
Inputs: a full-page screenshot of the page and a written buyer persona. Always use a “Thinking” model for audit prompts like this one.
Condensed from the verbatim prompt published by Andy Crestodina at Orbit Media[2].
The part worth copying most is the follow-up, the actual “argue with your marketing” move. After the bias pass returns recommendations, run one more prompt forcing the opposite conclusion: “Argue why this page should NOT be changed,” or “What would make this idea fail?” A direct check against the model telling you what you want to hear.
A related instinct: Dara Denney, an independent paid-social strategist, uses Claude Cowork with brands like Ridge[3]. She builds a persona from a brand’s actual reviews, separately builds the persona implied by its own live Meta ad library, and treats the gap between the two as the finding. She also runs a scheduled weekly Slack prompt asking Claude for blunt “do less of” recommendations on her own content. Her reaction to one batch: “you’re reading me to filth.” Worth disclosing: Denney runs her own agency, so a workflow that makes her look sharp is also good marketing for her services.
Both versions are fully non-technical, a screenshot and a persona for Crestodina’s audit, a Claude subscription for Denney’s. Neither comes with outcome data. Treat it as a real technique that surfaces real gaps, not proof that acting on it moves a number.
Ask an AI model to rate anything on a 1-to-5 scale and it will, reliably, cluster almost everything around a 3. Not laziness. A bare scale with no anchors gives it nothing to disagree with, so it hedges toward the middle, and the score becomes useless.
Kieran Flanagan, SVP of Marketing at HubSpot, built a fix specific enough to copy, documented in Maja Voje’s 2026 State of AI for GTM report[4]. He runs a customer “digital twin” in a Claude Project, in three layers:
When he tests a campaign idea against the twin, the model writes a natural-sounding response first, then matches it to the nearest anchor statement. You get the score and the reasoning, instead of a bare number with nothing to argue with.
Illustrative example for a fictional AI meeting-notes tool sold to mid-market operations teams, written to show the shape, not a real customer’s words.
| Position | Anchor statement, in customer language |
|---|---|
| 1. Hard no | “We already have someone taking notes. Not worth the switching cost.” |
| 2. Skeptical | “I’ve been burned by ‘AI notetaker’ tools before, they miss half the action items and nobody trusts the summary enough to skip the recording.” |
| 3. Undecided | “I’d need to see it handle a messy, multi-topic meeting before I’d trust it on a client call. Show me, don’t tell me.” |
| 4. Interested | “If it catches the decisions we made and not just a transcript, that solves a real problem. I’d want a trial on our actual meetings, not a demo.” |
| 5. Perfect fit | “This is exactly the gap. We lose action items every week because nobody wants to take notes. Where do I sign up.” |
Modeled on the anchor-statement structure documented for Kieran Flanagan’s digital-twin workflow[4], written by Future Factors as a worked example, not a real customer transcript.
If you can’t describe what a 1 sounds like and what a 5 sounds like, in the customer’s own words, don’t ask AI to rate anything on that scale.
Flanagan’s own caveat is worth keeping exactly as he said it, the honest limit of the method: “This system gives you accuracy on purchase intent, but only if you feed it real customer data. Synthetic data in, synthetic insights out.” If your “real customer data” is actually two reviews and a guess, the twin will confidently produce a number anyway, and it will be wrong in a way that looks convincing.
This is fully non-technical: a Claude Project, transcripts you already have, five sentences you write yourself. Be precise about what’s confirmed here: high-confidence that Flanagan runs this workflow, low-confidence that it’s independently proven, there’s no outside validation study, only his own qualitative account. Whoever owns positioning or messaging testing is the natural fit, more than whoever owns day-to-day execution.
Most sales call transcripts get recorded, filed, and never opened again. Jonathan Kvarfordt, at Momentum, treats them as the highest-signal material a marketing team owns, and built a NotebookLM system to use them, also documented in Maja Voje’s GTM report[4]. Worth disclosing: Momentum sells the call-recording product that produces the transcript corpus this depends on. The mechanics are still worth stealing regardless.
The specific move: upload instructions as source documents NotebookLM reads first, rather than retyping them each time, a master document cascading to the others, a brand-standards document with a pronunciation guide (Audio Overview kept mangling the company’s name), and an analysis protocol comparing customer words against internal words, every claim citing a transcript. Load twenty to thirty transcripts and it produces overviews, decks, or reports.
The detail worth copying alone: never delete the notebook. Add transcripts to the same one every quarter, and you watch which objections faded and which showed up, compounding instead of resetting. His stated saving, two to three hours versus fifteen-plus of manual review, is real but self-reported.
A related, separate habit: before a deliverable ships, check it against everything the client has already told you, not just the brief. The pain is exact: someone writes the post, having never seen the email from three weeks ago asking for a specific wording change. The Monday version needs no code, just a Claude Project.
What to feed it: the last 90 days of client email and Slack exports relevant to this account, plus the current brand guidelines document, loaded into a dedicated Claude Project for that client.
A Future Factors synthesis built for this article, in the spirit of a pattern that surfaces often in agency discussions of internal AI tooling: check new work against what the client already said, not just the brief.
Whoever owns client delivery, an agency account lead, a fractional marketer running several accounts, gets the most from this one. Neither this nor Kvarfordt’s version needs a developer. Both need the source material organized somewhere the model can read it, usually the part people skip.
Every workflow above shares a boundary worth stating plainly: read access is where these examples stay safe. Write access is where they stop being diagnostic.
In early-to-mid 2026, a wave of Meta ad account bans hit marketers who’d connected autonomous AI agents, usually Claude Code, directly to Meta’s Marketing API. Blend AI, which builds a reviewed, rate-limited Meta Ads connector, documented the pattern, including the most-shared account, a Reddit post titled “Claude Code got my Meta ads account permanently banned”[5]. The poster’s own explanation: “claude code was hammering the API too fast and tripped their fraud detection. the automated budget changes looked exactly like bot activity to meta’s system and the AI-generated creatives being published without human review violates their ad policies.” Dozens reported the same pattern, some losing years of campaign history. Worth naming: Blend sells the safer alternative, so their account also sells their product, though the mechanism they describe, high-volume API calls with no human approval on writes, matches what caused the bans regardless.
Read-only is where AI marketing is currently safe. The moment an agent gets write access, budgets, published creative, live audience edits, you’ve left diagnostic for a different risk category. See our fit test for when not to use AI at work.
The counterweight isn’t only about runaway automation. In May 2026, the FTC made Cox Media Group and two smaller firms pay a combined $930,000 over an AI product called “Active Listening,” marketed as targeting ads from conversations picked up by smart-device microphones[6]. The FTC’s complaint is blunt: “this service did not, in fact, listen in on consumers’ conversations or use voice data at all… the service the companies provided consisted of reselling, at a significant markup, email lists obtained from other data brokers.” A real, sold product, and the AI claim behind it was nothing of the sort. Clicking through a mandatory terms-of-service screen, the FTC noted, isn’t opt-in for something this invasive.
The third failure mode isn’t malicious, just what happens when a diagnostic system quietly turns fully generative and nobody’s watching. Content agency Animalz built an AI system for their LinkedIn service and, once it worked, asked for full autopilot, raw materials in, finished drafts out, no human touch[7]. Each post looked fine alone. Zoomed out, patterns showed: topics repeating, hooks templatized, tone drifting. Their fix was reintroducing friction on purpose, redesigning drafts to arrive “less like finished posts, more like working documents, with pointers, checks, and decisions for the team to make”[7], the diagnostic instinct, turned on their own system.
Three cautionary tales, one lesson: AI marketing claims and AI marketing automation both need a check before you trust them.
| Question to ask | Why it matters |
|---|---|
| Is there an actual number, self-reported or confirmed? | Most workflows here have a described method and no outcome data. Fine if you know that going in, a problem if “here’s how I did it” gets read as proof. |
| Does the source sell the thing that makes their workflow possible? | Not disqualifying alone. Momentum sells call recording, Blend sells a safer Meta connector, agency case studies sell agency services. |
| Is the figure about this workflow, or borrowed from elsewhere? | Numbers migrate. A stat from an unrelated study can end up glued to a different method just because the topics are adjacent. |
| Does the real build need code the story isn’t mentioning? | Several workflows here needed a terminal or a development background under the friendly description. Ask what the non-technical version looks like first. |
Built from the evidence gaps found while re-verifying the sources in this article, not a generic checklist.
Six workflows is too many for Monday. Pick one, based on what you actually own, not what sounds most impressive in a meeting.
Own the website, run the claims audit on one product line, not all 30,000 pages. Own the ad account, try the ad-library-versus-reviews check, or the bias audit if conversion is closer to your job. Own positioning, build five anchor statements for one real decision you’re unsure about. Own client delivery, load ninety days of one client’s history into a Project and run the pre-send check on the next deliverable before it ships.
Drew Bredvick, who leads GTM engineering at Vercel, has a useful filter: “anytime you hear ‘should’ in GTM, think extra hard, that’s probably a great spot for AI,” because the “we should analyze why we’re losing deals” list never happens, no time, no owner[8]. If you’ve been meaning to audit your site’s claims for months, that’s the signal.
Whichever you pick, be explicit about the split between what the model does and what you still decide, rather than letting “AI flagged it” quietly become “it’s fixed.”
| What AI does | What you still own | How it gets checked |
|---|---|---|
| Flags claims that lack visible backup, with reasoning for each flag | Deciding whether a flag is a real problem, a known one-off, or a false positive | A named owner spot-checks flagged pages against the live site before any fix ships |
| Matches a campaign idea to the nearest anchor statement, with reasoning | Deciding whether to trust that match, and whether the anchors still reflect real customers | Anchors get revisited whenever new transcripts or reviews come in, not left static |
The split changes by workflow. What doesn’t change is that something stays a human decision and someone actually checks the output.
Measuring whether it was worth keeping means going past whether you logged in once. Use: did you actually run it, rather than trying it once and going back to the old way. Persistence: are you still running it a month later without a reminder. Impact: did the site, campaign, or deliverable actually change. A lot of AI marketing experiments stop at the first of those three and call it a win. It isn’t one yet.
Run one workflow, on one real piece of work, this week. Check back in a month against use, persistence, impact, and decide honestly whether it earns a second run or was more interesting to read about than to operate. For teams still stuck at “we know we should be using AI for something,” our piece on marketers being told to use AI without being taught how is the right next stop.
The unusual, verified uses point AI at work that already exists rather than creating something new: auditing a live site with Screaming Frog’s AI connector, diffing a brand’s ad library against its review corpus to find who it thinks it’s talking to versus who’s buying, building a calibrated persona with five reference statements instead of a numeric scale, and turning call transcripts into a compounding NotebookLM archive. All six are covered above with the named practitioner and exact prompt.
Most of them, yes. The Screaming Frog audit needs a settings menu and an API key, no code. The anchor-statement persona needs a Claude Project and five sentences you write yourself. The NotebookLM cascade needs documents, not code. A couple of the practitioners built more advanced versions involving a terminal or AWS, and this article deliberately describes the non-technical version instead, flagging it where the two diverge.
Creating asks AI to produce something from a blank page: new ad copy, a persona invented from assumptions. Auditing asks it to examine something that already exists and answer a falsifiable question: does this page contain an unsupported claim, does our targeting match who actually buys, does this deliverable contradict something a client already told us. The generative uses have gotten crowded. The diagnostic ones are where this research found the real surprises.
Ground it in real data, actual call transcripts and objection notes, not invented assumptions, since a persona built on guesses agrees with whatever guess you feed it next. And replace numeric rating scales with five anchor statements written in real customer language, spanning genuine skepticism to enthusiasm, so the model has something concrete to compare a response against instead of defaulting to a safe middle score.
The mechanisms are usually real and checkable, you can verify Screaming Frog’s AI connector works as described. The outcome numbers attached are almost always self-reported, rarely with an independent baseline. Worth asking, for any claim: is there an actual figure, does the source sell the tool involved, and does the real version secretly need code the story isn’t mentioning.
This article draws on six verified practitioner sources (Aimee Peake at Workshop Digital, Andy Crestodina at Orbit Media, Dara Denney, Kieran Flanagan at HubSpot and Jonathan Kvarfordt at Momentum via Maja Voje’s GTM Strategist report, and blend-ai.com’s documentation of the 2026 Meta ad-ban pattern) plus the FTC’s May 2026 press release on the Cox Media Group settlement and Animalz’s own account of their AI content system, all fetched and re-verified directly during this writing session on September 8, 2026. Two candidates from the original research brief were cut: a synthetic-survey-panel workflow (PR-specific, evidence assertion-only) and a client-folder-of-markdown-files pattern sourced to Reddit, whose thread could not be re-opened this session. A third, the pre-send-check-against-client-history workflow, is presented as a Future Factors synthesis rather than a verified case study, because the Reddit thread it was originally drawn from was inaccessible in this session (Reddit is blocked by this environment’s browser policy and the specific thread did not surface through search) and the brief’s own standard is to drop rather than cite anything that cannot be reopened. The disputed statistic sometimes attached to the anchor-statements workflow, a 90% accuracy figure that actually belongs to an unrelated PyMC Labs and Colgate-Palmolive study, does not appear anywhere in this article.