Ask any chatbot for ten headlines and you'll get ten variations of the same forgettable sentence. That's a briefing problem, and it's fixable in about four lines.
Most people brief AI for headlines in one sentence and then blame the model for the output. The fix is a five-part brief: audience, promise, proof, constraint, and examples of your own headlines that worked. Beyond that, this guide covers the headline patterns that hold up under real testing (including a 22,743-experiment study finding each additional negative word lifted click-through by 2.3%[8]), the character limits Google and the email platforms actually publish, and why testing matters more than any generator: 75% of links shared on Facebook are shared without anyone clicking through[6], so your headline is frequently the entire experience.
Let’s be honest about what happens when you type “give me 10 headlines for a blog post about email automation” into ChatGPT. You get ten headlines. Roughly seven of them contain a colon. At least two use the word “ultimate.” All of them would work equally well for any company in any industry, which means none of them work for yours.
People conclude from this that AI is bad at headlines. It isn’t. It’s doing exactly what you asked, which was to write a headline about a topic. A headline about a topic is always going to be generic, because a topic is generic. Good headlines aren’t about topics, they’re about a specific person getting a specific thing.
I’ve been writing and testing headlines for over a decade, across email, paid social, search ads and blog content, and the single biggest shift in output quality I’ve seen from AI came from changing the brief rather than changing the model. Same tool, same afternoon, dramatically better results.
It’s worth knowing how much rides on this. Nielsen Norman Group has been measuring this since 1997, when its early research put the share of web users who scan rather than read at 79%[13]. Its later eyetracking work, spanning studies over 13 years with more than 500 participants, found people “rarely read online, they’re far more likely to scan than read word for word,” a finding they describe as unchanged in 23 years[14]. On an average page visit, users read only 28% of the words[15]. NNG’s guidance is specific about where that attention lands: “Headlines are particularly important for fast communication, and the first few words are even more important, given users’ tendency to scan”[15].
And in social, the headline is often the whole thing. A 2024 study in Nature Human Behaviour analysing over 35 million public Facebook posts with URLs shared between 2017 and 2020 found that “shares without clicks” made up around 75% of forwarded links[6]. Three quarters of the time, nobody opened it. The headline was the article.
So yes, this is worth spending twenty minutes on properly.
Here is the entire fix, and it takes about four lines to write. Give the model five things it cannot guess.
The briefing structure recommended in this article. Author’s framework, not survey data.
Part five is the one nearly everybody skips and the one that does the most work. The model has no idea what your audience responds to. It has a general sense of what headlines look like on the internet, which is precisely why its output reads like the average of the internet. Three of your own winners give it a target to aim at.
“Audience: heads of marketing at UK B2B software companies with 20 to 200 staff, who are being asked to prove pipeline contribution and don’t have a data analyst. Promise: they can build a working attribution view in a week without buying anything new. Proof: we did this with a client and cut reporting time from two days a month to about forty minutes. Constraint: blog headline, aim for 55 to 65 characters so it doesn’t truncate badly in search. Examples of our headlines that performed well: [three, with their CTR]. Give me 20 options. Vary the structure, don’t put a colon in more than five of them, and don’t use the words ultimate, unlock, or guide.”
The banned-word list at the end is not a stylistic preference. Those words are the model’s defaults, and defaults are what make output recognisable as AI. Banning them explicitly forces it somewhere less worn.
One more thing on proof. If you don’t have a number, say so and give it something else concrete: a named customer, a specific tool, a real timeframe. “Faster reporting” is not a promise. “Attribution reporting in forty minutes a month” is.
There’s an enormous amount of headline folklore in marketing, most of it recycled from a Buzzfeed-era blog post that nobody has re-tested. A small amount of it has genuine, published evidence. Here’s what survives.
This is the best-evidenced finding available, and it comes from an unusually good dataset. The Upworthy Research Archive contains 32,487 randomised headline experiments, 150,817 experiment arms and over 538 million participant assignments, published as an open dataset in Scientific Data[7]. Researchers analysing a confirmatory subset of 22,743 tests, covering around 105,000 headline variations, 5.7 million clicks and more than 370 million impressions, found that “for a headline of average length, each additional negative word increased the click-through rate by 2.3%”[8].
Use this carefully. It’s a real effect, and a headline built entirely out of dread will earn the click and cost you the relationship. My rule: one negative word, pointed at the problem rather than the reader. “The reporting mistake that’s hiding your best channel” works. “You’re doing attribution wrong and it’s costing you” is the same mechanic turned on the person you’re trying to win over.
Google says this outright in its own advertising guidance, in the context of responsive search ads: “Avoid generic language in your ads: Use specific calls to action. Generic calls to action often show decreased engagement with ads”[2]. Google also reports that advertisers who improve Ad Strength from “Poor” to “Excellent” see 15% more clicks and conversions on average[2].
In practice specificity means numbers, timeframes, named tools and named roles. “Cut your reporting time” is a category. “Forty minutes a month” is a headline.
NNG’s guidance on how people read online recommends “placing information up front, in other words, front-loading,” alongside clear headings that break up content[14]. Google’s own documentation on title links makes the stakes clear: the title link “is often the primary piece of information people use to decide which result to click”[3].
Which means: put the thing you want them to notice in the first three or four words, not after a clever preamble. This one is easy to instruct. Add “the key benefit must appear in the first four words” to your brief and regenerate.
The most common mistake after a bad brief is asking for too few options and then editing the least-bad one to death.
Ask for 20 to 30. It costs you nothing, and the value of a batch is not the average quality, it’s the two or three genuinely unexpected angles buried in it. Then cut hard. My working ratio is thirty generated, three kept, and of those three, one usually survives contact with the team.
What I’m scanning for in the cut:
Then, and this is the step that actually improves the batch, feed your three back in and ask for ten more in that direction. Second-round output beats first-round output almost every time, because now the model has examples of what good means to you specifically rather than in general.
“These three are closest: [paste]. Here’s why each one works: [one line each, be specific about the mechanic]. Now give me 10 more in that direction. Don’t repeat the same sentence structure more than twice across the set.”
Honestly, most people stop after round one, which is why most AI headline output is mediocre. Round two takes ninety seconds.
Half the “rules” circulating about headline length are invented. Here’s what the platforms themselves actually publish, because if you put the real constraint in your brief, the model will respect it and you’ll stop rewriting truncated headlines.
| Channel | What the platform actually says | Source |
|---|---|---|
| Google responsive search ads | Headline fields support up to 30 characters. Minimum 3 headlines, up to 15 per ad. Descriptions up to 90 characters. | Google Ads Help |
| Page titles in search | No published character limit. Truncated “as needed, typically to fit the device width”. | Google Search Central |
| Email subject lines | Mailchimp recommends no more than 9 words and 60 characters, and no more than 3 punctuation marks. | Mailchimp |
| Email subject lines | Klaviyo’s data puts the average at about 7 words, associated with roughly 30% open rates. | Klaviyo |
| Email preview text | Support varies from about 278 characters down to none; Litmus recommends keeping it under 90. | Litmus |
Length guidance as published by each platform, cited in full in the sources below. Figures are the platforms’ own recommendations, not independent test results.
A few notes on using these properly.
The Google Ads 30-character headline limit is a hard technical cap, not advice[1]. If you’re briefing AI for search ads, say “maximum 30 characters, count them” and then actually check, because models are unreliable at counting and will hand you a 34-character headline with total confidence.
Mailchimp’s recommendation of no more than 9 words and 60 characters, with a maximum of three punctuation marks and one emoji, is their own published best-practice guidance[10]. Klaviyo’s data shows subject lines averaging about 7 words driving open rates around 30%, with longer 15 to 20 word lines sitting closer to 25%[11]. Two vendors, two datasets, same broad direction: shorter wins in the inbox.
Preview text is the forgotten half of the subject line. Litmus notes support ranges from roughly 278 characters down to zero depending on client and device, and recommends staying under 90[12]. Their advice for the AI-summary era is worth repeating: “Prioritize the subject line. Since AI-generated summaries may alter preview text display, focus on crafting subject lines that get your main message across, even if preview text is truncated”[12].
If email is where most of this lands for you, our ChatGPT prompts for email marketing covers subject lines in more depth, and using AI for email marketing covers the wider campaign workflow.
Everything up to this point produces better candidates. It does not tell you which one wins, and no model can, because it has never met your list.
Mailchimp puts this better than I could, in their own documentation: their subject line helper “uses data from all Mailchimp users,” while “A/B and Multivariate tests can tell you what your specific contacts like best”[10]. That distinction is the entire argument for testing in one sentence, from a company that sells you the AI feature.
What I’d actually run, in order of return:
On that last point, be realistic about what search data will tell you at the moment. Advanced Web Ranking’s Q1 2026 analysis found click behaviour has “fundamentally decoupled by device type,” with the top five desktop positions gaining a combined 10.54 percentage points quarter on quarter while mobile position-one CTR fell 2.20 points[18]. Their own caution is the right one to carry into any headline test: “your specific performance may vary significantly. There’s no one-size-fits-all benchmark”[18].
Keep a document of your winners with their numbers. That document becomes part five of the brief, and this is the compounding bit: every test makes your next batch of AI headlines better, because you’re feeding it a growing file of evidence about your specific audience rather than the internet’s average taste. For the broader testing habit, writing Facebook ad copy with AI and AI for landing page copy both cover variant generation at the ad and page level.
Three failure modes, all of which I’ve either done myself or watched a team do.
You let it write the article too. A headline is a promise, and the fastest way to burn a list is to make promises the piece doesn’t keep. Worth knowing that even heavy AI users mostly don’t do this: HubSpot’s research found only 4% of marketers use AI to write entire pieces of content, with the vast majority using it “for inspiration or to give them an outline and a few paragraphs to build on”[17]. The same research found 46% are only somewhat confident they’d know if AI output was inaccurate[17], which is a good reason to keep a human on the promise.
You scale it. Generating headlines is fine. Generating hundreds of pages to hang them on is a different activity with a different outcome. Google’s spam policies define scaled content abuse as generating many pages “for the primary purpose of manipulating search rankings and not helping users,” and explicitly lists “using generative AI tools or other similar tools to generate many pages without adding value for users” as an example[4]. Google’s guidance on generative AI content says it’s “particularly useful when researching a topic, and to add structure to original content,” while warning against volume without value, and specifically names metadata including title elements as an area where accuracy and quality matter[5].
You assume AI-written means worse, or better. The evidence here is more interesting than either camp expects. Research from NYU Stern examining generative AI in advertising found ads created entirely by generative AI increased click-through rates by up to 19% compared with ads made by human experts, that using AI to modify human-made ads showed no significant improvement, and, notably, that telling consumers an ad was AI-made reduced click-through by 31.5%[9]. That study looked at visual creative rather than headline copy, so don’t over-read it. But “AI-generated performs worse” is not a safe assumption, and neither is the reverse.
The wider adoption picture supports the boring conclusion. Content Marketing Institute’s 2026 B2B research, covering 1,015 B2B marketers, found 95% say their organisations use AI-powered applications and almost nine in ten already use AI to produce written content, while 12% say the quality of their content decreased with AI and around a fifth don’t see it moving the needle on creativity or quality[16]. Nearly everyone is using it. Results vary enormously, and the variable is how it’s used.
Pulling it together, here’s what this looks like on a Tuesday when you have a piece to publish and forty minutes.
Step seven is the one that separates people who get value out of this from people who keep having the same disappointing experience. The brief is an asset, and it appreciates.
If you want a wider view on where AI earns its place in a content operation, AI copywriting versus human writers covers the honest trade-offs, and writing a blog post with AI covers what happens after the headline is settled.
Because a one-line prompt gives the model nothing to work with except general internet patterns, so it returns the average of everything it has seen: colons, the word “ultimate”, and a promise vague enough to apply to any company. The fix is a five-part brief that supplies what the model cannot guess: who the audience actually is (job title, company size, current worry), the specific outcome, one real piece of proof, the channel’s actual character limit, and three of your own headlines that performed well with their numbers attached. Adding an explicit banned-word list also helps, because those default words are precisely what makes output recognisable as AI.
It depends entirely on the channel, and several of the numbers people quote are invented. Google responsive search ad headline fields have a hard cap of 30 characters and accept between 3 and 15 headlines per ad. For page titles in search results, Google states there is no character limit on the title element and that it truncates as needed to fit the device width, so the widely-repeated 60-character rule is a display convention rather than a Google policy; 55 to 65 characters is a sensible working target. For email, Mailchimp recommends no more than 9 words and 60 characters, and Klaviyo’s data shows subject lines averaging about 7 words performing well. Preview text should stay under about 90 characters.
There is genuine evidence for measured negativity, from an unusually strong dataset. Researchers analysing 22,743 randomised headline experiments from the Upworthy Research Archive, covering around 105,000 headline variations and 5.7 million clicks, found that for a headline of average length each additional negative word increased click-through by 2.3%. That is a real and replicated effect. The practical caveat is that it measures clicks, not trust, revenue or whether anyone comes back. A useful working rule is one negative word, aimed at the problem rather than at the reader: “the reporting mistake hiding your best channel” rather than telling your audience they are doing their job badly.
Not for using AI as such. Google’s published guidance says generative AI is “particularly useful when researching a topic, and to add structure to original content.” What Google targets is scale without value: its spam policies define scaled content abuse as generating many pages primarily to manipulate rankings rather than help users, and explicitly name using generative AI to produce many pages without adding value as an example. The policy language is notably about outcome rather than method, describing unoriginal content that provides little value “no matter how it’s created.” Using AI to draft twenty headline options for a genuinely useful article sits comfortably outside that. Auto-generating hundreds of thin pages does not.
The tool matters far less than the brief, and any of ChatGPT, Claude or Gemini will produce good headlines given good input and mediocre ones given a topic. Whatever you use, keep the winning-headline document alongside it and paste it in every time, because that file is what actually improves the output. One workflow note: models are unreliable at counting characters, so if you are writing for a hard-capped field like a Google responsive search ad headline, check the length yourself rather than trusting the model’s assurance that it stayed under 30. The bigger differentiator in practice is whether you run a second generation round after identifying your favourites, which most people skip.