AI is genuinely good at reading the evidence you already own and telling you who actually buys. It is genuinely bad at inventing a customer from nothing, which is exactly what most people ask it for.
Prompt an AI assistant cold for an ideal customer profile and it will hand you a fluent, plausible, entirely fabricated composite of every ICP document it has ever read. There is published evidence for how badly that goes: researchers testing persona-conditioned language models against real World Values Survey respondents found persona prompting often degraded accuracy, with the best model matching real answers about 40% of the time against 27% for random guessing. Fed actual evidence, the same tools are excellent. This guide covers the six data sources you already own, four prompts that do real work (including the one that makes AI argue against your own profile), a campaign where the right-looking segment converted badly, how to validate a profile before you fund it, and the Ehrenberg-Bass research that says your ICP should not govern your reach.
Most ideal customer profiles are produced in a two-hour workshop by people who were already in the room. Somebody opens a template, somebody else says “I think our best customers are mid-market ops leaders,” and forty minutes later there is a slide with a stock photo, a first name, three bullets labelled Pains and a made-up quote in italics. It goes in the brand folder and nobody opens it again until the next workshop.
I have built that document, more than once, for clients who paid real money for it. It usually contains something true, because the people in the room have spoken to customers. You just cannot tell which parts are true and which are the most senior person’s hunch, because the output looks identical either way.
What has changed is that AI will now produce the same document in eleven seconds, in better prose, with more confidence, and no idea whether a word of it is true.
An ideal customer profile describes the kind of organisation worth spending money to pursue: size, sector, structure, situation, trigger. A buyer persona describes a person inside it. Merging the two is how you end up marketing to a fictional individual instead of a real account.
Forrester’s The State Of Business Buying, 2026 puts the typical business buying decision at 13 internal stakeholders plus nine external influencers, with procurement acting as a decision-maker in 53% of cycles.[2] Gartner found 99% of B2B purchases are driven by organisational change rather than by one person waking up wanting your product.[1] Forrester lists five buying group roles: champion, decision-maker, influencer, user and ratifier.[3] Marketing Mary is a champion. She is never the ratifier who kills the deal in week nine.
Try this, then close the tab without saving anything. Open any AI assistant and type: “Build an ideal customer profile for a B2B project management SaaS.”
You will get something that looks like research. Company size 50 to 500. Industries: technology, professional services, agencies. Pains: siloed communication, missed deadlines, tool sprawl. Triggers: rapid headcount growth, a failed competitor rollout. Structured, specific, immediately usable.
It is an average. The model has read thousands of ICP templates and category landing pages and is returning a composite of what your category says about itself: fluent, internally consistent, untethered from a single row of your revenue data.
The problem is not that those attributes are necessarily wrong. Plenty of them will be right. The problem is that you have no way to tell which ones describe your buyers and which ones merely describe the category, and the output gives you no signal either way.
There is now published work on precisely this failure mode. Researchers at IIT-CNR with the universities of Pisa and Florence tested whether conditioning a language model on a demographic persona makes it answer survey questions more like the real people it imitates, running two open-weight models plus a random-guesser baseline against United States microdata from the World Values Survey, across more than 70,000 respondent-item instances, published at the ACM Web Conference 2026.[5]
This is not an ICP study, and it would be sloppy to use it as one. What it is, is a warning about the underlying behaviour: giving a model demographic or persona cues does not reliably make it behave like the real people those cues describe. The finding is the honest kind. Persona prompting produced no clear aggregate improvement in alignment and in many cases significantly degraded it. The best persona-conditioned model matched real respondents’ actual answers about 40% of the time, against 27% for uniform random guessing.[5] Better than chance, nowhere near a substitute for asking a person.
What I keep returning to is what sits underneath that average. The distortion was not spread evenly: most questions barely moved, while a small subset of items and several underrepresented subgroups took disproportionate error.[5] Some underrepresented groups showed disproportionately large errors, and the model does not tell you when that is happening. If you are about to synthesise a persona for a segment you do not currently sell to, that is the sentence to sit with.
Most established teams I have worked with already hold more ICP evidence than they realise. It sits in six places, none of which talk to each other, and none of which are pleasant to read.
That last source has stopped being optional. G2’s 2026 Buyer Behavior Report, based on a survey of more than 1,000 B2B software buyers plus interviews with over 50 sales and marketing leaders, found review sites were the top source shaping which vendors make a shortlist, cited by 38% of buyers and just ahead of AI chatbots at 37%.[4]
Gong puts the call-recording case plainly: in B2B sales, the truth about deals and customer needs lives in conversations rather than in CRM fields.[10] If you have a year of recordings and have never asked what closed-won buyers said in the first ten minutes that closed-lost buyers did not, that is your highest-value afternoon.
| Question | Opinion ICP says | Evidence ICP says |
|---|---|---|
| Who is it? | A named persona with a stock photo | An account type, plus the count of real accounts that match |
| Where from? | A workshop, or a cold AI prompt | Closed-won, closed-lost, churn, tickets, calls, reviews |
| The trigger? | Generic pain points | An event you can observe and date in your own records |
| Who else decides? | Nobody else is named | Champion, decision-maker, influencer, user, ratifier |
| How is it wrong? | No stated way to find out | Written falsification tests, re-run each quarter |
The practical difference between the two documents. Buying group roles are Forrester's five-role framework.[3]
The usual objection is that the data is too messy to bother with. It is messy. Salesforce’s seventh State of Sales report, a double-anonymous survey of 4,050 sales professionals across 22 countries fielded in August and September 2025, found 74% working on data cleansing and 51% of sales leaders with AI saying disconnected systems were slowing them down.[7] HubSpot’s 2026 State of Marketing, drawing on more than 1,500 global marketers, found just 65% say they have high-quality audience data, unchanged year on year.[6]
Messy evidence still beats clean fiction. If feedback data is your messiest part, our guide on analysing customer feedback with AI covers getting structure out of tickets, and AI customer segmentation covers splitting a profile without inventing the splits.
These are the four I use, in order, and they are deliberately unglamorous. Before running any of them, replace personal names and email addresses with account IDs, and use a tool where your organisation has a contract and training on your data is off.
Export closed-won opportunities from the last 24 months: account ID, industry, employee count, contract value, sales cycle length, lead source, use case, reason-won. Then:
“Here are our closed-won deals from the last 24 months. Do not summarise them. Find recurring attributes in this set. For each one, give me the count and what share of the total it represents. Show both, and flag any pattern based on fewer than 10 records as low-confidence so I can inspect it manually. Do not infer that an attribute predicts winning until I give you a comparison set.”
The counts clause is the whole trick. Without it you get adjectives. With it you get “this appears in 41 of 180 deals” and can see which findings are real and which are two or three deals that happened to rhyme.
I used to phrase that first instruction as “more often than you would expect by chance,” which sounds more rigorous than the workflow actually is. The model is not running a significance test when you say that, it is pattern-matching and then agreeing with your framing. Worse, a closed-won set on its own cannot tell you what predicts winning, because you are only looking at winners. If 60% of your wins are mid-market, that means nothing until you know what share of your losses were too. That is what the last line of the prompt is holding back, and it is why the next prompt is the one that does the real work.
Ten is not a magic number either. Ten matching records out of fifteen accounts is a strong signal; ten out of ten thousand is nothing. The flag is there to make you look, not to make the decision for you.
“Here are closed-won and closed-lost deals in the same format. What is present in the won set and absent in the lost set? Where the two look identical on firmographics, what else differs? List anything appearing in both sets at similar rates, because those attributes are useless for targeting.”
That final sentence produces the most valuable output of the exercise. Half of what sits in a typical ICP describes your whole market rather than your winnable slice of it.
“These are first-30-day support tickets from accounts that later renewed, and from accounts that later churned. What problem does each group describe in their own words? Quote them directly. Do not paraphrase and do not tidy up the grammar.”
The no-paraphrasing instruction matters because paraphrase is where the model swaps your customer’s words for category words. You want the phrase your buyer typed at 11pm, not a tidy version that sounds like your website.
“Here is the ICP we have drafted from that data. Argue against it. What in the evidence I gave you contradicts it? What would have to be true for this profile to be wrong, and what specific check would tell me?”
This one earns its keep and almost nobody runs it. Models often follow the framing you give them, so if you only ever ask one to strengthen an ICP, it can end up reinforcing the assumptions already inside it. Asking for the counter-case gets you falsification tests in about ninety seconds.
One caveat on all four: the model does pattern description rather than causal inference. It can tell you your winners skew towards 200 to 400 employees, but not whether that is because they need you more or because that is where your two best reps have networks.
| Prompt | What the model does | What you still have to decide |
|---|---|---|
| 1. Closed-won pattern | Finds recurring attributes and returns counts and shares | Whether the sample is big enough, and whether the attribute is something you can actually target on |
| 2. Won and lost contrast | Finds what differs between the two sets | Whether the difference reflects your buyers or just reflects where your reps prospect and how you price |
| 3. Language | Extracts and quotes customer wording without tidying it | Whether those quotes are representative or just the most vivid ones |
| 4. Disconfirmation | Generates counterarguments and proposes tests | Which of those tests is worth the money and time to actually run |
The split across all four prompts. The right-hand column is the job, and it is the part that does not get faster.
A few years ago I ran an account-based programme for a B2B SaaS client whose ICP had come out of exactly the workshop I described at the top. Mid-market logistics firms, 200 to 1,000 employees, target buyer the operations manager. It felt right to everyone including me, because their three biggest logos were logistics companies and so were all the case studies.
We built a named list of about 180 accounts, wrote logistics-specific content, ran ads against the list, briefed the SDRs on logistics language. Meetings came in roughly at the rate we had forecast, which is why nobody panicked for four months. Then almost nothing closed. Deals reached a second call and evaporated, and the reason-lost field filled with variations of “no decision.”
What broke it open was one afternoon reading the closed-won records properly, including the small ones the team had ignored because the contract values were unimpressive. The pattern was not industry. Those companies had, in the previous nine months, hired their first dedicated operations analyst. Some were logistics. Several were manufacturing, one was a hospital group, two were e-commerce.
The firmographics were not the useful differentiator. The trigger was, and it had been sitting in our own data the whole time. The rebuilt list was smaller and looked wrong on a slide, a scatter of industries with nothing tidy in common. It converted. Not spectacularly, but more than the tidy list ever did.
A related habit: defining your best customer as your loudest customer. Advisory boards, the people who answer every survey, the accounts your CSM has a relationship with. They are wonderful and they are not a sample. I have watched whole campaign narratives get built around eight enthusiastic customers who were a rounding error of revenue, while a much larger quiet group described a different problem in tickets nobody read.
So ask yourself: when you say “our best customers want X,” can you name the number behind that sentence? If it is a feeling rather than a figure, that is an opinion ICP. The fastest way to find out which one you have is to put a small paid budget behind the profile and watch what happens, and our guide to running LinkedIn ads with AI covers how to set that test up targeting-first.
| Question | What a good answer looks like |
|---|---|
| Source. Which real records support this attribute? | Named: closed-won opportunities, first-30-day tickets, call recordings. Not “we all know that” |
| Count. How many accounts show it? | A number and a share, with the size of the set it came from |
| Contrast. Is it more common in won accounts than lost ones? | Two rates you can compare, not one rate on its own |
| Trigger. What observable event tends to come before demand? | Something a person could go and look up, like a new role appearing or a funding round |
| Counterexample. Which won accounts do not fit? | You can name them, and you know why they are exceptions |
| Test. What would prove this profile wrong? | A specific check with a number attached, agreed before you run it |
Run this before you call the document an ICP. Six questions, about twenty minutes. The two rows that get left blank are Contrast and Test, and those are the two that decide whether the profile is worth funding.
An attribute you cannot trace back to a record is a belief. Keep it if you like, but label it.
The gap between a document and a decision is validation, and it takes about a week.
The validation sequence described in this section. Every step runs on records you already hold.
Step three finds the problem. Send the draft to two or three reps and ask one question: name three deals we won last year that this profile would have told you not to bother with. If they can, your profile is too narrow.
If they cannot think of any, resist the obvious conclusion. That does not prove the profile is right, and it does not prove it is too broad either. It only tells you the profile does not currently contradict anything your reps remember. You still have to test whether it excludes enough of the low-probability market to be worth having, which is what step four is for.
Before you run that test, write down the signal you are looking for. Positive response rate, qualified meeting rate, progression to a second conversation, or whatever else matches how your sales motion actually works. Pick one, and pick the number that would count as a pass. Do it beforehand, because “it worked” decided afterwards is not a result, it is a mood.
Step four matters because the buying process has got harder to get through even when the targeting is right. G2 found evaluation is now the longest stage of the buying journey, surpassing research for the first time, with IT security review the biggest source of delay at 39% of buyers. Nearly half said a CFO vetoed an approved deal in the past year.[4] A thirty-account test tells you whether your profile is wrong before you spend a quarter’s budget finding out.
You do not need enrichment software for the first pass. Start with the records you already own: a CSV export and a general-purpose assistant will get you most of the way. Once the profile survives a small test, then decide which missing attributes are worth paying to enrich.
Here is the tension, and I would rather put it in front of you than pretend it away. The whole logic of an ICP is narrowing, and there is an important counterargument from B2B marketing science: narrowing too aggressively can limit growth.
Professor Jenni Romaniuk of the Ehrenberg-Bass Institute, in a study for the LinkedIn B2B Institute, found narrow targeting counter-productive and argued the best way to drive B2B growth is to target all customers in your category, because the duplication of purchase law holds in B2B: you share and acquire customers from other brands in proportion to their size rather than their similarity to you.[8] The B2B Institute’s Jon Lombardo put it bluntly in the same piece, arguing segmentation has been pushed down to the individual with no evidence that smaller and smaller works.[8]
Then the companion finding from Professor John Dawes: companies change providers of services like banking, legal, software or telecoms roughly every five years, so only around 20% are in the market in a given year and about 5% in a given quarter. Advertising works mainly by building memory links that get retrieved when they finally are.[9]
So how do you hold both? My honest position, after running campaigns on both sides of this argument and losing money on each: an ICP should govern where you spend effort, not the ceiling on your reach.
I used to think that split was a fudge. I no longer do, because the two decisions have different constraints and time horizons. I watched a team take an ICP built for sales prioritisation and use it to set audience exclusions on every ad campaign in the account. Pipeline looked beautifully efficient for two quarters, then thinned out, and by the time anyone noticed, the fix was twelve months away.
Four failure patterns, roughly in order of how often I see them.
Asking the model to imagine rather than to read. A prompt beginning “create an ICP for…” with nothing attached is asking for fiction, whereas “here is our data, find…” is asking for analysis. The interface makes them look like the same activity.
Treating a generated profile as a finished decision. The output looks complete, which is exactly the problem. If your sales team has not tried to break it, it has not been tested.
Building it around the people who answer your emails. Advisory boards, survey respondents, your three friendliest accounts. Enthusiastic customers are a biased sample by construction, and the quiet majority is usually describing a different problem somewhere you are not reading.
Freezing it. An ICP built in 2024 describes a market that has since had its buying process rearranged. G2 found finance involvement in software decisions jumped from 31% to 46% in a single year.[4] Forrester found more than 60% of business buyers now use a trial to de-risk a purchase, rising to 78% at $10 million or more.[2] If your profile mentions no ratifier and no security review, it was written for a different market.
Something worth saying plainly, because the prompts are the easy part. What teams get stuck on is not the wording. It is deciding which of their data is trustworthy enough to build on, separating a correlation from something they can actually target, agreeing in advance on what would prove the profile wrong, and then getting sales and marketing to make the same decisions off the back of it. None of that is a prompting problem, and none of it gets solved by a better model.
That is the kind of work we build into hands-on AI training: not “ask AI for an ICP,” but learning how to use AI to interrogate evidence you already own, without handing over the judgement part.
The thing to take from all of this: the value sits less in the document than in what building it from evidence tells you, which is which of your beliefs about your customers were ever true. That list is always shorter than you expect.
An ideal customer profile describes the type of organisation worth pursuing: size, sector, structure, situation and trigger. A buyer persona describes a person inside it. You need both, but they answer different questions. Forrester puts the typical business buying decision at 13 internal stakeholders plus nine external influencers, so a profile built around one imagined individual is describing a small slice of the real decision.
It will produce one, and it will look convincing, but it will be a composite of ICP documents and category marketing pages rather than a description of your customers. Research presented at the ACM Web Conference 2026 found persona-conditioned language models matched real survey respondents about 40% of the time against 27% for random guessing, and that persona prompting often degraded accuracy rather than improving it. Use AI to read your evidence, not to invent it.
Closed-won opportunities from the last 18 to 24 months, with the reason-won field, then closed-lost in the same format so the model can contrast them. After that, churned accounts, first-30-day support tickets, sales call notes and review site language. Strip personal names and use account IDs, and run it in a tool where training on your data is switched off.
Narrow enough to direct finite sales effort, and no narrower. Ehrenberg-Bass research for the LinkedIn B2B Institute found narrow targeting counter-productive to B2B growth, and that only around 5% of businesses are in the market for a given service in any quarter. Use the profile to prioritise named accounts, seller time and custom content. Do not use it to set exclusions on brand campaigns aimed at buyers who will enter the market later.
Quarterly, against new closed-won data, and write the date next to each assumption. Buying conditions move fast: G2 found finance involvement in software decisions rose from 31% to 46% in a single year and evaluation has overtaken research as the longest stage of the buying journey. A profile that has not been re-checked in a year is describing a market that has already changed shape.
Every figure here was checked against its original source in August 2026, and several tempting ones were dropped. Widely quoted ABM benchmark statistics were excluded because the underlying Momentum ITSMA study sits behind a client login and could not be verified. Apollo pricing was left out because the published plan table did not render on the vendor page, and vendor prices are quoted only where the vendor's own pricing page loaded. The campaign described in this article is drawn from Hina's own client work and is told without invented figures deliberately: the numbers that matter in it were never the point. This is general marketing guidance, not a substitute for reading your own data.