Home / AI in Marketing / A Practical Guide to Usin...

AI in Marketing

A Practical Guide to Using AI for Customer Research and Personas

August 25, 2026 · 11 min read
Two persona profiles side by side — one assembled from stacked documents of real customer evidence, the other formed from empty generated shapes with no source beneath it

There are two activities being sold under the same name right now, and they produce opposite results. One is analysis. The other is fabrication. Both involve typing into the same box.

The distinction everything depends on

Ask a model to describe your target customer and it will produce something. It will have a name, an age, a job title, three frustrations and a favourite app. It will read exactly like the personas in every marketing textbook, because that is what it was assembled from — an average of everything ever written about people who might resemble your customers.

Now upload two hundred sales call transcripts and ask the same model what objections come up most often, in what words, at what stage. That output is grounded in something. It can be checked. It can be wrong in ways you can detect.

Same tool, same interface, completely different epistemic status. The debate about whether AI personas are good or bad mostly collapses once you separate these two cases — and almost every guide on the subject picks one and pretends the other doesn't exist.

The whole test AI is excellent at finding patterns in customer evidence you already have. It is worthless at inventing customer evidence you don't have. Which one you're doing is the only question that matters.

Researchers have been fairly direct about the second case. A review of dozens of synthetic persona studies published across leading AI venues found weak ecological validity across much of the field — the personas frequently failed to reflect real populations, real interactions, or real domain data. Practitioner guidance from usability research bodies is consistent too: synthetic outputs can be useful for forming a hypothesis to test, and should not be presented as findings about real users.

There's a subtler failure mode worth naming as well. Models are trained to be agreeable, which means a simulated customer tends to describe the person they think they ought to be. Ask a synthetic persona how often they compare prices before buying and you'll get a diligent, rational answer. Real customers are considerably less impressive, and the gap between those two pictures is precisely where most marketing goes wrong.

What to feed it

The useful version of this work starts with evidence you already own and mostly aren't reading. Most organisations are sitting on far more customer language than they realise.

Sources ranked by how much they reveal, and what each is good for.
Source What it gives you Watch out for
Sales call recordings Objections, comparisons, exact vocabulary, decision process Skews to people who took a call — not your whole market
Support tickets and chat logs Where the product fails, in the customer's own framing Over-represents problems; absence of complaint isn't satisfaction
Reviews (yours and competitors') Purchase triggers, unmet needs, why people chose someone else Skews to extremes — the delighted and the furious
Win-loss notes The actual decision criteria, including the ones you'd rather not hear Usually recorded by the person who lost the deal
Cancellation reasons What stopped being worth paying for Stated reason often isn't the real one
Survey free-text Unprompted language and priorities Only the motivated respond
Site search and query data What people expected to find and didn't Tells you the what, never the why

Notice that every high-value source is qualitative and unstructured. That's exactly the material that used to sit unanalysed because reading four hundred support tickets was nobody's job. It's also the material language models handle genuinely well. This is the real unlock — not persona generation, but making a mountain of customer language readable for the first time.

The competitor review angle deserves particular attention. Reviews of the products your prospects considered instead of yours are public, plentiful, and full of people explaining exactly what they wanted and didn't get. Very few teams mine them systematically, and the language people use there feeds directly into keyword research as well as messaging.

Before you upload anything. Strip names, email addresses, phone numbers, account IDs and anything else that identifies an individual — most of this analysis works fine on anonymised text. Check your organisation's policy on which customer data may be processed by third-party tools, and whether your terms of service with customers permit it. If you're in a regulated sector or handling EU data, that check is not optional and should happen before the first upload rather than after.

The four jobs AI does well here

Not "make me a persona." Four narrower tasks, each of which produces something checkable.

1. Coding at scale

Qualitative coding — reading transcripts and tagging recurring themes — is slow, tedious, and the reason most customer research stops at twelve interviews. A model can do a first pass over hundreds of documents in minutes, returning themes with frequency counts and source references.

Treat that first pass as a draft by a fast, literal-minded junior researcher. It will catch things you'd have missed through fatigue and it will also confidently merge two themes that only look similar. You review and correct; you don't accept.

Featured Recommendation AD · AFFILIATE
Originality AI logo
4.0 / 5.0

Originality AI

Detects AI-generated content ensuring your blog posts pass Google's quality guidelines

Best for: AI Content Checker

2. Language extraction

The highest-value output and the most underused. Ask for the exact phrases customers use to describe the problem, sorted by frequency, with no paraphrasing. What comes back is the vocabulary your market actually uses, which is almost never the vocabulary your website uses.

This feeds messaging, ad copy, landing pages and search targeting simultaneously. It's also the one output that's hard to fabricate convincingly, because you can spot-check any phrase against the source.

3. Contradiction hunting

Point the model at what customers say they want and at what they actually do, and ask where the two disagree. Stated preferences and revealed behaviour diverge constantly, and that divergence is usually the most commercially useful thing in the dataset. Humans tend to smooth it over; a model asked directly for conflicts will surface it.

4. Interview preparation and synthesis

Before: generate question sets, identify what previous rounds failed to ask, flag leading questions in your draft. After: transcribe, summarise, compare across sessions. The interview itself stays human — that's where the unexpected happens — but the work either side of it compresses dramatically. Setting this up as a standing process rather than a one-off is straightforward if you already have an AI workflow that genuinely saves time, and it benefits more than most tasks from careful prompting — vague instructions here produce confident nonsense rather than obvious errors.

Building the persona

Once the analysis exists, the persona is assembly work. Four rules keep it honest.

Every claim carries a source. If the persona says price sensitivity is high, that line should trace to specific transcripts. Anything untraceable gets deleted, not softened into a hedge. This single rule eliminates most of what makes AI personas dangerous.

Keep the disagreements. The standard persona format smooths a population into one tidy individual, and models are unusually good at that smoothing. If a third of your customers behave differently, that's a second persona or a stated split — not an averaged compromise that describes nobody.

That smoothing problem has a commercial cost worth spelling out. If your data contains a price-driven segment and a support-driven segment, an averaged persona describes a moderately price-conscious customer who moderately values support — and that person doesn't exist. Every message written for them lands weakly with both real groups. This is also why persona work and funnel work belong together: the segments that matter are usually the ones that behave differently at a specific stage, which shows up clearly when you're auditing the funnel end to end.

Cut the decoration. Names, stock photos, invented hobbies, a favourite coffee order. None of it comes from data and all of it makes the fiction feel more solid than it is. The most useful persona documents in practice look like evidence summaries, not character sheets.

Write down what you don't know. A section listing open questions is more valuable than a section of confident detail, because it tells the next researcher where to point. It also prevents the persona being read as complete.

Validation, which is where teams stop

Almost every guide in this category ends at persona creation. That's the point at which the risk actually begins, because a well-formatted persona is persuasive regardless of whether it's true.

Three checks, in order, and the third is the one that gets skipped.

Traceability. Walk the document line by line and confirm each claim maps to source material. Expect to delete a meaningful proportion. That deletion is the process working, not failing.

Falsification. Ask what evidence would show this persona is wrong, then go looking for it specifically. Models will happily confirm whatever framing you bring them, so you have to introduce the challenge deliberately. Run the same analysis with an opposing prompt and see whether the conclusion survives.

Human contact. Take the persona to five or six actual customers and ask them where it's wrong about them. Not whether it's right — people are polite — but where it's wrong. This takes a week and catches the failures the first two checks can't, because a fabricated detail can be internally consistent and still describe nobody who exists.

If your organisation can't fund those six conversations, that's worth knowing before you build a year of messaging on the output. It's also a reasonable point at which outside research and content support costs less than the campaigns you'd otherwise aim at the wrong person.

What to do with it afterwards

A persona that sits in a shared drive has produced nothing. The output only matters where it changes a decision, and there are three places it reliably should.

Messaging and page copy. The extracted language goes straight into headlines, subheads and body copy — replacing the internal vocabulary that almost every organisation drifts into. This is the fastest measurable return from the whole exercise, and it's most visible on landing pages, where the gap between your words and your customer's words shows up directly in conversion rate.

Segmentation you actually act on. If the research surfaced two genuinely different groups, that split should reach your email programme rather than staying in a slide. Personalisation built on real behavioural segments performs very differently from personalisation built on demographic guesses — a distinction that matters more as AI reshapes email personalisation.

Knowing when it's stale. Personas rot. Markets move, competitors change, and in B2B the composition of buying committees has shifted noticeably — how those committees now weigh peer evidence against vendor claims is a good example of a change that invalidates persona assumptions written two years ago. Set a review date when you create the document. Six months is reasonable; twelve is the outer limit.

Where AI genuinely can't help

Four situations where the tool is the wrong instrument, regardless of how good your data is.

  • New markets and new products. No existing customer data means nothing to analyse. A model asked to imagine a market you haven't entered returns the internet's average assumptions about it, dressed as research.
  • Emotional and sensitive territory. Anything involving fear, shame, health, money stress or identity. The material that matters here surfaces through trust and silence in a real conversation, and does not appear in a transcript summary.
  • Genuinely novel behaviour. Models describe patterns from the past. Customers doing something new — a workaround, an unexpected use case — read as noise to be smoothed away rather than the signal they are.
  • High-stakes irreversible decisions. Pricing architecture, positioning, a product bet. Use AI to generate the hypothesis; use real research to make the call. Low-stakes and reversible is where synthetic exploration belongs.

A fifth caution, less about capability than about output quality: personas built the lazy way converge. If three competitors all ask the same model to describe the same market, they receive near-identical answers and produce near-identical messaging — the same dynamic behind the content sameness problem. Your own customer data is the thing competitors cannot replicate. Analysis grounded in it is differentiating by construction; analysis grounded in training data is differentiating by accident, which is to say not at all.

A workable first project

If you want to start this week, the smallest useful version:

  1. Gather one source type. Sales calls or support tickets. Six months' worth. Anonymised.
  2. Extract language before themes. Ask for the exact phrases people use, by frequency, unparaphrased. This produces value immediately and is trivially checkable.
  3. Code for themes, then verify a sample. Pull ten source documents at random and confirm the coding held up. If it didn't, your instructions were too vague.
  4. Hunt contradictions. Ask explicitly where stated wants diverge from observed behaviour.
  5. Draft one persona, sourced throughout. One, not five. Delete every unsourceable line.
  6. Take it to six customers. Ask where it's wrong. Revise.

That's a fortnight of work rather than an afternoon, and it produces something you can defend in a meeting. The afternoon version produces something that sounds identical and can't survive a single well-informed question — which, given how much gets built on top of a persona, is a poor trade.

Sitting on customer data nobody has time to read?

We turn sales calls, tickets and reviews into research your messaging can actually stand on.

Explore Branding & Strategy →

Frequently asked questions

Can AI create customer personas?

AI can build a persona from customer evidence you supply, and it does that well — reading hundreds of interview transcripts, reviews or support tickets and finding the patterns is exactly the kind of work it is suited to. What it cannot do is invent a persona from nothing useful. Asked to describe your target customer with no data attached, a model returns a composite of averaged internet content: plausible, confident, and unconnected to anyone who buys from you. The distinction is not whether AI is involved but whether real customer evidence is.

Are AI-generated or synthetic personas accurate?

Only in proportion to the data behind them. A persona built from a representative first-party customer dataset is far more reliable than one produced by a general-purpose model working from training data alone. Research reviewing synthetic persona studies has found weak ecological validity across much of the field, meaning the personas frequently fail to reflect real-world populations and behaviour. Models also tend toward agreeable, idealised answers, so a synthetic persona will often describe how someone thinks they should behave rather than how they actually do.

Can AI replace customer interviews?

No, and the guidance from research practitioners is consistent on this point. AI can help you prepare for interviews, analyse them afterwards at a scale no human could match, and generate hypotheses worth testing. It cannot produce the unexpected observation, the emotional detail, or the moment a customer contradicts themselves — which is usually where the real insight sits. The practical rule is that AI multiplies the value of research you have done, and cannot substitute for research you have not done.

What customer data should you feed an AI for persona research?

Anything where customers speak in their own words. Sales call recordings and transcripts, support tickets and chat logs, product reviews including your competitors', survey free-text responses, win-loss interview notes, cancellation reasons, and search queries that bring people to your site. Structured analytics tell you what people did but not why, so they support the picture rather than form it. Remove personal identifiers before uploading anything, and check your organisation's policy on what customer data may be processed by external tools.

How do you validate an AI-generated persona?

Three checks, in order. First, traceability: every claim in the persona should map back to specific source material, and anything that cannot be traced gets deleted rather than kept as a guess. Second, falsification: ask what evidence would prove the persona wrong, then look for it deliberately, since models tend to confirm rather than challenge. Third, human contact: take the persona to five or six actual customers and ask where it is wrong. That last step is the one teams skip, and it is the only one that catches confident fabrication.

THE LAB REPORT

Tactics that move metrics — every Tuesday.

Be an early subscriber. No spam, unsubscribe anytime.