Solutions
Experience measurement
Track sentiment and KPIs with AI-driven gap analysis
Strategic & foundational discovery
Uncover market whitespace with AI-led foundational studies
Journey & behavioral research
Connect user actions to motivations across the lifecycle
Market & consumer Insights
Understanding markets, audiences, & opportunity
Concept & prototype testing
Test designs and prototypes with rapid feedback
Agents
Design
Structure rigorous studies
Field
Run adaptive studies at scale
Analyze
Get statistical analysis you can trust
Synthesize
Turn results into research reports
Deploy
Email
Reach external audiences with native deliverability
Panels
Recruit from 300K+ verified participants
Web apps and websites
Embed studies in web experiences
Mobile apps
Run studies in iOS and Android apps
Customers
Community
Events
Join curated gatherings shaping the future of research
Blog
Insights on integrating AI into research craft
Book icon
Guides
Ultimate playbooks for enterprise survey research
Pricing
Sign in
Book a demo
Sign in
Book a demo
Guide

How to Run: Customer Persona Research Study

October 2, 2026

By The Sprig Team

Example H2
Example H3
Example H4
Example H5
Example H6

Introduction

Running persona research takes five decisions in order, and the first one generally decides whether the rest of the study is defensible.

Separate the two objects. A segment is a statistical partition of a population that can be sized, tested and shown to be wrong. A persona is a narrative design artifact that is singular, concrete and, by its own definition, invented.

Build the segmentation first and the personas on top of it. A persona built without a segmentation underneath has no size attached, and the arithmetic for what a richly specified composite covers is one line long.

Field a frozen battery and size it from the number of segmentation variables. A published simulation gives you a derived rule rather than a convention.

Then validate every persona attribute against the prevalence in your own data, because it is what the critique literature has asked for since 2006 and almost nobody does it.

A persona carrying 21 independent attributes describes roughly 134 people in a United States population of 281 million. That figure comes from a deliberately simplified model, which assumes the attributes are independent and each held by half the population. Real attributes correlate, which raises it, and real base rates are often lower, which lowers it. The guide shows the calculation so you can run it on your own personas and see the order of magnitude.

This guide covers:

  • Separate segments from personas, because they are different objects with different claims
  • Discover segments statistically, then build personas on top of the segments you found
  • Size the sample at 70 respondents per segmentation variable, from a derived rule
  • Prefer latent class analysis over k-means, because scaling decisions silently drive k-means
  • Validate every persona attribute against its prevalence in your data before anyone acts on it

Platform capabilities and documentation cited here were verified against sprig.com and docs.sprig.com on September 17, 2026. No prices appear anywhere in this guide.

What persona research actually is

The phrase typically covers two operations that get run together and should be kept apart. Separating them is the argument this whole guide is built around.

The two objects

A persona is a narrative artifact describing one invented person in concrete terms: their situation, what they are trying to do, what gets in their way, and enough specific detail that a design team can picture them. It is singular rather than aggregate, it is partly invented by construction, and its purpose is to move a team from designing for an average to designing for someone. It is not a population estimate and it does not carry a size.

A segment is a statistical partition. It has a size, a membership rule, and a stability property. It is a claim about a population and it can be wrong, which is what makes it useful.

A persona is a narrative design artifact. It is singular, concrete, and invented, which its own originators acknowledge. Its job is generally to get a design team reasoning about a specific person instead of an average.

| | Segment | Persona | |--------------|:--------------------------------------------------------:|:-----------------------------------------------------------------------:| | What it is | A statistical partition of a population | A narrative artifact describing one invented person | | Unit | A group | An individual | | Has a size | Yes, with membership probabilities | No, and asking how many people it describes is the critique | | Can be wrong | Yes, and that is the point | Not as a whole, though each attribute on it can be checked against data | | Built from | Survey or behavioural data and a clustering model | A segment, plus qualitative material and judgment | | Good for | Sizing an opportunity, targeting, prioritising | Making a design team reason about a person | | Bad for | Making a design team feel anything | Any claim about how many people are like this | | Fails when | The structure was constructed and reported as discovered | It is treated as a population estimate |

Where the confusion comes from

The two typically arrived from different traditions. Segmentation comes from quantitative marketing research, and the persona was introduced to software design by Cooper, A. (1999), The Inmates Are Running the Asylum, Sams Publishing, published in March 1999. The 1998 date that circulates widely is wrong.

Writing from inside Cooper's own firm, Brechin, E. (2002), "Reconciling Market Segments and Personas," Cooper Journal, put the relationship plainly: "Market segmentation provides a quantitative breakdown of the market, while personas provide a qualitative analysis of user behavior." She calls them complementary tools.

That is a practitioner source rather than a peer-reviewed one, and no peer-reviewed paper whose central contribution is this distinction was located for this guide. The distinction is nonetheless what makes the academic critique legible, as the critique section shows.

The committed recommendation

Run the segmentation. Build personas on top of the segments you find, where a team needs them, and the evidence in the critique section says that is more often than a purely quantitative instinct would suggest. Validate the persona attributes against the data underneath.

That order matters more than any other decision in the study. A persona built on a segmentation has a size and a membership rule behind it. A persona built on impressions commonly has neither, and the critique literature is aimed squarely at the second kind.

When to use persona research

Use it when a team is making design or product decisions for a population it cannot picture, when you have or can collect survey data on that population, and when somebody will act on the result.

Three conditions that make it the right approach

The first is a real audience question. Persona work generally earns its cost where the team genuinely disagrees about who they are building for, and it is typically wasted where everyone already agrees.

The second is data you can actually collect. Segmentation needs a battery fielded to a sample, and a project with no route to respondents typically produces personas assembled from opinion.

The third is a decision downstream. The critique literature's most uncomfortable finding is that personas are frequently used for communication rather than for design, so name the decision in advance.

What it answers well

Which groups exist in your user base, how large each one is, what distinguishes them, and what a designer should hold in mind when building for the largest or the most valuable.

The first four of those are segmentation outputs. The fifth is what a persona adds, and it is a communication contribution rather than a statistical one.

When not to use persona research

Five things persona research cannot tell you. Each names the instrument that answers the question instead.

It cannot tell you how many real people the persona describes

This is the limit the whole critique rests on, and the arithmetic is short enough to run on your own personas.

Chapman, C. N. and Milham, R. P. (2006), "The Personas' New Clothes: Methodological and Practical Arguments against a Popular Method," Proceedings of the Human Factors and Ergonomics Society Annual Meeting 50(5), 634 to 636, DOI 10.1177/154193120605000503, take a persona carrying 21 identified attributes.

In their words: "Imagine each of those represents a variable with uniform random distribution, independence, dichotomous values, and 0.50 base rate. Under those assumptions, the composite data for 'Patrick' would represent (0.5)21 * 100% = 0.000048% of the population, or approximately 134 people in the United States."

Their conclusion is blunt: "There is essentially no way to generalize from a well-specified persona to a population of interest, and thus no way to say anything about the users of interest."

The assumptions are deliberately simplifying, since real attributes correlate and base rates are not all one half. The direction of the result survives that, and the order of magnitude is what matters.

The instrument that answers this is a sized statistical segmentation, which returns explicit segment sizes and membership probabilities.

It cannot tell you whether the grouping is real

A clustering algorithm returns groups whether or not groups exist. The output looks identical in both cases.

Ernst, D. and Dolnicar, S. (2018), "How to Avoid Random Market Segmentation Solutions," Journal of Travel Research 57(1), 69 to 82, DOI 10.1177/0047287516684978, applied a structure test to 32 tourism survey datasets. Only two, about 6 percent, contained genuinely naturally occurring segments.

Their audit of published work is worse. Of 78 segmentation studies published between 2010 and May 2016, 53 were data-driven, and approximately 92 percent of those were at risk of having presented a random solution.

The instrument that answers this is bootstrap stability analysis using the Adjusted Rand Index, covered in the analysis chapter.

It cannot tell you what these people will choose or pay attention to

A persona encodes attributes and goals. It does not encode choice behaviour, and the two commonly come apart.

The honest counterweight belongs in the same sentence. Salminen, J., Jung, S.-G., Chowdhury, S., Şengün, S. and Jansen, B. J. (2020), "Personas and Analytics: A Comparative User Study of Efficiency and Effectiveness for a User Identification Task," CHI '20, DOI 10.1145/3313831.3376770, gave 34 participants two systems built on identical underlying data. On a user-identification task the persona system beat the analytics system on both speed and accuracy.

So the correct claim is that they answer different questions, not that one dominates. For preference, the instrument is a conjoint or MaxDiff study. For behaviour, it is product analytics or a live experiment.

It cannot tell you whether the persona changed any decision

Matthews, T., Judge, T. K. and Whittaker, S. (2012), "How do designers and user experience professionals actually perceive and use personas?" CHI 2012, 1219 to 1228, DOI 10.1145/2207676.2208573, interviewed 14 experienced practitioners, ten designers and four UX professionals.

Their headline finding, verbatim: "Practitioners used personas almost exclusively for communication, but not for design."

Friess, E. (2012), "Personas and decision making in the design process: an ethnographic case study," CHI 2012, DOI 10.1145/2207676.2208572, watched a design team's decision-making directly and found that personas made relatively few appearances in the designers' language during decision sessions, despite the investment in creating them.

Fourteen practitioners in one study and one team in the other. Both are small, and both generally point the same way.

The instrument that answers this is a controlled comparison of design outputs with and without the artifact, run on practitioners.

It cannot tell you whether it is encoding a pattern or a stereotype

Turner, P. and Turner, S. (2011), "Is stereotyping inevitable when designing with personas?" Design Studies 32(1), 30 to 44, DOI 10.1016/j.destud.2010.06.002, conclude that "In short, as soon as we picture the kind of people for whom we are designing (cf. Lippmann's definition of a stereotype) we may well be committed to a stereotype."

The complication is the interesting part and it should not be flattened. They also note that "stereotypes are not necessarily bad. As other authors have commented, they are often disconcertingly accurate."

Their own evidence makes the point. Their students' invented iPhone-user personas were 78.4 percent male against the 75 percent comScore reported for the actual UK iPhone user base in 2009, and 87 percent of the invented users fell in the 18 to 44 band the same market report described.

So the personas were stereotyped and approximately correct at the same time. That tension is typically not resolvable by being more careful, and the practical response is to check each attribute against data rather than to trust or distrust the composite.

The instrument that answers this is attribute-level validation against the underlying survey, which the interpreting chapter sets out as a procedure.

Study design, and the fork that decides everything else

The fork is whether you are discovering groups or describing ones you already have. Many teams assume they are doing the first, and in the one domain where this has been measured most work turned out to be the second or third.

Three kinds of segmentation, and most work is the third

Dolnicar, S., Grün, B. and Leisch, F. (2018), Market Segmentation Analysis: Understanding It, Doing It, and Making It Useful, Springer, DOI 10.1007/978-981-10-8818-6, open access, set out a three-way distinction that reframes the whole exercise.

Natural segmentation is where distinct market segments exist in the data and the job is to find them. This is generally what most teams assume they are doing.

Reproducible segmentation is where natural segments do not exist but the data are not unstructured either. In their words, the data "contain some structure, other than cluster structure, making it possible to generate the same segmentation solution repeatedly."

Constructive segmentation is where neither cluster structure nor any other usable structure exists, and the analyst is partitioning rather than discovering.

The empirical finding is the one worth carrying. Ernst, D. and Dolnicar, S. (2018) applied a structure test to 32 tourism survey datasets and found that only two, about 6 percent, contained natural segments, while roughly 22 percent had no usable structure at all.

The remaining 72 percent fell in the middle category, reproducible rather than natural, which is a more encouraging picture than the 6 percent figure alone suggests. That figure is domain-specific to tourism survey data and should not be quoted as a universal rate.

In those 32 datasets, most segmentation was construction rather than discovery. The authors' own position is that it is still worth doing, because targeting subgroups is "more promising" than attempting to satisfy the entire range of consumer needs.

What that changes about how you report

If your segmentation is constructive, say so. A constructed partition is often extremely useful and it is not a discovery, and reporting one as the other is typically the failure mode that makes segmentation work distrusted inside organisations.

The stability analysis in the analysis chapter is what tells you which of the three you have. Run it before you name anything.

Segment first, persona second

The committed design is two stages. Field a battery, discover or construct segments, size them, then build personas on top of the segments a team will actually design for.

That ordering is what gives each persona a size and a membership rule. Without it, a persona is generally an assertion, and the arithmetic in the limitations section is what it runs into.

Not every segment needs a persona. Build them for the two or three segments somebody is going to design for, and leave the rest as sized segments in the report.

What goes into the battery

Segmentation variables are the questions the clustering runs on, and choosing them is the most consequential decision in the study.

Pick variables that plausibly differ between groups and that relate to the decision you are making. Needs, attitudes, behaviours and contexts of use generally do this well.

Demographics usually do not. They are easy to collect and they frequently produce segments that are demographically clean and behaviourally identical, which is the classic failure of segmentation work.

Keep the count deliberate. Every additional segmentation variable raises the sample you need, on the rule in the sample-size chapter, and adds another dimension in which the structure can dissolve.

The variables that must not go in

Keep profiling variables out of the segmentation itself. Anything you want to describe the segments with afterwards, such as tenure, plan, or spend, should be held back and used for profiling rather than included in the clustering.

Including them means the segments are partly defined by the thing you were going to use to characterise them, and the profile then says nothing you did not put in.

That separation is typically easy to lose in a spreadsheet, so mark each field as segmentation or profiling before the analysis starts.

Writing the instrument

The segmentation battery is a frozen object. Once it is fielded and clustered, changing it means the segments from the next wave are not the segments from this one.

The battery

A battery is a set of items sharing one scale, answered by everyone, covering the needs and attitudes you expect to differ between groups.

Keep every item answerable by every respondent. A segmentation variable that only applies to some people produces missing data that the clustering has to handle, and handling it well is typically harder than avoiding it.

Write items that can genuinely split. An item everyone agrees with carries no information for clustering, however true it is, and a battery of consensus statements commonly produces segments that differ only by response style.

A worked battery

For a product team segmenting on needs, a battery might carry twelve items on one five-point agreement scale, each naming a concrete need rather than a value.

I need to get set up without talking to anyone.

I need to see exactly what changed before I approve it.

I need this to work the same way on my phone as on my laptop.

Those are specific enough that people genuinely differ on them. "I value quality" is not, and a battery of items like it will cluster on acquiescence.

Matrix questions and the freeze

A matrix bundles a battery into one table sharing a scale, which is what you want for segmentation work.

One documented Sprig constraint matters here and it cuts both ways. Matrix rows and columns cannot be deleted or reordered after launching a survey, though new rows and columns can be added.

That constraint cuts against you more than it helps. A battery item that turns out to be badly worded, or that everyone answers identically, cannot be removed for the life of the study.

It does happen to prevent one specific error, which is silently reordering or dropping items between waves and breaking the comparison. That is a narrow benefit against a real cost.

Plan around it. The battery has to be right before launch rather than after the first responses arrive, which makes the pilot in the fielding chapter load-bearing rather than optional.

What else the instrument needs

Add the profiling questions you held out of the battery, placed after it. Add one open-text probe, because the qualitative material is what turns a segment into something a design team can picture.

Keep the whole instrument short enough that people finish it. A twelve-item battery plus profiling plus one open text is already a longer survey than most in-product instruments, which is a fielding constraint covered later.

Scale and scoring choices

One decision in this chapter typically has more influence on your results than most analysts realise, and it is not the one people argue about.

Use one scale for the whole battery

Every item on the same scale, same length, same labels, same direction. This is not an aesthetic preference.

A battery on mixed scales cannot be clustered without rescaling, and rescaling is often the decision that quietly drives the answer.

Scaling is a weighting decision in disguise

This is the part worth understanding properly, because it is generally where segmentation studies go wrong invisibly.

k-means minimises within-cluster sum of squared distances. Each variable's influence on that objective is proportional to its variance in whatever units it happens to be measured in.

Put a one-to-five item next to age in years and the standard deviations differ by roughly a factor of fifteen, which means their contributions to the distance calculation differ by roughly a factor of two hundred. Unstandardised, the solution is close to a partition on age, and the battery is nearly inert.

Dolnicar and colleagues make the same point with a starker example, noting that with an unstandardised spend variable a difference of one dollar per person per day is weighted equally against the difference between liking to dine out and not.

Nothing in the output typically tells you this happened. The clusters usually look fine.

And standardising is also a choice

Do not read the above as an instruction to z-score everything and move on. Standardising makes the weighting explicit instead of removing it.

Milligan, G. W. and Cooper, M. C. (1988), "A study of standardization of variables in cluster analysis," Journal of Classification 5(2), 181 to 204, DOI 10.1007/BF01897163, compared standardisation approaches by simulation and found that dividing by the range recovered the underlying cluster structure consistently better than the traditional z-score transformation, which was less effective in several situations.

So the practical instruction is to keep the battery on one scale so the question mostly does not arise, to standardise by range where you must mix, and to record what you did.

The alternative that sidesteps it

Latent class analysis avoids the problem rather than managing it, which is one of the reasons the analysis chapter prefers it. The detail belongs there.

Sample size and precision

The expectation going into this topic is commonly that segmentation sample size is unsourced folklore. That expectation is wrong, and correcting it is worth doing loudly, because a derived rule exists and almost nobody cites it.

The rule, and where it comes from

Dolnicar, S., Grün, B., Leisch, F. and Schmidt, K. (2014), "Required Sample Sizes for Data-Driven Market Segmentation Analyses in Tourism," Journal of Travel Research 53(3), 296 to 306, DOI 10.1177/0047287513496475, ran a simulation study using artificial data of known structure.

Their abstract states the result plainly: "Under all simulated data circumstances, a sample size of 70 times the number of variables proves to be adequate."

The method is what makes it citable. They generated data where the true segment structure was known, varied the difficulty, ran k-means with thirty random initialisations, and measured how well the recovered segments matched the truth using the Adjusted Rand Index.

One honest caveat, since this guide recommends latent class analysis as the primary method. The rule was derived under k-means, so applying it to a latent class design is an approximation instead of a result, and it is the same kind of cross-framework transfer this guide criticises two paragraphs below. Use it as a planning floor and say which method you actually ran.

That is a derivation against ground truth rather than a convention someone repeated. The paper also reports 60 times the number of variables, and the scope matters. That figure applies to an aggregated-data setting instead of to easier data generally, so 70 is the planning number and 60 is not a licence to field smaller.

The table

Work from the number of segmentation variables, which is the count of items the clustering runs on and not the length of your questionnaire.

| Segmentation variables | n at 60 per variable | n at 70 per variable | n at 100 per variable | |:----------------------:|:--------------------:|:--------------------:|:---------------------:| | 5 | 300 | 350 | 500 | | 8 | 480 | 560 | 800 | | 10 | 600 | 700 | 1,000 | | 12 | 720 | 840 | 1,200 | | 15 | 900 | 1,050 | 1,500 | | 20 | 1,200 | 1,400 | 2,000 | | 25 | 1,500 | 1,750 | 2,500 | | 30 | 1,800 | 2,100 | 3,000 |

The 100 column is the same authors' later practical guidance in the 2018 book, which recommends at least 100 respondents per segmentation variable. Treat it as the conservative version rather than as a separate finding.

Read the table as completed responses, not invitations. Apply your response rate on top.

The competing rules, and their actual provenance

Three other figures commonly circulate. Each has a real source and each was derived for something other than planning your study, which is worth knowing before somebody quotes one at you.

| Rule | Source | What it was derived for | How to treat it | |:---------------------------------------------------:|:----------------------:|:--------------------------------------------------------------------------:|:-----------------------------------------------------------------------------------------------------------------------------------------------------------------------:| | n at least 2 to the power p, better five times that | Formann (1984) | Goodness-of-fit testing in latent class analysis, on binary variables only | Narrow and frequently misapplied. Dolnicar and colleagues note it "can therefore not be assumed to be generalisable to other algorithms, inference methods, and scales" | | n at least 10 times p times k | Qiu and Joe (2015) | Constructing artificial datasets to benchmark clustering algorithms | Repurposed. Cite only with the caveat | | n at least 60 to 70 times p | Dolnicar et al. (2014) | Simulation against known ground truth | Best-sourced. This is the one to plan from | | n at least 100 times p | Dolnicar et al. (2018) | Practical extension of the same programme | The conservative version |

The first rule is typically the one that causes trouble. Applied to a battery of twelve five-point items it produces a number that has nothing to do with the study being planned, because it was derived for binary variables in a different modelling framework.

What the rule does not cover

The rule sizes the study for recovering segment structure. It does not size it for reporting on the smallest segment you find.

If your solution returns a segment holding 8 percent of the sample and somebody wants to characterise it, that subgroup carries the precision its own count supports instead of the precision of the whole study. Check the smallest segment's n before anyone quotes a percentage inside it.

And the rule assumes you are clustering on the variables you planned to cluster on. Adding three more after seeing the data changes both the rule's input and the honesty of the result.

Audience, targeting and screening

Segmentation partitions whatever population you field to. Getting that population wrong is generally not recoverable at analysis, because the segments will faithfully describe the wrong people.

Define the population before the battery

Write down who the segments are meant to describe. Current users, all users including dormant ones, the addressable market, or buyers in a category you do not yet serve are four different studies.

The most common mistake is typically fielding in-product to active users and then presenting the result as a market segmentation. Those segments describe people who already chose you, which is a useful population and not the market.

Screening

Screen on category behaviour rather than on stated interest, and keep the screener out of the battery. A screener that previews the battery's themes primes the answers you are about to cluster on.

Where the population is your own users, the behavioural definition is usually available without asking. Field to the cohort instead of screening for it.

Quotas, carefully

Quotas generally make the sample match a known population on the quota variables. They also constrain what the segmentation can find.

Set quotas on variables you are not segmenting on. Quota-filling on a variable that is also in the battery partly determines the segment sizes you will recover, which then get reported as a finding.

Where the sample comes from

A study about your own users can be fielded to them directly. A study about a market you do not serve needs a panel, which brings its own composition question and its own professional-respondent problem.

Whichever it is, record it in the report next to the segment sizes. A segment holding 30 percent of a panel sample and 30 percent of your user base are different claims.

Fielding and delivery

A segmentation battery is typically longer than most in-product instruments, which constrains where it can run.

Channel and length

Twelve battery items plus profiling is a survey rather than an intercept. An in-product widget can carry it, and completion will generally be lower than for a two-question instrument, so plan the invitation volume from the sample table instead of from habit.

A link survey, emailed or panel-distributed, is generally the better channel for a full battery. In-product works well where the battery is short and the population is exactly your active users.

Fielding window

Field the whole battery in one window. A segmentation assembled from responses collected months apart mixes people answering in different contexts, and the clustering typically has no way to know.

Where the population is seasonal, cover the season or state which part of it you covered.

One thing to check before launch

Pilot the battery on twenty to thirty people and look at the item distributions before the full launch. An item everyone answers the same way is dead weight in the clustering, and it is much cheaper to replace before the freeze than after.

This is also the moment to confirm the battery renders as you expect on mobile, where a matrix displays as an accordion rather than a full table.

Timing and cadence

Segmentation is episodic. Run it when the audience question is live, not on a schedule.

Re-run when the population or the product has changed enough that the battery no longer describes the decisions people make. That is usually years instead of quarters.

When you do re-run, field the identical frozen battery. The comparison between waves is only meaningful if the segmentation variables are the same, and a battery someone improved between waves produces a change you cannot attribute.

And expect the segments to move even so. A re-run on a frozen battery can return a different number of segments, which is itself a finding about how stable the original structure was.

Quality control and data hygiene

Segmentation is typically unusually sensitive to bad responses, because a straight-liner produces a perfectly consistent profile that clusters cleanly into its own group.

The checks that matter here

Straight-lining first. Compute the variance of each respondent's answers across the battery and flag the ones near zero. A respondent who gave the same answer to twelve items has told you nothing and will anchor a cluster.

Speed second. Flag the fastest few percent and read some of their records before deciding.

Then check for a satisficing cluster. If one recovered segment is defined mainly by low-variance answering rather than by content, it is a response-style artifact and not a market segment.

Note the sequencing. If you excluded zero-variance respondents before clustering, the extreme cases are already gone, and what remains to look for is a segment built on near-flat instead of perfectly flat answering. Run the check on the answer variance within each segment rather than on the exclusions.

Missing data

Decide the rule before you analyse. Most clustering implementations either drop incomplete cases or require imputation, and both decisions change the result.

Dropping is usually cleaner for a battery everyone could answer. Where missingness is substantial, report how many cases you dropped and whether the dropped ones differed on the profiling variables.

The exclusion rule goes in writing

Write it down before you look at the segments. A rule applied after seeing which respondents produced an inconvenient cluster is not quality control.

Analysis: the core calculation

Sprig does not do this step. No statistical segmentation of any kind is documented in the product, so the clustering happens in a connected AI client or a statistics tool against exported data. The product chapter states that plainly and this chapter assumes it.

Prefer latent class analysis over k-means

Both typically partition respondents. They differ in one way that matters for survey batteries.

Vermunt, J. K. and Magidson, J. (2002), "Latent Class Cluster Analysis," in Hagenaars, J. A. and McCutcheon, A. L. (Eds.), Applied Latent Class Analysis, Chapter 3, pages 89 to 106, Cambridge University Press, state it with a condition worth keeping: "no decisions have to be made about the scaling of the observed variables: for instance, when working with normal distributions with unknown variances, the results will be the same irrespective of whether the variables are normalized or not. This is very different from standard non-hierarchical cluster methods, where scaling is always an issue."

Read the condition instead of the headline. The scale invariance is a property of the Gaussian case with unknown variances, not a general property of every latent class model, so it reduces the weighting problem for a battery of that kind rather than removing it everywhere.

Two further advantages hold more broadly. It gives formal model selection through information criteria such as AIC, BIC and CAIC, and it handles variables of mixed measurement levels natively.

k-means generally remains a reasonable choice for a clean single-scale battery, and it needs the discipline in the rest of this chapter to be trustworthy.

Choosing the number of segments

Choose the number of segments using the variance ratio criterion, also called Calinski-Harabasz, or BIC, or the gap statistic. With latent class analysis, BIC does this job as part of the model fit. Then check stability across bootstrap replications before committing to a value.

Do not use the elbow method. Schubert, E. (2023), "Stop using the elbow criterion for k-means and how to choose the number of clusters instead," ACM SIGKDD Explorations Newsletter 25(1), 36 to 42, DOI 10.1145/3606274.3606278, argues that it "severely lacks theoretic support."

His mechanical objection is typically the one that lands. Changing the scaling of the axes "may well change the human interpretation of an 'elbow'," and he asks what meaning "angles", "distances", "elbows", and "slopes" have on a graph comparing two quantities of different scales.

In his tests on a dataset with many clusters and on uniform data, all the elbow-based methods failed. Use the variance ratio criterion, also called Calinski-Harabasz, or BIC, or the gap statistic instead. With latent class analysis, BIC does this job as part of the model.

Stability is the step nobody runs

A single clustering run is generally not a result. k-means starts from randomly selected points, and, in Ernst and Dolnicar's words, "different starting points lead to different segmentation solutions."

Their audit found that of the 53 data-driven studies, 30 used start-dependent algorithms and only three reported how many starting points they tried. Their conclusion: "one single calculation is never enough in data-driven market segmentation."

The remedy is bootstrap stability analysis, and the procedure has a step that is easy to leave out and makes the rest meaningless without it.

Draw two bootstrap resamples and cluster each one. Then use each solution to assign every case in the full original dataset to a segment, and compare those two sets of assignments using the Adjusted Rand Index.

The prediction step is what makes the comparison possible. Two bootstrap samples contain different observations with different duplicates, so their partitions are not directly comparable, and an index computed without projecting both back onto the same cases is measuring nothing. Repeat across many pairs and read the distribution.

The validation checklist

Run these in order before anyone names a segment. The reason each one is here is that published work routinely skips it.

  • Mark every variable as segmentation or profiling, and cluster only on the first set
  • Record the scaling decision, including the decision not to rescale
  • Run at least 100 random restarts and record the number
  • Compute the solution across a range of k instead of a single value
  • Choose k by variance ratio criterion, BIC or the gap statistic, never by eyeballing an elbow
  • Bootstrap the solution and report the Adjusted Rand Index across replications
  • Classify the result as natural, reproducible or constructive, and say which in the report
  • Check the smallest segment's count before reporting any percentage inside it
  • Check whether any segment is defined mainly by response style rather than content
  • Profile the segments on the held-out variables, not on the clustering variables

What the Adjusted Rand Index is telling you

It measures agreement between two partitions, corrected for chance. High agreement across bootstrap replications means the same solution keeps appearing, which is the reproducibility property.

There is no universal threshold that licenses reporting, and anyone offering one has invented it. What the index gives you is a comparison across values of k, and reading it requires care. Global stability is easiest to achieve at small k, so a rule of simply taking the most stable solution will select two segments almost every time.

Read stability per segment instead of only in aggregate, and choose among the solutions that clear a stability bar on the grounds of whether the segments are substantively distinguishable. Stability narrows the candidates and does not pick the winner.

Benchmarks and what a good result looks like

There is no sourced cross-industry benchmark for persona research. No published norm exists for how many personas a team should have, what share of a user base a persona should cover, or what a good segmentation looks like across industries.

A field whose own 2022 systematic review opened by noting that nobody had yet reviewed how personas deliver their claimed benefits does not have benchmarks to publish.

What exists instead are methodological floors

Three numbers in this guide look like benchmarks and are not. Label them correctly whenever you use them.

At least 70 respondents per segmentation variable is a sample-size floor derived by simulation. Six percent of 32 tourism datasets containing natural segments is a finding about one domain's data. Approximately 92 percent of 53 published data-driven studies being at risk of a random solution is an audit result about published practice.

None of those is a target you hit or miss. They are constraints on how the study is built and read.

Do not invent a number of personas

There is no published basis for three, or five, or any other count. The number of personas you build should follow from the number of segments your analysis supports and the number a team can hold in mind, which is a communication constraint rather than a statistical one.

If someone asks what the typical number is, the honest answer is that nobody has published one.

What good actually looks like

A segmentation that reproduces across bootstrap replications, has segments large enough to report on, profiles differently on variables that were not used to build it, and comes with an explicit statement of whether the structure was discovered or constructed.

Then, on top of it, personas whose every attribute has a prevalence attached from the underlying data.

Interpreting and acting on the result

The output is a set of sized segments with profiles. Turning that into personas is where the value is added and where the credibility is usually lost.

Read the profiles before naming anything

Profile each segment on the variables you held out, and read the differences before you invent a name. A name commonly arrives with a story attached, and the story starts shaping how people read the numbers within about a minute.

Look for segments that differ on the held-out variables. A partition that separates cleanly on the battery and not at all on anything else has found response patterns instead of people.

The attribute validation spec

This is the step the critique literature has asked for since 2006, and it is the step this guide most wants you to run. For every attribute you are about to put on a persona, record its prevalence in the segment underneath it.

| Persona attribute | Prevalence in the segment | Prevalence in the rest of the sample | Verdict | |:-----------------------------------------:|:-------------------------:|:------------------------------------:|:------------------------------:| | Needs to set up without talking to anyone | 81 percent | 44 percent | Keep. Distinctive and dominant | | Works primarily on mobile | 52 percent | 49 percent | Cut. Not distinctive | | Approves purchases themselves | 63 percent | 22 percent | Keep. Distinctive | | Has used a competitor before | 38 percent | 35 percent | Cut. Not distinctive | | Is named Patrick and cycles to work | Not measured | Not measured | Label as invented |

Written as a rule: to validate a persona, take every attribute it asserts, compute how common that attribute is inside the segment and how common it is among everyone outside the segment, and keep only the attributes where the two differ. Attributes that are common everywhere describe your whole user base rather than this persona. Attributes you never measured are inventions and should be visibly labelled as such on the artifact.

Three verdicts, applied strictly. Compare against the rest of the sample instead of the full sample, since the full sample contains the segment and a large segment inflates its own comparator.

An attribute that is high in the segment and similar outside it is typically not a persona attribute, it is a fact about everyone. An attribute you never measured is invented, and the persona should say so on its face.

Chapman, C. N., Love, E., Milham, R. P., ElRif, P. and Alford, J. L. (2008), "Quantitative Evaluation of Personas as Information," Proceedings of the HFES Annual Meeting 52(16), 1107 to 1111, DOI 10.1177/154193120805201602, tested the 2006 prediction against eight datasets, six of them surveys ranging from 268 to 10,307 participants and two of them simulated. Correlations between observed and predicted attribute prevalence ran from r = 0.394 to r = 0.713 across the six surveys, with the two simulated datasets reaching 0.857 and 0.869.

Their recommendation is exactly this procedure: personas should be "assessed empirically before they are assumed to describe real groups of people."

What a finished persona contains

The artifact itself is short. A page is plenty, and the fields below are the ones that earn their place.

| Field | Source | Notes | |:----------------------------------:|:------------------------:|:----------------------------------------------------------------------:| | Segment name and size | The segmentation | The number of people and the share of the base, stated on the artifact | | Membership rule | The segmentation | What puts someone in this segment rather than another | | Three to five validated attributes | The prevalence table | Each with its rate inside the segment and outside it | | What they are trying to do | Open text and interviews | The job, in the respondents' own language where possible | | What gets in their way | Open text and interviews | The friction, specific enough to design against | | One representative verbatim | Open text | Quoted, not paraphrased | | Invented detail | Nobody | Name, photograph, commute. Visibly separated from everything above | | Date and study | The study record | So a reader in a year knows what this rests on |

The size line at the top is what separates this artifact from the ones the critique literature is aimed at. A persona that cannot state how many people it describes is an assertion.

Mark the invented parts as invented

A persona generally needs concrete detail to do its communication job, and that detail is fiction. The fix is not to remove it, it is to label it.

Put the measured attributes and the invented ones in visually different places on the artifact. Anyone reading it should be able to tell in two seconds which parts came from data.

This also defuses the stereotyping problem in the only practical way available. An invented detail that is labelled as invented cannot be mistaken for a finding about a group.

What to hand the team

The segment sizes with their membership rules. The profiles on held-out variables. Two or three personas with prevalence attached to every measured attribute and the invented material marked. And one sentence saying whether the structure was natural, reproducible or constructive.

That last sentence is what keeps the work honest six months later, when somebody quotes a segment size in a planning meeting.

Running the analysis with Claude or ChatGPT

Sprig documents no clustering, latent class analysis, or statistical segmentation of any kind, so this analysis runs entirely outside the product against exported data. This chapter covers the whole workflow.

What the analysis produces

Three artifacts. A segmentation solution with segment sizes and a stability statistic. A profile table comparing segments on the held-out variables. And an attribute prevalence table for each persona you intend to build.

The first is the finding. The second is what makes it interpretable. The third is what makes the personas defensible.

Setup and getting your data in

Export the study as CSV with responses and the attributes you intend to profile on. Keep the labelled export alongside the coded one, since the labels are your codebook.

Separate the battery columns from everything else before you start, because the whole discipline of this analysis depends on clustering only on the segmentation variables.

The MCP connector reaches studies and responses directly from a supported client and retrieves up to 1,000 responses at a time. For a segmentation study sized from the table earlier in this guide, that means the data arrives in pages, and the CSV export is usually the simpler route.

Prompt one: the segmentation with stability

What this prompt does: fits a segmentation to a battery and tests whether
the solution is stable.
What it returns: a recommended number of segments, segment sizes, and an
Adjusted Rand Index across bootstrap replications for each candidate k.

The file is [FILENAME, default the exported study CSV]. The segmentation variables are the columns
[COLUMN LIST, default all columns beginning "Q" in the battery block].
Do not cluster on any other column.

Use code to calculate this, not estimation.

Rules:
1. Exclude any respondent with a blank or non-response on any segmentation
   variable, and report how many you excluded.
2. Also exclude any respondent whose answers across the battery have zero
   variance, which is straight-lining, and report that count separately.
3. Fit solutions for k from [K MIN, default 2] to [K MAX, default 8]. For
   each k, run at least [RESTARTS, default 100] random restarts and keep the
   best solution by within-cluster sum of squares.
4. For each k, run [BOOTSTRAPS, default 100] pairs of bootstrap resamples.
   Cluster each resample, then use each solution to assign every case in the
   FULL original dataset to a segment, and compute the Adjusted Rand Index
   between those two sets of assignments. Report the mean and the spread
   across pairs. Do not compute the index between the bootstrap partitions
   directly, which is undefined.
5. Report the variance ratio criterion and BIC for each k. Do not use or
   mention the elbow method.
6. Any segment holding fewer than [MIN SEGMENT, default 30] respondents is
   reported as "too small to profile" instead of as a percentage. A segment
   with zero members is also reported as "too small to profile", not omitted.
7. Recompute the segment sizes for the recommended k a second time by a
   genuinely different route, for example a latent class model on the same
   variables, and compare the two partitions with the Adjusted Rand Index.
8. If the two routes disagree by more than [TOLERANCE, default 0.15] on that
   index, report the disagreement explicitly. Do not reconcile it silently
   and do not pick one.
9. Output one markdown table named "Segmentation stability" with one row per
   k, plus a short list of every exclusion and every disagreement.

Prompt two: the profile and prevalence table

What this prompt does: profiles the segments on held-out variables and
computes attribute prevalence for persona building.
What it returns: a profile table and a prevalence table with base rates.

Using [FILENAME, default the same exported study CSV] and the segment
assignment from the previous step, the profiling variables are
[PROFILE COLUMN LIST, default every column not used in the clustering]. These were not used to build
the segments and must not be added to the clustering.

Use code to calculate this, not estimation.

Rules:
1. Exclude blanks and non-responses per variable, and report the count
   excluded for each.
2. Any segment with fewer than [MIN SEGMENT, default 30] respondents is
   reported as "too small to profile". A segment with zero members is also
   reported as "too small to profile".
3. For each profiling variable, report the value within each segment and the
   value among all respondents NOT in that segment, so the comparator sits
   beside every figure. Do not compare against the full sample, which
   contains the segment.
4. For each attribute, compute the difference between the segment rate and
   the base rate, and flag any attribute where the difference is smaller than
   [DISTINCTIVENESS, default 10] percentage points as "not distinctive". This
   default is an arbitrary working convention instead of a derived threshold.
   Say so in the output.
5. Recompute every segment-level figure a second time by a genuinely
   different route, for example aggregating from the raw respondent rows
   rather than from the grouped summary, and compare.
6. If the two routes disagree by more than [TOLERANCE, default 1] percentage
   point on any figure, report the disagreement instead of resolving it.
7. Output two markdown tables named "Segment profiles" and "Attribute
   prevalence", the second sorted by distinctiveness descending.

Reading the output

Read the stability table before the segment sizes. A solution with a high Adjusted Rand Index across replications is one you can report. A solution whose index collapses as k rises is telling you the structure thins out, and the honest report says so.

Read the prevalence table by the distinctiveness column rather than by the raw segment rate. An attribute at 80 percent in a segment and 78 percent in the base describes your whole user base.

What to verify before reporting

Confirm the exclusion counts against the raw response total. Confirm no reported percentage sits inside a segment below the minimum. Confirm the profiling variables never entered the clustering, which is the error that quietly invalidates the whole profile. And confirm the recommended k came from a criterion instead of from a chart.

Pitfalls

Models asked to segment data will typically produce segments, always, including from noise. Requiring the stability statistic is what distinguishes a finding from an output.

Models asked to double-check will re-run the same computation and agree with themselves, which is why both prompts require a genuinely different route.

And models are generally obliging about naming things. Ask for segment profiles before you ask for segment names, because a name supplied early will shape how the profile is described.

The critique you should know

Persona research carries an unusually pointed published critique, and a guide that hides it is worth less than one that uses it.

The arithmetic that started it

Chapman and Milham's 2006 calculation is the most-quoted item in this critique literature and it takes a line to reproduce. For a persona of independent dichotomous attributes each at a 50 percent base rate, the share of the population matching all of them is one half raised to the number of attributes.

| Attributes on the persona | Share of population matching | People in a 281 million population | |:-------------------------:|:----------------------------:|:----------------------------------:| | 5 | 3.125 percent | 8,781,250 | | 10 | 0.098 percent | 274,414 | | 15 | 0.0031 percent | 8,575 | | 20 | 0.000095 percent | 268 | | 21 | 0.000048 percent | 134 | | 25 | 0.0000030 percent | 8 |

Run it on your own persona. Count the attributes it asserts, and read the row.

The assumptions are generous to the persona in one direction and harsh in another. Real attributes correlate, which makes the true figure larger. Real base rates are frequently well below half, which makes it smaller. The order of magnitude is the finding, and it holds.

What the critique is actually aimed at

This matters, because the critique is often quoted as though it kills personas outright.

Their target is the persona as a claim about a population. It is not the persona as a design artifact, and the distinction this guide opens with is what lets you use the strongest available critique without landing in nihilism.

A persona that never claims to describe a population is not touched by this particular objection, though the others in this section still apply to it. A persona presented as "our typical user" is.

The other four objections

Coverage is indeterminate. Chapman and Milham ask how many users a given persona describes, and the method provides no way to answer.

Personas are not falsifiable. There is no result that would show a persona to be wrong, which is a strange property for a research output.

They may be used for communication rather than design. Matthews and colleagues found exactly that in fourteen practitioners, and Friess found personas making few appearances in a design team's actual decision talk.

They are organisationally political. Chapman and Milham warn that "We believe personas are likely to lead to political conflicts and to undermine the ability for researchers to resolve questions with data."

A documented case runs alongside it, though it points somewhere subtler than the warning does. Rönkkö, K., Hellman, M., Kilander, B. and Dittrich, Y. (2004), "Personas is not applicable: local remedies interpreted in a wider context," Proceedings of PDC 2004, Volume 1, 112 to 120, DOI 10.1145/1011870.1011884, report a persona effort on a mass-market telecom project failing "because of patterns of dominance in the telecom branch which were unrecognised at the time."

The politics were already there. The method is credited elsewhere with surfacing social and political issues, and in this case it failed to surface ones that predated it. Read it as a limit on what the artifact can neutralise rather than as evidence that personas generate conflict.

Marsden, N. and Haag, M. (2016), "Stereotypes and Politics: Reflections on Personas," CHI 2016, 4017 to 4031, DOI 10.1145/2858036.2858151, make the related argument that personas risk reinscribing existing stereotypes and follow an I-methodological rather than a user-centred approach, with practitioners knowingly compromising under organisational pressure.

So who writes the persona, and whose view of the customer it encodes, is an organisational question as much as a research one.

The counter-critique, and it is real

Three findings run the other way, and the guide is weaker without them.

Personas measurably corrected false beliefs. Salminen, J., Jung, S.-G., Chowdhury, S., Robillos, D. R. and Jansen, B. (2021), "The ability of personas: An empirical evaluation of altering incorrect preconceptions about users," International Journal of Human-Computer Studies 153, 102645, DOI 10.1016/j.ijhcs.2021.102645, ran a within-participant experiment with 31 employees of an international airline drawn from seven functions including customer relations, ecommerce, loyalty, IT, management, marketing and resource management.

They report that "25 (80.6%) of the participants changed their preconceptions of the typical user after interaction with the persona system," and that "94% of the participants maintained or increased the accuracy of their perceptions." Report the n of 31 alongside those percentages every time, and note that two participants held onto incorrect preconceptions despite contradictory facts.

The format itself helped, with the data held constant. The Salminen and colleagues 2020 study described earlier is the strongest counter-evidence available, because both systems ran on identical underlying data and only the presentation varied. Participants using the persona system completed the task faster, at a mean of 417 seconds against 553 seconds, and located the target segment correctly 25 times out of 34 against 8 out of 34 with analytics.

That is a large difference on a task about identifying users, and it is evidence that the narrative format carries information a dashboard does not. The authors are careful about scope, noting that the analytics system affords insights and capabilities personas cannot, so the finding is about one task rather than about which tool is better.

And the field has been asking the right question. Salminen, J., Guan, K. W., Jung, S.-G. and Jansen, B. (2022), "Use Cases for Design Personas: A Systematic Review and New Frontiers," CHI 2022, Article 543, DOI 10.1145/3491102.3517589, reviewed 95 articles and opened by noting that "While personas have been lauded for their benefits, we could locate no prior review of persona use cases in design, prompting the question: how are personas actually used to achieve these benefits?"

Read that carefully. The field's own systematic review, in 2022, began by observing that the mechanism had never been reviewed. That is not a defence of personas and it is not a dismissal, it is an accurate statement of where the evidence stands.

Where this leaves the method

Personas generally communicate well and describe populations badly. The segmentation underneath is what carries the population claim, and the persona is what gets a team to act on it.

Build both and keep them labelled, and the arithmetic objection stops applying to your work. The other objections do not. Stereotyping, organisational politics, and the finding that personas are used for communication more than for design all survive a segmentation and a label, and they are managed rather than solved.

Common mistakes

Nine that recur, in rough order of how much damage they do. Each one has a tell, which is what to look for when you suspect it rather than what proves it.

Most of them happen during the analysis rather than during the design, which is why a study that was planned carefully can still arrive at a result nobody should act on.

Building personas without a segmentation underneath. The arithmetic in the critique section is what this runs into. A persona assembled from impressions asserts a composite with no size attached, and the more detail it carries the smaller that size gets.

The tell is that nobody can say how many users the persona covers. Ask, and if the answer is a shrug, there is no segmentation underneath.

Reporting a constructed partition as a discovery. Only about 6 percent of one set of 32 datasets contained natural segments. Most segmentation constructs groups, which is legitimate, and describing it as finding groups is not.

The tell is the absence of a stability statistic. A report that never mentions how the solution held up across replications has not checked.

Clustering on demographics. They are easy to collect and they frequently produce segments that are demographically distinct and behaviourally identical.

The tell is a profile table where the segments differ on age and income and on nothing anyone can act on.

Including profiling variables in the battery. The segments then partly encode the thing you were going to characterise them with, and the profile confirms what you built in.

The tell is a profile that is suspiciously clean. If every segment differs sharply on the variable you most wanted to talk about, check whether it was in the clustering.

Choosing k by looking at a chart. The elbow method lacks theoretical support, and the appearance of an elbow changes with the scaling of the axes.

The tell is a methods note saying the number of segments was chosen by inspection.

Running the clustering once. Different starting points give different solutions, and only three of the thirty start-dependent studies in one published audit reported how many starts they used.

The tell is a single number of restarts, or none at all, in the write-up.

Naming the segments before reading the profiles. A name carries a story, and once it is attached people read the profile through it. The order matters more than it sounds like it should.

The tell is that everyone in the room uses the persona name and nobody can recall the segment's size.

Putting invented detail on a persona without labelling it. The name, the commute, the photograph are generally communication devices. Unlabelled, they commonly read as findings, and they are the material the stereotyping critique is aimed at.

The tell is a persona artifact where measured and invented attributes sit in the same visual treatment.

Reporting a segment percentage from a segment too small to carry one. A solution can return a segment holding 4 percent of the sample, and somebody will quote a percentage from inside it. That inner figure rests on its own small count rather than on the study's total.

The tell is a confident claim about the smallest segment. Check its n before the claim leaves the room.

Synthetic respondents

A synthetic panel can produce a full battery of plausible answers, and for this method that is a specific problem rather than a general one.

Segmentation reads structure out of the covariance between items. A model generating responses produces covariance from the way it was prompted and from its own regularities, and a clustering algorithm will find groups in it.

So the failure mode is not that the segments are weak. It is that they are clean. Synthetic data typically produces a tidier structure than real data, and tidier structure is exactly what a segmentation study is trying to detect.

That makes synthetic respondents unusually risky here. A team that fields a synthetic battery will get a stable-looking solution with a good stability statistic, and the stability will be a property of the generator.

Use them to pilot the battery, which is a genuine use, and to check that items are answerable and discriminating. Do not use them to produce the segmentation itself.

How Sprig supports this method

This is the largest capability concession in this guide and it belongs at the top of the section rather than buried.

Sprig documents no clustering, latent class analysis, k-means, or statistical segmentation of any kind. The central analytical step this guide teaches does not exist in the product, and the clustering happens in a connected AI client or a statistics tool against exported data.

Two naming collisions are worth defusing before they mislead anyone.

Sprig's open-text analysis describes themes regenerating to account for shifts in response clusters, and Sprig has published on thematic clustering elsewhere. That groups responses, not respondents, and it is a different operation from segmentation despite the shared word.

Sprig also uses the word segments in its cross-tab content. Those are pre-existing attribute values you supply, and comparing results across them is segment comparison rather than segment discovery. On attributes specifically, the documentation is explicit. Attributes cannot fire or trigger a study, and in Sprig's own words "they can only Filter."

Filtering on fields you already know is not discovering groups you did not.

What does help

Fielding the battery, including matrix questions for multi-item batteries. Constraint: matrix rows and columns cannot be deleted or reordered after launch, though new ones can be added, so a broken item stays in the instrument for the life of the study.

The MCP connector into a client that can run the clustering, which is the honest path given no native capability. Constraint: it retrieves up to 1,000 responses at a time and the analysis runs in the client rather than in Sprig.

Open-text AI analysis for the qualitative layer that turns a segment into something a design team can picture. Constraint: no published accuracy claim accompanies it. Separately, Sprig's AI insights feed documents its own eligibility thresholds of at least 30 responses, with correlations requiring 50, which apply to that feature rather than to open-text analysis.

Panels for a segmentation study against a frame you do not own, with more than 300 documented targeting attributes. Constraint: no incidence rate, sample minimum or maximum, or fielding turnaround is documented.

The absences that matter for this method

No statistical segmentation. No significance testing anywhere in the product, so profiling a segment against the rest happens outside it. And export is CSV only, capped at 950,000 rows with a one-week link expiry.

Alternatives and adjacent methods

A statistical segmentation without personas where the question is how large each group is and what to build for. This is most of the value, and the personas are optional.

Cross-tabulation across segments you already have where the groups are known and the question is how they differ. That compares segments rather than discovering them, and it is a different and cheaper operation.

Conjoint or MaxDiff where the question is what people will choose or prioritise. A persona encodes attributes and goals, not trade-offs.

Qualitative interviews where the question is what a group's experience is actually like. Persona interview templates exist for exactly this, for B2B and B2C audiences and for unmoderated formats, and the material they produce is what makes a segment legible to a designer.

Product analytics where the question is what people did. Segmentation on stated attitudes and segmentation on observed behaviour are different studies and frequently disagree.

Jobs-to-be-done style interviewing where the question is what people are trying to accomplish. It produces a different kind of grouping, built from situations rather than from people.

Frequently asked questions

What is persona research?

Persona research is the work of understanding who your users are and turning that understanding into artifacts a team can design against. It covers two different objects that are often conflated. A segment is a statistical partition of a population, with a size and a membership rule, that can be shown to be wrong. A persona is a narrative artifact describing one invented person, built to make a design team reason about someone specific. The defensible sequence is to discover segments first and build personas on top of them.

What is the difference between a persona and a segment?

A segment is a claim about a population and a persona is a communication device. A segment has a size, a membership rule and a stability property, and it is falsifiable. A persona is singular, concrete and partly invented, and there is no result that would show it to be wrong. The two are complementary rather than competing, and treating a persona as though it had a segment's properties is the error the critique literature targets.

How many people does a persona actually describe?

Far fewer than teams assume, and the arithmetic is short. For a persona asserting independent attributes each held by half the population, the matching share is one half raised to the number of attributes. Chapman and Milham showed in 2006 that a persona carrying 21 such attributes describes about 0.000048 percent of the population, which against the roughly 281 million United States population of the time is about 134 people. Real attributes correlate, which raises the figure, and real base rates are often well below half, which lowers it. The order of magnitude survives.

How many responses do I need for a segmentation study?

Plan for at least 70 completed responses per segmentation variable, rising toward 100 where segments are likely to overlap. That comes from Dolnicar and colleagues (2014) in the Journal of Travel Research, derived by simulation on artificial data of known structure rather than as a rule of thumb. For a twelve-item battery that means roughly 840 completed responses. Count segmentation variables, not questionnaire length.

Should I use k-means or latent class analysis?

Prefer latent class analysis for survey batteries. Its decisive advantage is that no decisions need be made about the scaling of the observed variables, which removes a choice that silently drives k-means results. It also supports formal model selection through AIC and BIC, and handles mixed measurement levels natively. k-means remains reasonable for a clean single-scale battery, provided you run many restarts and test stability.

How do I choose the number of segments?

Choose k with the variance ratio criterion, also called Calinski-Harabasz, or with BIC or the gap statistic. Do not use the elbow method, which Schubert (2023) argues severely lacks theoretical support and whose apparent elbow moves when you change the scaling of the axes. Use the variance ratio criterion, also called Calinski-Harabasz, or BIC, or the gap statistic. With latent class analysis, BIC does this as part of the model. Then check stability across bootstrap replications before committing to any value.

How do I know my segments are real?

Bootstrap the solution and compare the partitions using the Adjusted Rand Index. A solution that reproduces across resamples is reproducible, which is a weaker and more common property than the segments being naturally present in the data. In one study of 32 tourism datasets, only about 6 percent contained naturally occurring segments, and an audit of 53 published data-driven studies found approximately 92 percent at risk of presenting a random solution.

Does Sprig do the segmentation?

No. Sprig documents no clustering, latent class analysis, or statistical segmentation of any kind. Sprig fields the battery and exports the responses, and the segmentation runs in a connected AI client or a statistics tool. Two terms in the product can mislead: "clustering" there refers to grouping open-text responses into themes, and "segments" in the cross-tab features refers to attribute values you already supply.

Are personas just stereotypes?

Sometimes, and the honest answer is more complicated than yes or no. Turner and Turner (2011) argue that picturing the kind of people you are designing for may commit you to a stereotype, and also note that stereotypes are often disconcertingly accurate. In their own study, students' invented iPhone-user personas were 78.4 percent male against the 75 percent comScore reported for the UK user base in 2009. The practical response is to validate each attribute against data and to visually label the invented parts as invented.

Do personas actually change design decisions?

The evidence is mixed and the studies are small. Matthews and colleagues (2012) interviewed 14 practitioners and found personas used almost exclusively for communication rather than for design, and Friess (2012) observed personas making few appearances in a design team's decision talk. Against that, Salminen and colleagues (2020) found that 34 participants using a persona system located target users faster and far more accurately than the same people using an analytics system built on identical data. The formats appear to answer different questions.

The bottom line

Persona research answers two questions, and separating them is what makes the work defensible.

Which groups exist in your population and how large each one is, which is a segmentation question with a statistical answer that can be checked.

And what a designer should hold in mind while building for one of those groups, which is a communication question that a persona answers well and a table of percentages answers badly.

If your goal is to size an opportunity or decide what to build first, the segmentation is the deliverable and the personas are optional.

If your goal is to get a team reasoning about a specific person instead of an average, build the personas, and build them on a segmentation so that every attribute on them has a number behind it.

Your next study is the validation. Take the personas you already have, count the attributes each one asserts, and check each attribute's prevalence in your own data. It is the step the critique literature has asked for since 2006, almost nobody runs it, and it will tell you in an afternoon whether the personas on your wall describe anyone at all.

Back to top
Solutions
Experience measurementStrategic & foundational discoveryJourney & behavioral researchMarket & consumer insightsConcept & prototype testing
Agents
DesignFieldAnalyzeSynthesize
Deploy
EmailPanelsWeb apps and websitesMobile app
Pricing
Community
EventsBlogGuides
CustomersIntegrationsCompare
Company
About usCareersService agreementPrivacy policyData addendumSystem status
Socials
LinkedInX