Solutions
Experience measurement
Track sentiment and KPIs with AI-driven gap analysis
Strategic & foundational discovery
Uncover market whitespace with AI-led foundational studies
Journey & behavioral research
Connect user actions to motivations across the lifecycle
Market & consumer Insights
Understanding markets, audiences, & opportunity
Concept & prototype testing
Test designs and prototypes with rapid feedback
Agents
Design
Structure rigorous studies
Field
Run adaptive studies at scale
Analyze
Get statistical analysis you can trust
Synthesize
Turn results into research reports
Deploy
Email
Reach external audiences with native deliverability
Panels
Recruit from 300K+ verified participants
Web apps and websites
Embed studies in web experiences
Mobile apps
Run studies in iOS and Android apps
Customers
Community
Events
Join curated gatherings shaping the future of research
Blog
Insights on integrating AI into research craft
Book icon
Guides
Ultimate playbooks for enterprise survey research
Pricing
Sign in
Book a demo
Sign in
Book a demo
Guide

How to Run: Product Concept Test

October 2, 2026

By The Sprig Team

Example H2
Example H3
Example H4
Example H5
Example H6

Introduction

A concept test puts a written or visual description of a product in front of people who might buy it, and measures what they think before you build anything.

Run it monadic by default. Each respondent sees one concept and one concept only, assigned at random, which is what keeps the score interpretable and lets you build norms you can reuse.

Sprig fields this as a single study with randomized assignment across concepts, and it balances the cells for you.

Derive your sample from the gap you need to detect, instead of copying a figure from another page. No methodology source publishes a concept-test sample-size rule with a stated derivation behind it, and the two figures commonly in circulation, 100 and 200 per concept, read one concept's score to roughly plus or minus 9 and plus or minus 7 points respectively.

Neither can separate two concepts less than about 20 and 14 points apart respectively.

Report the full-scale mean and the top-two-box share together, and add an 11-point probability scale when the result will feed any kind of forecast.

The honest statement of what you get: a concept test separates concepts that are far apart, forecasts demand badly, and cannot resolve close calls at realistic sample sizes.

Use it to kill weak concepts and to build a shortlist worth prototyping. Do not use it to pick between two concepts separated by less than your own study's minimum detectable difference, which at realistic sample sizes lands somewhere between twelve and twenty points, and do not use it to size a market.

The next study after this one is a prototype test on whatever survived.

What a concept test actually measures

A concept test measures stated reaction to a description, and typically nothing beyond that. That is a narrower thing than it sounds, and most of the method's failure modes come from forgetting it.

The respondent is not choosing between real options with real money attached. Nothing is at stake, no budget is spent, no alternative is foregone, and the product does not exist. What you are collecting is an attitude toward a paragraph, or toward a prototype if you have one.

That attitude is generally informative, within limits worth knowing. It reliably separates a concept people understand and want from one they do not, and it does so cheaply and before the expensive part starts.

What it does not do, on the available evidence, is forecast.

The field's canonical meta-analysis, Morwitz, Steckel and Gupta in the International Journal of Forecasting in 2007, is the source most often cited as validation for the method, and the number usually quoted comes from their Study 1: across 40 studies, "the average correlation between purchase intentions and behavior across these 40 studies is .49 and the standard deviation is .31."

Their Study 2 is a different sample and a different finding, and it is the one that matters here.

Analysing 60 comparable product tests, they report that "the intent-sales correlation is .751 for the 18 existing products and -.177 for the eight new products."

Read the second half of that sentence again. For new products, the correlation between stated intent and actual sales was slightly negative. A concept test is, by definition, a measurement of a product that does not exist yet.

Morwitz is widely cited as validation for the method. The negative eight-product figure travels much less, and it is the single most important number on this page.

The four designs, briefly

The literature and the vendor pages use four names, and they differ in what each respondent sees.

| Design | What one respondent sees | Gives you | Main cost | |:------------------:|:-----------------------------------------:|:-------------------------------------------------------:|:------------------------------------------------------------------------------:| | Monadic | One concept, assigned at random | An absolute read per concept, comparable across studies | The most sample, since every cell is separate | | Sequential monadic | Several concepts, one after another | More data per respondent | Order and contrast effects, and scores that no longer compare to monadic norms | | Comparative | All concepts together, side by side | A clean relative winner | No absolute read, so it cannot tell you all of them are weak | | Protomonadic | One concept alone, then the rest revealed | Both an absolute and a relative read | The longest instrument, and the absolute read only covers the first concept |

Most guides list these four and then decline to recommend one of them. The section on the design fork below makes the call and explains what the evidence supports.

When to run a concept test

The clearest case is generally a shortlist you need to cut down. You have four or five candidate concepts, you can afford to build one, and you want to eliminate the weak ones before committing engineering time.

It fits when the concepts differ substantially from each other, because that is the situation the instrument can actually resolve.

It fits before a prototype, as the cheap filter deciding what is worth prototyping at all. That sequencing is the method's best use and it is how the guide recommends running it.

And it fits when you need a defensible artifact for a decision that is going to be contested.

A ranked set of scores with a stated margin of error is a better input to that argument than a room full of opinions, even when the scores are imperfect.

Where it does not fit is as a go or no-go gate on a single concept in isolation, because you have nothing to compare the number against.

A 38 percent top-two-box means nothing without either a second concept or a set of prior concepts scored on the same instrument.

A note on what makes the method cheap, since that is the real argument for it. Everything else on the product development path costs engineering time, and a concept test costs a few hundred respondents and a week.

Even a noisy instrument is generally worth running when the alternative is finding out after a quarter of build time, and that asymmetry is why the method survives its own critique literature.

When not to run a concept test

Five situations, each naming the instrument that is actually right.

You need to know whether people will buy it. This is the mandatory one. Morwitz, Steckel and Gupta's negative correlation for new products is the finding, and no design change fixes it, because the problem is that stated intent about a thing that does not exist is weakly related to behaviour toward a thing that does. The better instruments attach a real action: a landing page with a genuine click-through, a waitlist with an email cost, a preorder with a card. Where a forecast is genuinely required, use a calibrated pretest-market model instead of a raw score.

There is a version of this limitation that is easy to miss. Even where a concept test does separate concepts correctly, it separates them on appeal rather than on adoption, and those come apart whenever the barrier to adoption is not appeal.

A concept can win the test and fail in market because switching cost, procurement, habit or an incumbent contract sits between wanting it and having it, and none of those appear anywhere in the instrument.

You need to know what people will pay. Hypothetical bias is well measured and the two meta-analyses disagree instructively. List and Gallet, in Environmental and Resource Economics, found subjects "overstate their preferences by a factor of about 3." Murphy and colleagues, in the same journal, found "the median ratio of hypothetical to actual value of only 1.35, and the distribution has severe positive skewness." Both are right about different statistics, and the skew is the useful part: most studies show modest bias and a few show enormous bias. State the scope limit, because it matters: both are contingent-valuation economics rather than concept testing. For a price, use Gabor-Granger or Van Westendorp.

You need to know which attribute is driving preference. A concept test scores the bundle, so a concept that wins tells you the package worked and not which part of it did. Conjoint decomposes a bundle into attribute-level utilities, and MaxDiff ranks a list of features or benefits.

You are testing something genuinely novel. This is the second mandatory limitation and it is the one most likely to change what you do. Jerry Thomas of Decision Analyst describes the pattern directly: "uniqueness and purchase intent scores tend to be negatively correlated," and "we see many highly unique concepts killed by companies every year because purchase intent scores are marginal." The method is systematically biased against exactly the concepts most worth pursuing, because unfamiliarity reads as low intent. The better instrument is comprehension probing and a prototype, before the concept is scored at all.

You need to know how it performs against the real competitive set. A respondent can express high intent for four competing concepts at once, which is impossible at a checkout. A constant-sum allocation across the consideration set, or a conjoint with a none option, imposes the constraint that a concept test lacks.

One more framing is worth holding before the designs themselves. A concept test is cheap in money and expensive in commitment, because the decision it feeds is usually irreversible within the planning cycle.

That asymmetry is the argument for spending the extra sample on getting the comparison right rather than fielding something fast that returns a ranking nobody should act on.

The design fork

Four designs are in circulation and most guides enumerate all four without choosing. The research supports a choice, and it also supports a real tension that the choice has to resolve rather than hide.

The tension

The case for showing a respondent more than one concept is empirical and strong.

In Study 2 of the same Morwitz, Steckel and Gupta paper, the intent-behaviour correlation was .530 across the 40 products evaluated in comparative mode against .088 across the 20 evaluated in non-comparative mode. Note this is the same Study 2 as the sales finding above, and a different sample from the 40-study meta-analysis in Study 1. Comparative measurement predicted behaviour better, and by a wide margin.

The case against showing more than one is also peer-reviewed.

Friedman and Schillewaert, in the Journal of Marketing Theory and Practice, found two distinct confounds in real-world evaluations of three concepts: concepts are rated higher when shown first, and a concept shown after a strong one is rated lower. Position and contrast are two separate effects, and they compound on each other.

An Ipsos study by Nikolai Reynolds put numbers on the position half, across 22 product tests run between 2010 and 2011 in ten countries and five categories.

The second-position penalty is negative in 21 of the 28 published position comparisons, drawn from a fourteen-study subset, and exceeds 0.7 points in only three of them. Averaging those 28 values gives roughly a quarter of a scale point, which is a calculation on the published table rather than a figure the paper itself reports. The penalty is consistent in direction and typically small in magnitude.

The load-bearing finding from that paper is not the magnitude. The load-bearing part is what rotation actually does to those effects. Reynolds reports that with rotation, order effects are "distributed more evenly and not... being 'reduced', 'cancelled out' or 'avoided'."

Rotation redistributes the bias across concepts rather than removing it, which is the opposite of the usual claim.

There is a third cost that is easy to miss. Sequential monadic produces a suppression effect, so all scores run lower than they would monadically, and an interaction effect, where one exceptional concept depresses the others.

Sequential monadic results are therefore not comparable to monadic norms, which matters a great deal if you intend to build norms at all.

There is a reading of those two findings that reconciles them, and it is worth holding. Comparative measurement predicts behaviour better because a real purchase is comparative: people choose among options, and an instrument that mirrors that structure captures more of what drives the choice.

Monadic measurement is cleaner because it isolates the concept from its neighbours. You are choosing between an instrument that resembles the decision and one that measures the object, and there is no design that does both without cost.

The rule

Monadic is the default. Each respondent sees one concept, assigned at random. Use it when the decision is go or no-go on one concept, when you intend to build norms across studies, or when the concepts are close enough that carryover would swamp the difference. It costs the most sample, because every cell needs its own respondents.

Sequential monadic only when sample or budget genuinely binds, and only with three concessions stated together: scores run lower than monadic, they are not comparable to monadic norms, and rotation redistributes the position bias rather than removing it.

Comparative when the only question is which one wins and no absolute read is needed. It will not tell you whether all of them are bad, which is a real possibility you should want the instrument to be able to surface.

Protomonadic when the decision needs an absolute score and a relative ranking together. Show one concept monadically, collect the absolute score, then reveal the others for a relative read. The term appears in practitioner taxonomies rather than in the peer-reviewed literature, so define it for your stakeholders rather than assuming it.

The ordering above is the recommendation, and it holds for most product teams most of the time. Start monadic, and add a comparative block only when a relative read is genuinely needed and you can afford the design.

The tree

Work down the questions in order until a branch fires, and then stop reading.

Is the question whether this one concept is good enough to build? Yes, run monadic. You need an absolute read and there is nothing to compare against except your own prior norms.

Do you intend to compare these results against concepts you test later? Yes, run monadic. It is the only design whose scores stay comparable across studies, and one sequential monadic wave contaminates the norm set permanently.

Are the concepts close enough that a few points would decide it? Yes, run monadic, and check the sample-size chapter before you field, because the answer may be that no design you can afford will separate them.

Is the only question which one wins, with no need to know whether any of them is any good? Yes, comparative is legitimate and cheaper. Accept that a unanimous winner among four weak concepts looks identical to a winner among four strong ones.

Do you need both an absolute score and a relative ranking? Yes, protomonadic. Field the monadic block first, capture the absolute read, then reveal the remaining concepts.

Is sample or budget the binding constraint, and none of the above applies? Then sequential monadic, with the three concessions stated in the readout instead of omitted: scores run low, they do not compare to monadic norms, and rotation spreads the position bias rather than removing it.

How this fields in Sprig

Sprig runs the monadic design as a single study rather than as parallel ones. Respondents are assigned across concepts at random, and the assignment is balanced, so cells fill at roughly equal rates rather than drifting apart.

The monadic concept test template is the fastest starting point. It describes the mechanism as randomized skip logic, and states that when you upload multiple concepts "each participant will only see one concept, in isolation."

One documentation note worth knowing rather than discovering. Sprig's published survey logic documentation describes two kinds of randomization, order randomization of pages and questions, and response-option randomization.

The random assignment that powers the monadic template is not described there, and the template names the mechanism without explaining it. Sprig's product team confirms that assignment is random and that cells are balanced, which is the basis for this guide's recommendation.

Because assignment happens inside one study, a protomonadic design is fieldable: the monadic block runs first, and a comparative block can follow once the absolute read is captured.

For the comparative block, images can be attached to individual answer options. That is documented for Multiple Choice Single Select and Multi Select, on Surveys, across all plans, accepting JPEG, PNG or WebP files up to 10MB, rendered as a grid.

An image-choice comparative question is therefore buildable, which is worth knowing because it is the natural shape for a visual concept.

Two logic constraints are worth designing around from the very start of the build. Skip Logic "cannot be used to redirect respondents from one question to another on the same page," so each concept's block needs its own page.

And Response Piping disables randomization outright. Sprig's documentation notes that where both piping types are present, randomization is unavailable "because it can't be used alongside Response Piping," so the rule to hold is simpler than it first looks: do not use Response Piping anywhere in a monadic concept test, because it removes the random assignment the whole design rests on.

One practical consequence of the balanced assignment worth planning for. Because Sprig fills cells at roughly equal rates, the study's completion time is governed by the total sample across all concepts rather than by any one cell.

Five concepts take roughly five times as long to field as one at the same per-cell target, which is obvious in arithmetic and routinely missed in project plans.

Writing the concept and the instrument

The concept statement is the stimulus, and differences between statements typically become differences in your results whether or not they reflect differences in the underlying ideas.

The concept statement

A workable concept statement generally has four parts and typically runs between 50 and 100 words. That length is a working convention rather than a sourced finding, and what matters far more is that every concept in the study matches.

The situation, meaning the problem or occasion the product addresses, stated in the customer's own terms. The product, described in plain language with no feature list.

The benefit, meaning what actually changes for the person using it. And the reason to believe, which is whatever makes the benefit credible.

Hold every concept in the test to the same length, the same structure and the same level of specificity.

A concept described in 120 words with three proof points will generally beat one described in 60 words with none, and you will have measured copywriting.

Generally avoid brand names unless brand is the variable, avoid price unless price is the variable, and avoid superlatives everywhere. A statement that sells will score better than one that describes, and the lift is not information about the concept.

Where the concept is visual, prototype embedding is generally a better stimulus than a paragraph, and Sprig supports Figma, Sketch, Framer, Marvel, ProtoPie, Webflow and several others by URL.

The instrument

Five components make up the instrument, and the order they appear in matters.

Comprehension first, before anything evaluative. An open text item asking the respondent to describe the product in their own words. It costs one question and it is the only way to distinguish a concept people dislike from a concept people did not understand.

Purchase intent, the five-point item with the conventional labels: definitely would buy, probably would buy, might or might not buy, probably would not buy, definitely would not buy.

Purchase probability, the eleven-point Juster scale, when the result will feed a forecast. The scale runs from 0 to 10 with each point carrying both a verbal anchor and a numeric gloss, from no chance at the bottom through a fair possibility in the middle to practically certain at the top. Take the exact anchor wording from Juster's own text rather than from a vendor page, because the glosses are what make the scale a probability instead of another intent scale. The anchors are reproduced in the NBER volume chapter, which is freely available. Sprig's Rating Scale supports the eleven points this scale requires.

Uniqueness and relevance, as two short rating items at the end. Uniqueness is the diagnostic that rescues a novel concept from a weak intent score, per the limitation above, and relevance tells you whether you reached the right audience.

Reasons why, an open text item immediately after the intent question while the reasoning is still available to the respondent.

Keep the instrument short, which generally means five or six items. A long instrument completes worse than a short one, and past the fifth or sixth item little typically changes the decision.

Sprig's guidance on writing survey questions and the three rules for effective questions cover the general hygiene on top of this.

A word on the open-text items before the scales, because they carry more of the study's value than their length suggests.

The comprehension item is a gate and the reasons-why item is the explanation, and between them they are typically the difference between a readout that says concept B scored 39 and one that says concept B scored 39 because people could not tell what it replaced. The second version of that sentence is actionable and the first one is not.

Keep both of those items genuinely open rather than offering choices, and let open-text theming do the structuring afterwards. A comprehension check offered as multiple choice measures recognition rather than understanding, and a reasons-why item with preset options measures your hypotheses rather than their reasoning.

Scales and scoring

Whether to report the top-two-box share or the full-scale mean is a live disagreement, and a guide that presents it as settled is misrepresenting the field. Stage it properly and the reader can make the call.

Against top-two-box. Faro and Ohana, writing in Quirk's, argue that dichotomising a continuous trait misrepresents it, that collapsing discards individual differences, and that a top-box focus ignores bottom-box risk entirely. Their recommendation is to stop collapsing the scale and to be deliberate about which statistic gets reported and why.

For top-two-box, conditionally. Jeff Sauro's analysis at MeasuringU found that collapsing an eleven-point scale to two categories "lost 4% of the information." But in his 2012 wireless-carrier work the top-box transformation raised explanatory power substantially, from 8 to 20 percent and from 14 to 27 percent R squared. Note the scope before borrowing it: that work used usability and recommendation measures to predict return rates for wireless carriers, not purchase intent on a concept. And in the same piece he carries the counter-evidence, that Morgan and Rego found top-two-box measures perform slightly worse than average satisfaction.

Against the five-point intent item entirely. Joel Rubinson argues in Greenbook that the item has no budget trade-off and no competitive context, and that only "definitely will not buy" is strongly predictive. Where a budget constraint is the missing ingredient, conjoint supplies one and a concept test cannot.

The probability alternative. Juster found that for automobile purchases, probability explains roughly twice the variance of intentions. Once probability is controlled, the F ratio on the intentions variable falls to 0.8 for new cars, which he reads as "indicating that the intentions variables are essentially random numbers," and to 0.4 for used ones. Those passages sit in the NBER monograph chapter, which is the text to check rather than the journal version. Wright and MacRae independently corroborate that probability scales outperform intent scales.

One practical note on scale direction that causes more damage than it should. Keep every rating item in the instrument running the same way, with the positive pole in the same position.

An instrument mixing a definitely-would-buy-at-the-top intent item with a uniqueness item scaled the other way produces a straight-lining check that flags careful respondents and misses careless ones, and it produces a correlation between the two items whose sign is an artefact of layout.

What to do

Report the full-scale mean and the top-two-box share together, every time, and run the probability scale alongside the intent item whenever the result will feed a forecast. Sprig's guidance on writing studies covers the scale hygiene sitting underneath all of this.

The reason to report both is practical rather than theoretical. They answer different questions, and a reader who presents only top-two-box will be asked for the mean by the first researcher who sees the deck.

Watch carefully for the case where the two numbers disagree with each other. A top-two-box that rises while the mean falls means the distribution is polarising, which is a finding and not an improvement, and it is invisible if you report one number.

Sample size

No methodology source publishes a concept-test sample-size rule with a stated derivation. Not the meta-analyses, not the pretest-market literature, and not any vendor page located.

The two figures in circulation are 100 per concept, which is Sprig's own template guidance, and an unsourced 200.

So derive the number from the comparison you actually need to make. The arithmetic is standard and the two tables below are computed, not copied.

Precision on one concept

How tightly you can state a single concept's top-two-box share, at 95 percent confidence.

| n per cell | Margin at T2B near 35% | Margin at T2B = 50% | |:----------:|:------------------------:|:-------------------:| | 100 | plus or minus 9.3 points | plus or minus 9.8 | | 150 | plus or minus 7.6 | plus or minus 8.0 | | 200 | plus or minus 6.6 | plus or minus 6.9 | | 250 | plus or minus 5.9 | plus or minus 6.2 | | 385 | plus or minus 4.8 | plus or minus 5.0 |

Separating two concepts

How much sample each cell needs to detect a given gap, at 95 percent confidence and 80 percent power.

| Comparison | n per cell | |:---------------:|:----------:| | 30% against 50% | 93 | | 30% against 45% | 163 | | 20% against 30% | 294 | | 30% against 40% | 356 | | 40% against 50% | 388 | | 30% against 35% | 1,377 |

What that means in practice

Read the second table as a warning, not a menu. At 100 per cell you cannot reliably detect a gap smaller than about 20 points. At 200 you cannot detect one smaller than about 14. At 250 per cell the floor comes down to about 12 points.

Most real concept comparisons fall somewhere inside those bands.

Two concepts separated by eight points, which feels like a clear winner in a readout, needs roughly 600 respondents per cell to establish and will not be established by any study anyone reading this is likely to run.

A concept test is a filter, and it is not a discriminator between close options.

A worked example

A team has four concepts and expects the best two to land within about ten points of each other on top-two-box.

Read the second table against that expectation. Detecting a ten-point gap around the 30 to 40 percent range needs roughly 356 per cell. Four concepts at 356 is 1,424 completed responses, before any screening loss.

If that is out of reach, the honest options are to accept the study as a filter that will separate the top two from the bottom two and call the top pair a tie, or to cut to two concepts and spend the same sample on the comparison that matters.

What is not an option is fielding 150 per cell and reporting the leader.

At 150 the minimum detectable difference is about 16 points, so a ten-point gap is indistinguishable from no gap at all, and the ranking will reverse itself on a rerun often enough to be embarrassing.

Adjusting for incidence

If your concept appeals to a narrow slice of the population, the sample that matters is the interested base rather than the total.

Jerry Thomas advises that where a product "establishes a new category or redefines the category," the sample should go "up to 400 or 500, so that you have at least 150 respondents likely to express interest."

That rule assumes an appeal rate between about 30 and 37 percent, which is worth making explicit because it is not general.

The general form of the adjustment is one line of arithmetic. Required sample equals your target interested base divided by the expected appeal rate. At 15 percent appeal, a 150-person interested base needs 1,000 respondents, and a 200-person test leaves the forecast resting on 30 people.

Sprig assigns respondents across concepts in balanced cells, so plan the total as your per-cell figure multiplied by the number of concepts, and use quotas if you also need to hold a demographic composition, noting that quotas are documented as Enterprise-only.

Audience and screening

Screen on the behaviour the concept addresses, not on whether someone already uses your product.

A concept for a new reporting workflow should reach people who do that work, including the ones doing it somewhere else.

Existing customers are a convenient frame and typically a biased one, because they have already selected into your way of doing things.

Use a short screener ahead of the instrument itself, and the demographic screener template is a reasonable starting structure. Note that the documented three-question limit applies to screeners attached to response-based quotas rather than to screening generally.

Keep the screener behavioural and not attitudinal. Asking whether someone is interested in the category selects for people who will score the concept highly, which commonly manufactures the result.

For an audience beyond your own users, Panels is the route and the changelog names concept testing explicitly.

Note the constraints: no incidence rate, no minimum or maximum sample, and no fielding turnaround time is documented, so build in schedule slack, since turnaround commonly varies.

If you are recruiting from inside the product, the in-product prompt template is built for exactly that handoff.

Fielding

A concept test with a written stimulus fields comfortably as a link survey or by email, and either sets a better expectation than an in-product intercept does for a multi-question instrument.

Where the concept is a prototype, the prototype study is the right vehicle, and Sprig's prototype-studies documentation names monadic concept testing among the supported study types.

One constraint will catch you if the prototype came out of Claude Artifacts specifically.

Sprig documents that "Artifacts run in a sandbox that blocks most outgoing calls to third-party services, including the calls Sprig's SDK makes to api.sprig.com." The related guidance to publish rather than preview applies to AI-built prototypes more generally. Test the link end to end before you launch, whichever builder produced it.

Field all concepts simultaneously, inside one study. Fielding them in sequence introduces a time confound on top of everything else, and the balanced assignment only works within a single running study.

One fielding decision that is easy to get wrong in the other direction. Do not stagger concept launches to manage sample cost, even though it looks like a sensible way to spread the spend.

Concepts fielded a fortnight apart differ by whatever moved in that fortnight, and you will have introduced a time confound into the one comparison the study exists to make. Field them all together at once, or else field fewer of them.

Timing

Run a concept test when the decision it feeds is actually open and before anything has been built.

The method's value is entirely in the cheapness of killing a bad concept early, and that value falls to nothing once the engineering has started. A test run after the roadmap is committed produces a number that gets cited selectively.

Leave the study live until the smallest cell clears its target, not until the total does. Balanced assignment makes cells fill at similar rates, and the smallest one is what governs every comparison you will make.

Re-testing a revised concept against the original is legitimate and is one of the better uses of the method. Keep the instrument and the audience identical between the two waves, because those are the only conditions under which the comparison means anything.

Quality control

Three checks are worth running, and the first is specific to this method and frequently skipped.

Comprehension failures. Read the open-text comprehension item before you look at any score. A respondent who cannot describe the product has not evaluated your concept, they have evaluated their confusion. Report that rate per concept, because a concept with a high comprehension-failure rate has not been tested yet and its intent score should not be compared to anything.

Speeders. Time the instrument yourself, take roughly a third of the pilot median as a floor, and review rather than delete automatically. That fraction is a convention, so check what it actually excludes.

Straight-lining across the rating battery, meaning identical answers to intent, uniqueness and relevance. Flag zero within-respondent variance and look at whether those respondents cluster in one cell.

Report the exclusions alongside the result itself, every single time. State the starting count, the number removed, the rule, and the per-cell counts after cleaning, because an unequal cell after exclusions changes the arithmetic in the sample-size chapter.

Bot detection covers the automated end, which matters most for panel and link fielding.

A last note on the instrument, on something most guides omit. Decide before launch how you will handle a respondent who completes the intent item and abandons before the open text, because that pattern is common and the two obvious policies give different denominators.

Counting them in the intent calculation and out of the theming is generally the right call, and the important part is choosing rather than discovering the choice halfway through the analysis.

The core calculation

Three numbers per concept, and one statistical test between concepts.

Top-two-box share is the count of respondents choosing definitely would buy or probably would buy, divided by the count who answered the item at all. Non-responses are excluded from the denominator rather than counted as negative, and the distinction matters when response rates differ across cells.

Full-scale mean treats the five points as 1 through 5 and averages them. That treats an ordinal scale as interval, which is a convention the field generally accepts because it is useful rather than because it is justified, and it is the reason to report the distribution alongside it.

Comprehension rate is the share of respondents whose open-text description matched the concept. It is the gate that both of the other two numbers depend on.

The between-concept test is a two-proportion comparison on the top-two-box shares, and it is the step teams most often skip. A gap that sits inside the margin of error is not a result, however clean it looks on a slide.

A note on the mean that matters when cells differ in size. Compute each concept's mean within its own cell and never pool across concepts, and when you report an overall figure across the study, say which it is.

A weighted average across unequal cells is a different number from the average of the cell means, and the two get confused in decks more often than they should.

Identifying which concept a respondent saw

There is typically no export column recording the assigned branch. The documented export fields cover the study, the visitor, timestamps, question text, responses, themes and attributes, and none of them names the arm.

You recover it from the shape of the data instead. Each concept's block has its own questions, so a respondent assigned to concept B has values in concept B's Q#_Response columns and blanks in the others.

Group on which block is populated and you have your cells.

That is a reader technique and not a product feature, so build it into the analysis deliberately.

If you would rather have a clean identifier, add a hidden item or an attribute per branch at design time, which costs nothing and removes the inference step entirely.

Benchmarks

There is no published, sourced cross-industry benchmark for concept-test purchase intent. Every top-two-box norm table in circulation is uncited, including the ones published by firms that simultaneously argue against using norms.

A search across the published tables returned none with a named study, sample or database behind it.

It is worth being precise about what is missing, because "no benchmark exists" sounds like a gap someone could fill. The problem is structural, and not a gap that anyone could simply fill in. A benchmark needs a stable instrument, a defined population and a disclosed sample, and concept testing has none of the three in common across organisations.

Different scales, different label sets, different designs, different categories, different countries and different recruitment all move the number, and none of them is reported alongside the tables in circulation.

Why cross-industry norms do not work

Clement Dargent of PRS IN VIVO gives the clearest statement of the problem available, organised around five points covering internal and external validity, outdated data, over-reliance on attitudinal measures and inadequate benchmarking. Four of them bear directly on concept-test norms. Categories are simply not comparable to one another.

Countries differ in scale use, because "consumers use scales very differently from one country to another." Sample composition rarely matches between the norm and your study. And norms decay over time, because the population's expectations keep moving.

Add one more that follows from this guide's own scoring chapter.

A norm collected on a five-point intent item is not comparable to one collected on an eleven-point probability scale, and a norm from a sequential monadic design is not comparable to a monadic score at all, because sequential monadic suppresses scores systematically.

What to do instead

Build your own norms, which is Jerry Thomas's recommendation and the only defensible one available. Sprig holds no cross-study numeric comparison, so that archive lives on the export rather than in the platform. Benchmark against your own prior concepts, in the same category, on the same instrument, fielded to the same audience.

A norm generally becomes useful once you have enough prior concepts that a new score can be placed in a distribution rather than compared against a single point.

State how many you have when you report against it, and let the reader judge whether that is enough. Nothing published derives a minimum count, and inventing one would be the same error as the tables above.

It is also a recommendation that sits awkwardly with publishing a benchmark table, which may be why it is rare.

Reading the results

Work through the numbers in a fixed order, because the order is what stops you drawing a conclusion the data does not support.

Comprehension first, always. Any concept below your comprehension threshold is not in the comparison. It has not been tested, and its intent score is measuring confusion. This is the step most frequently skipped, and it is the one that invalidates everything downstream.

Then the gaps, against the margin of error. Compute the minimum detectable difference for your actual cell size before you look at the ranking. If the top two concepts are inside it, you have a tie, and saying so is the correct result rather than a failure of the study.

Then the distribution, not just the top box. A concept with 35 percent top-two-box and 10 percent definitely-would-not is a different proposition from one with 35 percent top-two-box and 30 percent definitely-would-not. The second has a constituency that actively rejects it, which matters for anything with a public launch.

Then uniqueness against intent. A concept scoring low on intent and high on uniqueness is the pattern Thomas describes, and it is the one worth a second look instead of a kill. Novelty depresses stated intent, so an unfamiliar concept is often penalised by the instrument in a way that has nothing to do with its merit.

Then the open text. Themes on the reasons-why item will usually explain a score faster than any cross-tab will, and Sprig's open-text analysis handles the theming.

A worked readout

Four concepts, 200 per cell, comprehension threshold at 70 percent.

| Concept | Comprehension | T2B | Mean | Def. not | Uniqueness | n | |:-------:|:-------------:|:---:|:----:|:--------:|:----------:|:---:| | A | 91% | 44% | 3.4 | 8% | 2.9 | 198 | | B | 88% | 39% | 3.3 | 11% | 3.1 | 203 | | C | 62% | 21% | 2.6 | 24% | 4.2 | 196 | | D | 93% | 19% | 2.4 | 31% | 2.2 | 201 |

Concept C is out of the comparison before any scoring is read, because 62 percent comprehension is below the threshold.

Its low intent score is not evidence against the concept, and note its uniqueness score is the highest on the page, which is the pattern that says rewrite the statement and retest rather than kill it.

Concepts A and B differ by five points on top-two-box. At 200 per cell the minimum detectable difference is about 14, so that is a tie and should be reported as one however much the ordering flatters A.

Concept D is a genuine loser and can be killed on this evidence. High comprehension, lowest intent, the largest definitely-would-not share, and the lowest uniqueness. People understood it perfectly well, and they did not want it.

So the study's actual output is: kill D, rewrite and retest C, and take A and B both forward to a prototype test. That is a useful result and it is not the ranking anyone came in expecting.

What the result licenses

Be precise about the decision the study actually supports, because this is where concept tests get overread.

A clear separation generally licenses killing off the bottom concepts. That is the method working exactly as intended, and it is generally worth a great deal.

A clear separation does not license a demand forecast of any kind. The Morwitz finding on new products means the number typically tells you about relative appeal and close to nothing about volume.

A narrow separation licenses nothing at all beyond taking both forward. Take both concepts forward to a prototype test and let a better instrument decide.

And a uniformly low set of scores licenses going back to discovery, which is a legitimate and underused outcome.

The instrument can detect that every concept is weak, and a team that treats the highest of five poor scores as a winner has misused it.

Running the analysis with Claude

The arithmetic is simple enough to do in a spreadsheet and easy enough to get wrong that it is worth automating with a check on top. Sprig documents no significance testing, so the comparison between concepts happens outside the platform either way.

Pull the responses through the Claude integration or work from a CSV export. Replace the bracketed placeholders before running each step, and defaults are given where a sensible one exists.

Step 1. Load and identify the cells

You are analyzing a monadic concept test exported from Sprig.

WHAT THIS STEP DOES: reads the file, works out which concept each respondent
was assigned to, and reports the cell structure back to me for confirmation
before any scoring runs.
WHAT IT RETURNS: a table of concept, inferred from the populated question
block, with the respondent count in each cell.

File: [PATH TO EXPORT FILE]
Number of concepts: [N, default 3]

There is no column naming the assigned concept. Each concept has its own block
of questions, so a respondent assigned to a concept has values in that block's
Q#_Response columns and blanks in the other blocks. Group respondents on which
block is populated.

Use code to calculate this, not estimation. Do not infer the structure from a
sample of rows.

Report any respondent populated in more than one block, or in none, as an
anomaly rather than assigning them. Do not guess.

Save the cell map as concept_cells.csv.

Step 2. Clean and exclude

WHAT THIS STEP DOES: removes responses that cannot contribute, and reports
exactly what was removed.
WHAT IT RETURNS: a cleaned dataset plus an exclusion log.

Rules:
1. Exclude a respondent from the intent calculation when the intent item is
   blank. Blank means empty, null, or a non-response code. Blank is NOT the
   lowest scale point. The scale runs [SCALE MINIMUM, default 1] to
   [SCALE MAXIMUM, default 5], and the minimum is a real rating that counts as
   a value below any threshold I set later. Any value outside that range is a
   coding error, so report it and do not treat it as a rating.
2. Flag respondents who completed in under [SPEED FLOOR IN SECONDS]. If I leave
   that empty, derive it as a third of the median completion time and tell me
   the number you used.
3. Flag respondents whose ratings across intent, uniqueness and relevance have
   zero variance.

Use code to calculate this, not estimation.

Do not delete flagged respondents. Write two datasets, one including them and
one excluding them, and carry both through every later step.

Save as concept_clean_all.csv, concept_clean_strict.csv, and
concept_exclusions.csv.

Step 3. Score the comprehension gate

WHAT THIS STEP DOES: decides which concepts are eligible for comparison.
WHAT IT RETURNS: a comprehension rate per concept and a pass or fail.

For each respondent, judge whether their open-text description of the product
matches the concept they were shown. Report the share matching, per concept.

Concept descriptions: [PASTE EACH CONCEPT STATEMENT, LABELLED]
Comprehension threshold: [THRESHOLD, default 70 percent]. No source derives this
  figure, so set it deliberately for your own study rather than taking the default.

Use code to count, and your own judgement to classify. Show me twenty
classified examples per concept so I can check your judgement before you
report a rate.

Any concept below the threshold is reported as not tested, and is excluded from
every comparison in later steps. Say so explicitly rather than quietly dropping
it.

Save as concept_comprehension.csv.

Step 4. Compute the scores

WHAT THIS STEP DOES: computes the headline numbers per concept.
WHAT IT RETURNS: one row per concept with every number the readout needs.

Per concept, compute: top-two-box share, full-scale mean, the full five-point
distribution, the definitely-would-not share, mean uniqueness, mean relevance,
and n.

Top-two-box is the count choosing the top two points divided by the count who
answered the item. Non-responses are excluded from the denominator, never
counted as negative.

Use code to calculate this, not estimation. Do not compute any value by
reasoning about it in prose.

Save as concept_scores.csv, and run it on both datasets from Step 2.

Step 5. Test the differences

WHAT THIS STEP DOES: decides which gaps between concepts are real.
WHAT IT RETURNS: a pairwise comparison table with a verdict per pair.

For every pair of eligible concepts, run a two-proportion z test on the
top-two-box shares at [CONFIDENCE, default 95 percent], and report the
confidence interval on the difference.

Then compute, from the actual cell sizes, the minimum detectable difference at
80 percent power, and report it.

Use code to calculate this, not estimation. Do not eyeball whether a gap is
significant.

Verify the result a second way, by bootstrapping the difference in top-two-box
between each pair with [ITERATIONS, default 2000] resamples of the respondent
rows, and compare the bootstrap interval to the analytic one. These are
different methods and they should broadly agree. If any pair disagrees in
direction, or if the intervals differ by more than [TOLERANCE, default 3
points], say that you cannot reconcile them and show me the rows.

Do not resolve a mismatch by picking the answer that looks more reasonable.

Save as concept_comparisons.csv.

Step 6. Theme the reasons why

WHAT THIS STEP DOES: explains the scores.
WHAT IT RETURNS: themes per concept, with counts and example verbatims.

Theme the reasons-why open text separately within each concept, then report
which themes appear across concepts and which are specific to one.

Minimum respondents before a theme is reportable: derive it as
[THEME FLOOR PERCENT, default 5] percent of that concept's cell size, and state
the number you used. That percentage is a convention, not a derived threshold. Do not report a theme below it, and do not merge small
themes to clear the floor.

Use code to count theme frequencies, not estimation.

Report an Other or unclear bucket and its size. If it exceeds a quarter of
responses for any concept, say the theming did not work rather than presenting
the themes.

Save as concept_themes.csv.

Step 7. Challenge the readout

WHAT THIS STEP DOES: attacks my interpretation before anyone else does.
WHAT IT RETURNS: specific objections with the rows supporting them.

Here is the scores table, the comparison table and the interpretation I plan to
present: [PASTE YOUR INTERPRETATION].

Use code to check any claim that can be checked numerically.

Find every claim the data does not support. Check in particular:
1. Any winner declared on a gap inside the minimum detectable difference.
2. Any concept compared while below the comprehension threshold.
3. Any statement about volume, demand or market size, which this instrument
   cannot support at all.
4. Any concept with low intent and high uniqueness that I have described as a
   loser.
5. Any cell whose n after exclusions differs from the others by more than
   [CELL IMBALANCE, default 10 percent].

Quote the specific claim and give the contradicting rows.

If the interpretation holds, say so plainly and say what you checked. Do not
manufacture objections to appear thorough, and do not soften a real one.

Save as concept_challenge.md.

Sprig's note on prompting for research covers why the preamble and the escape hatch carry as much weight as the instruction.

For cutting results by segment, the cross-tab walkthrough is the companion piece, and theming mechanics at depth belong to the advanced version.

The critique you should know

Four objections, and a defence that is generally stronger than they make it sound. Every widely used research instrument eventually accumulates a literature like this, and the case against NPS is the version most product teams have already argued through.

Take these in order of how much they should change your behaviour, which is not the same as how damaging they sound.

The forecasting objection

Covered at the top of this guide and restated here because it is the one a senior stakeholder will typically raise.

Morwitz, Steckel and Gupta's Study 2 found an intent-sales correlation of .751 for 18 existing products and -.177 for eight new products.

Note the sample sizes in that sentence before leaning on it. Eight products is generally a small basis for a strong claim, and the authors present it as such.

What it establishes is not that concept tests are worthless but that the validation evidence people cite for them comes overwhelmingly from existing products, which is not the case a concept test is ever used for.

The self-generated validity objection

This is the sharpest one available, because it undercuts the evidence rather than the method.

Chandon, Morwitz and Reinartz, in the Journal of Marketing, found that the intent-behaviour correlation is "58% greater among surveyed consumers than it is among similar nonsurveyed consumers."

Asking the question appears to change the behaviour.

Which means every validation study measuring how well intent predicts behaviour is measuring a population that was made more predictable by being asked, and the published correlations are inflated relative to the population you actually care about.

The wrong-battery objection

Kato and colleagues, in a conference proceedings chapter rather than a journal article, tested ten evaluation indices against actual sales across Japanese cosmetics and food categories. Comprehension scored highest among the indices and predicted sales worst.

Suitability scored lowest among the indices and correlated most strongly with sales. Only two of the ten reached significance at the 5 percent level, suitability and empathy.

If that generalises, and it comes from one recent study in two categories so treat it carefully, the standard battery may be measuring the things that are easiest to ask rather than the things that predict. A voice of customer study is generally the cheaper way to find out which attributes your market actually talks about before you commit a battery to them.

The failure-rate statistic that props up the whole category

Worth a paragraph because you will frequently meet it and because it is wrong.

The claim that roughly 95 percent of the 30,000 new products launched each year fail, usually attributed to Clayton Christensen, circulates on second-hand sourcing. Where it is footnoted at all, the footnote leads to another secondary retelling, and no primary Christensen publication containing that figure could be found.

The peer-reviewed position is materially different from it. Castellion and Markham, in the Journal of Product Innovation Management, call the common assertion that 80 to 90 percent of products fail "an 'urban legend'" and report that "the actual product failure rate is around 40%," with nineteen peer-reviewed studies between 1945 and 2004 putting it in a range of 30 to 49 percent.

The same OpenStax passage that repeats the 95 percent claim also reports a PDMA study ranging from 35 percent for health care to 49 percent for consumer goods, on the same footnote, which is worth noting as a sign of how loosely the figures travel together.

The method does not need the inflated number, and repeating it typically costs credibility with anyone who checks.

Before the defence, one more thing the critique does not say. None of these objections is an argument that concept tests produce random output. They are arguments about what that output actually means. A concept test reliably measures stated reaction to a description, and stated reaction to a description is a real quantity that varies in interpretable ways.

The objections are all about the distance between that quantity and the commercial question, which is why the guide's recommendations are about how far to read the result rather than whether to run the study.

What survives

The defence is real, and it is unusually specific for this literature.

Urban and Katz validated the ASSESSOR pretest-market model against actual test-market results and reported a standard deviation between predicted and actual share of 1.99 points before adjustment and 1.12 after.

State that as a distribution and not an accuracy guarantee. The adjusted 1.12-point standard deviation works out to roughly plus or minus 2.2 points at 95 percent, which is a very different claim from one-point accuracy on any single forecast. It remains the strongest evidence that the method works when it is modelled instead of read raw.

Armstrong, Morwitz and Kumar found four intentions-based methods each beat extrapolation of past sales across four datasets, with combination cutting error by roughly a third against extrapolation. Note the scope limit, which is in the title: existing products.

And Wright and MacRae, across two meta-analyses, found that "purchase intention scales are empirically unbiased," that "the variability is much less than previously assumed," and that "purchase probability scales performed even better than purchase intention scales."

That is the strongest source in the purchase-intent scale literature and it contradicts the widespread practitioner claim that intent scales systematically overstate.

So here is the position this guide lands on. The instrument is empirically unbiased as a scale and it is noisy. Note that this is a different sense of bias from the one in the limitations chapter: the scale does not systematically over or understate, and it does systematically penalise unfamiliar concepts, and both statements are true at once. Modelled properly against a calibrated benchmark it forecasts usefully.

Read raw off a five-point item at 200 per cell it typically separates far-apart concepts and nothing else. Both halves of that are true and most content commonly carries only one.

One further observation about the defence, because it changes how the critique should be used in a readout. Most of the objections above are about the gap between stated reaction and commercial outcome, and none of them is reduced by running a bigger study.

Sample size fixes the noise and does nothing about the bias, which is why the guide's recommendations split cleanly into two kinds: derive your sample to handle the noise, and limit your claims to handle the bias.

Common mistakes

Declaring a winner on a gap inside the margin of error. The single most common misuse, and the reason the sample-size chapter exists. Compute the minimum detectable difference before you look at the ranking, not after.

Skipping comprehension. A concept nobody understood produces a low intent score that looks exactly like a concept nobody wanted. Without the comprehension item you cannot tell them apart, and you will often kill the wrong one.

Killing a novel concept on intent alone. Uniqueness and intent are negatively correlated, so the instrument typically penalises unfamiliarity. Read the two together or you will systematically select for the familiar.

Writing concepts of different lengths. The longer, more detailed, more polished statement generally wins. You have measured your copywriter.

Using sequential monadic and comparing to monadic norms. Sequential monadic suppresses scores. The comparison is invalid and the direction of the error makes every concept look worse than it is.

Treating rotation as a fix for order effects. Rotation distributes the bias more evenly across concepts. It does not remove it, and the Ipsos work says so directly.

Reporting only top-two-box. You will usually be asked for the mean, and a polarising distribution is invisible without it.

Reading a concept test as a demand forecast. The instrument does not support it for new products and the literature is clear. Size the market some other way.

Fielding before the decision is open. A test run after the roadmap is committed typically becomes a document to cite selectively instead of an input.

Synthetic respondents

Readers are already asking, and the two relevant studies are not actually in conflict once you see what each measured.

Maier and colleagues report that a semantic-similarity method reproduced human purchase-intent ratings across 57 personal-care concept surveys, reaching 90 percent of the human test-retest ceiling on their headline result, and 92 percent for one model in an ablation. Two flags belong in the same breath rather than a footnote.

It is an arXiv preprint and remains one, and the author list is PyMC Labs together with Colgate-Palmolive, meaning a vendor and its client validating the vendor's method on the client's single-category data. The paper also reports its own failure: removing demographics from the prompt collapsed the correlation from 92 percent to 50 percent.

Bisbee and colleagues, peer-reviewed in Political Analysis, compared 7,530 real respondents against 3.6 million synthetic ones. Aggregate means corresponded closely.

But 48 percent of regression coefficients differed significantly from the human data and 32 percent of those flipped sign, with artificially deflated variance that breaks any power calculation built on it.

The apparent contradiction dissolves once you notice the level each study works at. Maier measured whether simulated ratings reproduce the aggregate distribution of human ratings on one category of product, and found they largely do.

Bisbee measured whether relationships between variables survive, and found that roughly half do not. A concept test's headline number is an aggregate and its decisions are comparisons, so the first result is encouraging about the number and the second is disqualifying for the decision.

Both findings point the same way in the end. Aggregate means generally track, and subgroup structure does not.

The usable position: synthetic respondents may often help screen a long concept list down to a short one before you field anything.

They are not usable for segment reads, for significance testing, or for any result that has to survive review. And a concept test's whole value is the comparison between cells, which is exactly the thing the synthetic data gets wrong.

The pre-launch checklist

Work through this list in order, from start to finish, before you launch anything. Most of it cannot be fixed after launch, so the order matters.

Confirm every concept statement is the same structure, the same level of specificity and within roughly a tenth of the others on word count.

Confirm no concept statement contains a brand name, a price or a superlative unless that variable is the thing you are testing.

Confirm the comprehension item comes before anything evaluative.

Confirm the intent item uses the five conventional labels and that the scale direction is the same on every rating item in the instrument.

Add the probability scale if the result will feed a forecast, and confirm the eleven points are configured with their anchors rather than bare numbers.

Compute the minimum detectable difference for your planned cell size, and write it in the study plan before you field.

If it is larger than the gap you expect to find, either raise the sample or accept that the study is a filter and not a comparison.

Derive the total sample as per-cell multiplied by the number of concepts, then divide by the expected appeal rate if incidence is low.

Confirm each concept's questions sit on their own page, because skip logic cannot route between questions on the same page.

Confirm Response Piping is not used anywhere in the study, because it disables randomization and therefore the random assignment across concepts.

Add a hidden branch identifier if you would rather not infer the cell from populated columns at analysis time.

Test the full path end to end, including any prototype link, on a real device. Published rather than previewed, if the prototype came from an AI builder.

Field to a handful of colleagues first and record the median completion time, which becomes your speeder floor, which is a convention rather than a rule.

Then launch, and leave it live until the smallest cell clears its target rather than until the total does.

How Sprig supports concept testing

The monadic design runs natively as a single study, with respondents assigned across concepts at random and cells kept balanced. The monadic concept test template is the starting point and the concept validation template is the adjacent one.

Prototype embedding is the deepest documented capability here and it is directly on-method.

Figma, Adobe XD, Sketch, Miro, Marvel, Webflow, JustInMind, ProtoPie, Framer and others paste in by URL, and there is a documented path for AI-built prototypes covering Lovable, Claude Artifacts, Figma Make, Replit and Vercel v0. The Figma integration is typically the most common route.

Images attach to individual answer options on multiple choice questions, across all plans, which makes a visual comparative block buildable without leaving the instrument. That feature is documented for Surveys, so it is not available inside a prototype study.

The Rating Scale carries the five-point intent item and the eleven-point probability scale, with configurable labels and endpoint descriptions.

Open-text analysis themes the reasons-why, and AI Study Reports will summarise a study once it holds at least ten responses.

A note on attribution that is easy to get wrong when writing this up internally. Theming, AI Study Reports and the Explorer are Sprig capabilities.

The conjoint, MaxDiff and Gabor-Granger analysis recipes that sit alongside them in the documentation are executed by a connected AI client rather than by Sprig, even though they live in the same section of the docs. Describing the analysis as something Sprig performs would be a factual error, and it is the kind a reviewer catches.

What happens outside Sprig

Significance testing of any kind happens outside the platform. Sprig documents none anywhere, so the two-proportion comparison between concepts, the minimum detectable difference and any confidence interval are computed on the export or in the connected client.

The branch identifier, as covered in the analysis chapter, since no export field records it.

Cross-study comparison against your own prior concepts, which is how the benchmarks chapter says to build norms. Export is CSV only, capped at 950,000 rows, with a link that expires after a week.

One availability note worth stating plainly rather than burying. Panels, email delivery, Conjoint, MaxDiff and Rank Order are all Enterprise features, so a reader on another plan generally fields to their own audience.

One more capability note that shapes what the readout can contain. Because Sprig holds no cross-study numeric comparison, building norms from your own prior concepts means keeping those results somewhere outside the platform, in a spreadsheet or a warehouse, with the instrument version recorded alongside each score.

That sounds like overhead and it is the only route to a benchmark you can defend, so set it up on the first study rather than on the fourth.

Alternatives

Conjoint analysis when the question is which attributes drive preference rather than which whole concept wins, and when you need to simulate a configuration you have not built.

MaxDiff when the list is features or benefits rather than concepts, and you need an ordering across more items than a concept test can carry.

Gabor-Granger when the question is price.

Message testing when the variable is how you describe the thing rather than what the thing is. The two are frequently confused, and a concept test with differently worded statements is an accidental message test.

A prototype test once the concept has survived the filter. That is the honest next step after this method rather than an alternative to it, and it is where a shortlisted concept should go.

One last thing about the checklist that is worth saying plainly. Almost every item on it is cheap before launch and expensive or impossible afterwards, which is an unusual property and the reason the list exists at all. A concept statement cannot be rewritten mid-field without invalidating the responses already collected, a cell size cannot be raised retroactively, and a missing comprehension item cannot be added to data that has already been gathered. Spend the extra hour here.

Frequently asked questions

How many respondents do I need for a concept test?

No source publishes a rule with a derivation behind it, so derive your own. At 100 per concept you can state one concept's top-two-box share to about plus or minus 9 points and cannot detect a gap smaller than roughly 20.

At 200 the figures are about plus or minus 7 and 14. Decide the gap you need to detect, then read the sample off that.

Is monadic or sequential monadic better?

Monadic, unless sample size is genuinely the binding constraint. Sequential monadic typically suppresses all scores, breaks comparability with monadic norms, and rotation redistributes position bias rather than removing it.

Can each respondent see only one concept in Sprig?

Yes. The monadic template assigns each respondent to one concept at random within a single study, and the cells stay balanced.

What is a good purchase intent score?

There is no sourced cross-industry benchmark, and every norm table in circulation is uncited. Build your own norms from your own prior concepts on the same instrument, and compare against those. The concept validation template is a reasonable place to standardise the instrument you reuse.

Should I report top-two-box or the mean?

Report both of them, every single time. They answer different questions, and a top-two-box that rises while the mean falls means the distribution is polarising rather than improving.

Does purchase intent predict sales?

Weakly overall, and for new products barely at all. The canonical meta-analysis found an intent-sales correlation of .751 for existing products and -.177 for new ones. Use the method to shortlist, not to forecast.

How many concepts can I test at once?

As many concepts as you can afford to sample properly. Each concept needs its own full cell, so five concepts at 200 per cell is a 1,000-respondent study before any incidence adjustment, and quotas, which are Enterprise-only, are how you hold composition across it.

How long should the instrument be?

Keep the instrument short and hold the line on it under pressure. Comprehension, intent, probability if you need it, uniqueness, relevance, reasons why. Anything past that rarely changes the decision and costs completion.

What if the concept is visual rather than written?

Use a prototype study and embed the prototype file directly in it. That is the better stimulus for anything with an interface, and it is Sprig's deepest documented capability for this method.

What if all the concepts score badly?

That is itself a finding, and it is a perfectly legitimate outcome. Going back to discovery is the right response, and treating the highest of five poor scores as a winner is the misuse the instrument most often enables.

Can I use the same concept test to set a price?

No, and the two questions should not share an instrument. A concept test measures reaction to a description, and a price inside that description changes what is being evaluated.

Use Gabor-Granger or Van Westendorp for the price question and keep the two studies separate.

Do I need a control concept?

A control concept generally helps more than teams expect. Including an existing product, or a concept you already shipped, gives every other score something to be read against and partially substitutes for the norms you do not have yet.

Can I run a concept test on my own users?

Yes, though it is worth knowing what it costs you. Existing users have selected into your way of working, so they will generally rate an extension of it higher than the market would.

For anything genuinely new, field to an audience that has not already chosen you.

How do I stop the concept statements from becoming a copywriting test?

Fix the structure before you write any of them, without exception. Same four parts, same word count within about ten percent, same number of proof points, no superlatives anywhere. Then have someone who did not write them check for tone drift.

One closing note on sequencing that applies whatever you find. A concept test is the second cheapest study in the sequence, behind desk research and ahead of everything involving a build. Running it before you have talked to anyone about the problem produces concepts written from inside the building, and those tend to score respectably against each other while missing the thing customers would actually have asked for. The order that works is discovery, then concepts, then this test, then a prototype.

The bottom line

A concept test is a filter, and typically nothing more. It separates concepts that are far apart, cheaply and early, and that is worth doing before anyone writes code.

It is not a forecast, and the evidence on that point is consistent across the literature.

The evidence that intent predicts behaviour comes overwhelmingly from existing products, the one analysis covering new products found a slightly negative correlation, and the act of asking inflates the relationship by more than half.

Run it monadic, inside one study, with balanced cells across the concepts. Write every concept to the same length and structure. Put the comprehension item before the intent item, without exception.

Report the mean and the top-two-box together, and add the probability scale if anyone downstream will treat the number as a forecast.

Compute your minimum detectable difference before you field, and say out loud that gaps below it are ties. Most concept comparisons that feel decisive in a readout are typically ties.

Then take whatever survived to a prototype test, which is the instrument that answers the question a concept test only gestures at. That is the honest next step, and it is also the cheapest one available to you at that point.

Back to top
Solutions
Experience measurementStrategic & foundational discoveryJourney & behavioral researchMarket & consumer insightsConcept & prototype testing
Agents
DesignFieldAnalyzeSynthesize
Deploy
EmailPanelsWeb apps and websitesMobile app
Pricing
Community
EventsBlogGuides
CustomersIntegrationsCompare
Company
About usCareersService agreementPrivacy policyData addendumSystem status
Socials
LinkedInX