The short answer
A message testing survey shows candidate messages to a sample of your target audience and measures reaction on a fixed scale, so you can rank the messages before you spend money putting one of them in market. Run it in four moves. Decide how many messages you are testing, because the count determines the design. Assign respondents to one message each, which is the monadic design, and size each cell from the difference you need to detect, not from a habit.
Field the same instrument every time, including an unchanged control message. Then test the difference between cells with a two-proportion test and a correction for the number of comparisons you ran.
Two things separate a message test that holds up from one that does not.
The first is power, because the conventional cell size of 100 respondents detects roughly a 19-point difference and nothing smaller, while real message differences typically run 5 to 10 points.
The second is scope. A message test is a screening instrument, not a forecast of in-market performance. Screen with the survey, decide with the experiment.
What a message testing survey is
A message testing survey is a quantitative study that exposes respondents to one or more written messages and captures their reaction on closed scales, producing a ranking you can defend with an arithmetic test.
The message is the stimulus. The scale is the measurement. The comparison between messages is the output.
The method belongs to the same family as concept testing and claims testing, and practitioners often use the three terms loosely. The distinction that matters is what is being varied.
Concept testing varies the product idea. Claims testing varies a single factual assertion about the product. Message testing varies how a value proposition is expressed while holding the underlying offer constant.
What gets measured
A message test typically carries three layers of measurement, and confusing them is the most common design error in the method.
The primary measure is a single scale that decides the winner. Purchase intent, likelihood to click, or likelihood to sign up are typically the choices, and the study reports it as a top-two-box proportion.
The diagnostic battery is a short set of scales explaining why a message scored the way it did. Relevance, believability, uniqueness and clarity are generally the four that earn their space.
The open-text probe is one question asking what the message made the respondent think. It produces the language you will reuse and it frequently surfaces the objection nobody anticipated.
What the output looks like
The output of a message test is a ranked table of messages with a top-two-box proportion for each, a stated confidence interval, and an explicit statement of which differences are large enough to act on.
That last part is the half most reports generally omit. A ranking without a minimum detectable difference beside it is a ranking of noise.
When to use a message testing survey
Use a message testing survey when you have more candidate messages than you can afford to put in market, the messages describe the same underlying offer, and you need to cut the set down before committing budget.
The method earns its cost in four situations.
- Narrowing a long list before creative production
- Choosing a headline for a paid campaign
- Testing positioning for a new audience segment
- Screening claims before legal review
The conditions that make the method work
A message test generally works when the messages are genuinely comparable, which means they describe the same product, at the same level of abstraction, in roughly the same number of words.
Messages that differ in length, specificity or format are not being compared on their content. They are being compared on their format, and the format will often win.
The method also needs an audience that could plausibly buy the thing. Testing enterprise positioning on a consumer panel produces a clean number about the wrong population, which is more dangerous than no number.
The decision the study supports
A message test supports a narrowing decision, not a launch decision. The output is a shortlist, and the shortlist is what you take into the next study.
It is a cheaper way of buying information a live test would eventually give you, and it is not a substitute for that test.
The practical test of whether you need this study is simple. If you can afford to put every candidate message in market and measure it, do that instead, because the experiment answers a harder question than the survey does.
If you cannot, the survey is how you decide which ones get that chance.
Teams that treat the survey as the decision and skip the live test are making the most expensive mistake commonly made in this method.
When not to use a message testing survey
A message testing survey cannot tell you what people will actually do, how large the real-world effect is, whether the message survives a competitive feed, whether the message or the execution moved the number, how the message wears out, or why it failed.
Each of those needs a different instrument, and each one is named below.
The method cannot tell you what people will do
Stated intent overpredicts behavior by roughly half, and the size of the gap is measured, not asserted.
Webb and Sheeran (2006), a meta-analysis of 47 experimental tests in Psychological Bulletin, found that a medium-to-large change in intention of d = 0.66 produced only a small-to-medium change in behavior of d = 0.36.
The better instrument is typically a randomized holdout or a geo-split incrementality test, which measures behavior, not a report of intended behavior. It is the reason the method is a screen and not a forecast.
The method cannot size the real-world effect
Survey-based estimates of advertising effect are frequently wrong by a large multiple.
Gordon, Zettelmeyer, Bhargava and Chapsky (2019), reporting 15 United States experiments covering 500 million user-experiment observations in Marketing Science, found that the bias can be large, and that in half of the studies the estimated percentage increase in purchase outcomes was off by a factor of three across all methods.
The better instrument is generally a platform-level randomized lift test. A survey is structurally unable to estimate incremental revenue.
The method cannot tell you whether the message survives real exposure
Survey exposure is forced and real exposure is not, and the difference is large enough to change conclusions.
Schmalzle and colleagues (2023), in PLOS ONE, note that micro-level research takes place under highly controlled laboratory conditions with forced message exposure, and largely ignores how recipients attend selectively to messages under more natural conditions.
Their own result is the useful part.
Recall fell from 6.45 of 20 billboards under free viewing to 2.95 once attentional demand was added, and billboards that were never looked at are practically never brought up during the post-drive memory test.
The better instrument is typically attention measurement under competitive conditions, or an in-feed split test with organic scroll.
The method cannot separate message from execution
A single-execution message test confounds what you said with how it was rendered.
Chen, Yang and Smith (2016), in the Journal of the Academy of Marketing Science, decompose creativity into divergence, which is execution, and relevance, which is message, and find a three-way interaction with repetition across six dependent variables.
Their finding is that the classic inverted U-shape is observed only for ads with low divergence and relevance, while creative ads wear in immediately and show little sign of wearing out.
The better instrument is generally a factorial design crossing message against execution, or a conjoint study with execution as an attribute.
The method cannot tell you about wear-out
A message test is a single-exposure snapshot and wear-out is a function of repetition, so the design has no way to observe it. The same Chen, Yang and Smith result applies directly.
The better instrument is typically in-market frequency and decay analysis, always-on tracking, or a marketing-mix model carrying an adstock term.
Pechmann and Stewart (1988) is the canonical review of advertising repetition effects and is cited here as the reference and not quoted.
The method cannot tell you why a message failed
A rating scale measures whether a message landed, not why it did not. An open-text probe gets closer without getting there, because a respondent typing two sentences is not being asked follow-up questions.
The better instrument is generally a moderated or artificial-intelligence-moderated depth interview. Those are qualitative instruments in a different category, answering a different question from few participants, and they are not survey tools at all.
A guide that lists them alongside survey platforms is making a category error that will cost the reader real money in tooling.
Study design, and the fork that decides everything
The design fork in a message test is how many messages each respondent sees. Monadic shows one per respondent, sequential monadic shows several in rotation, and comparative shows them side by side.
The count of messages you are testing typically decides which one you can afford.
The monadic design, and why it is the default
A monadic design assigns each respondent to exactly one message, so each message gets its own independent cell drawn from a separate sample of people.
The argument for monadic is that it mirrors typical real exposure, because a customer encounters one headline and not five in a grid.
The cost is arithmetic. Six messages at an adequate cell size means fielding six full samples.
The evidence against the reflexive monadic argument
The published predictive-validity evidence does not support the claim that monadic collection predicts purchasing better than comparative collection, and it points the other way. Morwitz, Steckel and Gupta (2007), in the International Journal of Forecasting, found that purchase intentions correlate more strongly with actual purchasing when collected in a comparative mode than when collected monadically. The reflexive argument that monadic is more predictive because it is more realistic has no published support behind it.
That meta-analysis identified six factors associated with a stronger intent-purchase correlation, and the sixth is the one that matters here.
Verbatim from the abstract, intentions are more correlated with purchase when purchase intentions are collected in a comparative mode than when they are collected monadically.
The counterargument you will commonly hear is that comparative exposure inflates discrimination and produces choices respondents would not make.
No peer-reviewed source demonstrating that was found, and searches returned vendor blog assertions with no study behind them.
State it plainly. The claim that comparative exaggerates differences is practitioner folklore with no published support, and the one meta-analysis that regressed intent-purchase correlation on design features found comparative collection correlated better with purchasing.
Hedge it honestly too. That paper's internal statistics sit behind a paywall and were not read for this guide.
The direction of the finding is quoted from the abstract. No number from it appears in this guide, and you should treat any source that attaches one with suspicion.
The sequential monadic design, and why its reputation is better than its evidence
Sequential monadic shows each respondent several messages one at a time, rotating the order, which buys efficiency by reusing respondents across cells. The published evidence on whether that efficiency is real is weaker than practitioners generally assume.
The most specific document available is a vendor white paper rather than a peer-reviewed study, and that qualification belongs in the same sentence as the finding.
Reynolds (2013), for Ipsos InnoQuest, reported across 22 two-product tests in ten countries that arithmetic means of second position ratings are generally lower, which is a systematic second-position penalty.
The same paper states that sequential design does not offer demonstrable improvement in discrimination or statistical sensitivity when examining data in total or by first position.
The counterintuitive part is that the discrimination which does appear lives in the second position, precisely the position contaminated by carryover from the first.
One peer-reviewed paper addresses this question directly. Friedman and Schillewaert (2012), "Order and Quality Effects in Sequential Monadic Concept Testing," in the Journal of Marketing Theory and Practice, is the right citation to chase.
No abstract is carried anywhere retrievable and every publisher route refused access, so this guide cites it as addressing the question and makes no claim about what it concluded.
The comparative design, and when it is the right answer
A comparative design puts messages side by side and asks the respondent to choose or rank, which typically produces a clean ordering from far fewer respondents.
Use comparative when the decision is genuinely a choice between alternatives you control, and when you accept a preference order instead of an absolute level.
It tells you message B beat message A, and it cannot tell you whether either would clear a bar.
The design tree
Use this as the decision, keyed to message count, the difference you need to detect, and whether the messages trade off on components.
| Messages | Recommended design | Why | Constraint to plan around |
|:--------------------:|:-------------------------------------------------------------------------:|:------------------------------------------------------:|:------------------------------------------------------------------:|
| 2 to 4 | Monadic, cells sized from the power table | Closest to real exposure, and affordable at this count | Total sample is message count times the per-cell figure |
| 5 to 8 | Monadic if budget allows, otherwise sequential monadic with full rotation | Monadic cost scales linearly and gets expensive fast | Analyze by position, never pooled, if you fall back to sequential |
| Above 8 | MaxDiff, also called best-worst scaling | Recovers a stable ranking from many items efficiently | Answers which claim ranks highest, not what a claim does in market |
| Components trade off | Discrete choice or conjoint | The unit of analysis is the component, not the message | This is no longer a message test, and the analysis is different |
That threshold of roughly eight messages is reasoning, not a sourced rule, because no professional body publishes a message-count threshold and this guide will not invent one.
Assumption flagged, and it changes the design rather than a sentence. No documentation was found establishing that respondents are randomly assigned to monadic arms with verifiable balance on any survey platform reviewed for this guide, including the one this guide's publisher sells. This guide therefore assumes arm assignment is not documented, and writes the design so it holds either way. Verify arm balance yourself after collection instead of relying on platform randomization. Cross-tabulate each arm against the attributes you already hold, confirm the distributions match, and report the check alongside the result. That step is cheap, it is the only way to know, and it is safe under either answer.
The honest note on MaxDiff, which matters because vendors sell it
MaxDiff, also called best-worst scaling, asks respondents to pick the best and worst item from small rotating subsets, and it produces a stable ranking across many items while fixing the scale-use bias that flattens rating grids.
The underlying method is attributable to Finn and Louviere (1992) in the Journal of Public Policy and Marketing.
MaxDiff does not fix the gap between what people say matters and what they choose.
Mueller, Lockshin and Louviere (2010), in Marketing Letters, found that both direct methods gave low packaging importance scores contrary to anecdotal industry evidence, while the discrete choice results revealed much higher impacts from packaging-related attributes, and that their results suggest considerable caution in using direct importance measures.
MaxDiff is itself a direct-elicitation method, so it fixes scale-use bias without fixing the say-do gap, and recommending it as the cure for stated-preference problems oversells it.
Sawtooth Software's MaxDiff technical paper is the vendor-canonical methodological reference. Its public page supports its superiority claim by noting that papers won best-presentation awards, and an award is not evidence.
The platform constraint, stated here and not buried in the product chapter
Every design above can be fielded on a modern survey platform, and the test that decides the winner generally cannot be run inside one.
Sprig documents randomization at three levels, display logic, matrix, rank order, MaxDiff and conjoint as question types, and publishes no significance testing anywhere.
For a method whose entire output is whether message A beat message B, that is a structural gap, and this guide handles it by running the test in the analysis chapter.
Where the platform helps, and where it stops. The question-type coverage spans the whole fork, so moving from four messages to twelve does not typically mean changing tools. Randomization, display logic, matrix, rank order, MaxDiff and conjoint are all documented from one study definition. The limitation is two-sided and both halves matter. The escalation path above eight messages runs straight into plan gating, because MaxDiff is Enterprise and Surveys only, conjoint is Enterprise and link surveys only, and Rank Order and Display Logic are Enterprise. And the arithmetic that names the winner happens outside the product, which is why the analysis chapter below carries it.
Writing the instrument
A message testing instrument has five parts in a fixed order: a screener, an exposure screen, the primary measure, the diagnostic battery, and one open-text probe.
Write all five once, then freeze them, because the instrument is what makes this study comparable to the next one.
The exposure screen
The exposure screen shows one message and nothing else. No logo, no screenshot, no supporting copy, and no answer options on the same screen.
Hold the messages to a comparable length, so the comparison is about content rather than about which message gave the respondent more to read.
Add a minimum display time if the platform supports it. A respondent who advanced in under two seconds did not read the message.
The primary measure
The primary measure is one scale, chosen before fielding, and it decides the winner. Every other question in the instrument is diagnostic.
Write it as a five-point or seven-point scale with fully labeled points, and report it as a top-two-box proportion.
Purchase intent is the conventional choice and it carries a hypothetical-bias problem. Likelihood to click or to sign up is often the better primary measure for digital messages, because it names a smaller and more plausible action.
State the action in the question rather than asking about a feeling. "How likely are you to sign up after reading this" is a measurable claim, and "How appealing is this message" is not.
The diagnostic battery
The diagnostic battery explains a result rather than deciding it, and four scales cover the ground that matters.
- Relevance to the respondent
- Believability of the claim
- Uniqueness against alternatives
- Clarity of the wording
Keep the battery to four items and keep the wording identical across every study. A battery that changes between studies cannot build a norm base, and a norm base is the only benchmark this method will ever have.
The open-text probe
One open-text question, placed after the scales so it does not anchor them, asking what the message made the respondent think. That is the whole probe, and one open question is generally the right number to ask.
Two probes typically produce two thin sets of verbatims, not one usable set. Resist the second one, and resist asking respondents to explain their rating, because that question produces rationalization, not reaction.
The in-study control, which most teams skip
Every message test ships with a control message fielded in the same study, in its own cell, under the identical instrument.
The control is your current live message, or the last study's winner if you do not have a live one.
Without it you are comparing this month's proportion against last month's, across two different samples drawn at two different times, and attributing the difference to the message.
A moving control tells you the sample changed and a stable control tells you the comparison is sound. That single cell is the highest-value line in the spec and the one most often cut for budget.
The instrument spec, ready to lift
Copy this structure directly. The bracketed parts are the only things that change between studies.
| Part | Question | Scale | Changes between studies |
|:------------:|:---------------------------------------------------------------------------:|:----------------------:|:-----------------------:|
| Screener | Category and role qualification, 2 to 3 items | Single select | Yes, per audience |
| Exposure | "Please read the following carefully." Message shown alone | None | Yes, the message |
| Primary | "How likely are you to [ACTION] after reading this?" | 5-point, fully labeled | Never |
| Diagnostic 1 | "This message is relevant to me." | 5-point agreement | Never |
| Diagnostic 2 | "I believe what this message claims." | 5-point agreement | Never |
| Diagnostic 3 | "This message says something other companies do not say." | 5-point agreement | Never |
| Diagnostic 4 | "This message is easy to understand." | 5-point agreement | Never |
| Probe | "What did this message make you think? Please answer in a sentence or two." | Open text | Never |
| Control cell | The identical instrument against the current live message | As above | Never |
Writing survey questions well is a separate discipline from designing the study, and the rules that keep a rating scale unbiased apply here without modification.
Where a design agent helps, and where it becomes a trap. Building a multi-cell instrument by hand is slow, and an agent that generates a fully programmed study from an uploaded document, with response options, logic and randomization already in place, removes real work from a design that has six or eight arms. The limitation is specific to this method and it is serious. A norm base requires a frozen instrument, and regenerating the questionnaire between studies silently breaks comparability with every prior test you ran. Generate the first instrument, then version it and reuse it rather than regenerating it.
Scale and scoring choices
Use a five-point or seven-point fully labeled scale for every measure, report the primary measure as top-two-box, and never change the scale once a norm base has started.
Why top-two-box instead of the mean
Top-two-box collapses the scale into a proportion, which is what the two-proportion test operates on, and it is the conventional reporting unit in this method.
A mean carries more information and it is generally harder to defend, because it treats the distance between "somewhat likely" and "likely" as equal to the distance between "unlikely" and "somewhat unlikely," which nobody has established.
Report both if you like. Decide on the proportion, because that is the statistic the arithmetic is built for.
How many scale points
Five points with full labels is generally the safe default and seven is defensible. Below five the scale cannot discriminate, and above seven the labels stop meaning distinct things while respondents typically cluster on the anchors anyway.
Label every point rather than labeling only the ends.
Numeric-only scales with labeled endpoints frequently invite respondents to treat the numbers as a grading scale they remember from school, which is a different instrument than the one you wrote.
Whether to include a midpoint
Include a midpoint and label it, because forcing a choice on respondents who genuinely have no reaction manufactures a preference that is not there.
The argument against a midpoint is that it becomes a parking space for disengaged respondents.
That is a real effect, and the fix is the speeder and straightliner screening in the quality control section, not the removal of a legitimate answer option.
A message test is frequently the study where the midpoint carries the most information.
A message that produces mostly midpoints has not offended anyone and it has not moved anyone either, which is a finding you would lose entirely on a forced-choice scale.
Why the scale is frozen
Changing a scale changes the numbers by an amount nobody can estimate afterward. A five-point scale converted to seven points mid-programme makes every prior study incomparable.
This applies to labels as much as to points.
Changing "very likely" to "extremely likely" is a change to the instrument, so freeze the labels and record them in the study documentation where the next person will find them.
Run every scale in the same direction across the whole instrument, with the positive pole in the same position on every screen. A scale silently flipped between studies produces a norm base that looks stable and is inverted.
Handling non-responses and skips
Exclude blanks and skips from the denominator of the top-two-box proportion rather than treating them as a negative response. A skip is missing data, and counting it as a low rating systematically biases every cell downward.
Record the exclusion rule in the study documentation so the next analyst applies the same denominator. Report the completion rate per cell alongside the result. A cell with a materially lower completion rate is telling you something about the message, and it is typically that the message confused people.
Sample size and precision
Size each cell in a message testing survey from the smallest difference you need to detect, not from a conventional figure. At a 30 percent baseline top-two-box rate, 100 respondents per cell detects a difference of about 19 percentage points and nothing smaller, while real message-test differences typically run 5 to 10 points. The derivation below shows the arithmetic behind that figure.
The conventional cell size is underpowered for this method by roughly a factor of four.
Where the hundred-per-cell rule comes from
No professional body publishes a per-cell sample size rule for message testing. ESOMAR, the American Association for Public Opinion Research and the ICC and ESOMAR code address disclosure, sample quality and response-rate definitions, not cell sizes.
The figures in circulation are platform capacity dressed as method.
One vendor methodology post states that its platform offers surveys with 1,000 respondents, giving ample room for up to 10 stimuli at 100 per cell with statistically stable results, and no source is given for that claim.
One more figure in this section is worth labelling honestly. No published source establishes the typical size of a true difference between two messages, so the 5 to 10 point range used above is the band teams commonly plan against and not a measured quantity. The argument does not depend on it, because a 19-point detection floor is too coarse for any difference a team would spend money resolving.
A 2018 Quirk's article uses 250 per cell inside a worked example rather than as a recommendation, and one major platform's own guidance gives no figure at all.
The rule is convention, not evidence, and convention is a poor reason to accept a 19-point detection floor.
The derivation, shown rather than asserted
The tables below are derived for this guide from the standard two-proportion sample size formula, at a significance level of 0.05 two-sided and 80 percent power, for two independent cells of equal size.
They are labelled derived because they are.
The formula, where p1 and p2 are the two top-two-box proportions and p-bar is their average:
n per cell equals the square of [ 1.96 times the square root of (2 times p-bar times (1 minus p-bar)) plus 0.842 times the square root of (p1 times (1 minus p1) plus p2 times (1 minus p2)) ], all divided by the square of (p2 minus p1).
Every cell below was computed from that formula and then independently checked by simulating 40,000 replicate studies at the resulting sample size, which returned achieved power of 80.2 to 80.3 percent against the 80 percent target.
Table one: respondents per cell required to detect a given difference
| Baseline top-two-box | +3 pts | +5 pts | +7.5 pts | +10 pts | +15 pts | +20 pts |
|:--------------------:|:------:|:------:|:--------:|:-------:|:-------:|:-------:|
| 20 percent | 2,943 | 1,094 | 505 | 294 | 138 | 82 |
| 30 percent | 3,763 | 1,377 | 623 | 356 | 163 | 93 |
| 40 percent | 4,234 | 1,534 | 686 | 388 | 173 | 97 |
| 50 percent | 4,356 | 1,565 | 693 | 388 | 170 | 93 |
Read one row and the scale of the problem is obvious.
Detecting a 5-point difference from a 30 percent baseline needs 1,377 respondents in each cell, so a four-message test needs more than 5,500 completes before the control cell is counted.
Table two: what conventional cell sizes actually buy
This inverts the same formula. At a 30 percent baseline, the figures below are the smallest difference a cell of that size can detect at 80 percent power.
| Respondents per cell | Minimum detectable difference |
|:--------------------:|:-----------------------------:|
| 100 | 19.3 points |
| 150 | 15.6 points |
| 200 | 13.5 points |
| 300 | 10.9 points |
| 400 | 9.4 points |
| 500 | 8.4 points |
| 750 | 6.8 points |
| 1,000 | 5.9 points |
The practical reading is that a 100-respondent cell can typically only find differences so large you would have seen them without a survey.
That arithmetic sits behind every message-test winner that failed to replicate, and it is not arguable.
The multiple-comparisons caveat, which makes the tables optimistic
Both tables describe the two-cell case. Testing k messages means k times (k minus 1), divided by 2, pairwise comparisons.
Six messages is fifteen comparisons, and at an uncorrected 0.05 threshold a set of fifteen comparisons will typically produce a winner whether or not one exists.
Correcting for that raises the required sample further. The table below applies a Bonferroni correction to the same derivation, for a 10-point difference from a 30 percent baseline.
| Messages tested | Pairwise comparisons | Per-cell n after correction | Total completes |
|:---------------:|:--------------------:|:---------------------------:|:---------------:|
| 2 | 1 | 356 | 712 |
| 3 | 3 | 475 | 1,425 |
| 4 | 6 | 550 | 2,200 |
| 5 | 10 | 605 | 3,025 |
| 6 | 15 | 648 | 3,888 |
| 8 | 28 | 714 | 5,712 |
Bonferroni is the conservative correction and other corrections are generally less punitive.
The point of the table is the shape, not the exact figure: message count raises the cost of the study twice, once through cells and once through the correction.
You will encounter different sample-size figures elsewhere, including in survey templates. Published guidance for this family of study ranges from roughly 10 to 30 participants for directional reads up to at least 100 participants per concept, and this guide's own derivation contradicts both. That disagreement is real and this guide is not going to paper over it. The derivation above is shown in full precisely so you can check it rather than taking it on authority. If you need a directional read and can accept a 19-point detection floor, a small cell is a defensible choice as long as you say so in the report. What is not defensible is a small cell and a claim that one message beat another.
What to do when you cannot afford the sample
Reduce the number of messages, not the cell size.
Four messages at 550 per cell is a study that can find a 10-point difference, while eight messages at 275 per cell is a study that cannot, and the second one costs the same.
Or change the design. A comparative or MaxDiff design recovers an ordering from far fewer respondents, and the predictive-validity evidence in the design chapter is not the argument against it that practitioners assume.
Or accept a directional read and label it as one. A ranking with a stated minimum detectable difference beside it is honest, and a ranking presented as a result is not.
Segments multiply everything
Every segment you intend to report is a separate sample size problem. Reporting a message test by three segments means each segment within each cell needs enough respondents to support its own comparison.
The common failure is fielding to a total that looks adequate, then discovering at analysis that the segment cells hold 40 respondents each. Decide the reporting segments before fielding and size to the smallest one.
Audience, targeting and screening
Field a message test to the population the message is meant to persuade, which is frequently not your current customers.
Screen on category behavior and role rather than demographics alone, and stamp every segment variable onto the response at collection.
Who should see the message
The audience for a message test is the audience the message targets in market. Testing acquisition messaging on existing users measures how your current customers react to copy written for people who have never heard of you.
Existing customers are the right audience when the message is about expansion, renewal or a new feature for people already inside the product.
They are the wrong audience for positioning aimed at a category you are trying to enter.
Screening rules
Keep the screener to two or three items and put the disqualifying question first, because a respondent who fails on question three has already spent time you paid for.
Screen on behavior rather than on self-reported attitude. "Which of these have you purchased in the last six months" is checkable against the rest of the response, and "How interested are you in this category" is not.
Avoid naming your brand in the screener. A respondent who learns the study is about your company typically evaluates the message differently, and that cannot be removed from the data afterward.
Quotas, and the constraint worth planning around
Quotas generally keep the cells balanced on the attributes that predict response, and they are the mechanism that stops one arm filling with a different population than another.
Quotas on the platform this guide's publisher sells are response-based, which means they evaluate answers as respondents submit them and can only be built from a short list of closed question types.
They cannot be set on pre-existing attributes held before the study begins.
Plan the screener and the quota together, because the quota can only act on answers the screener collects. The practical consequence is that balancing arms on something you already know about a respondent is not a quota operation on this platform. It is a targeting operation, and targeting is a weaker guarantee.
Stamping the arm onto the response is the step that makes the analysis possible. Attribute piping and passing attributes through the study URL let you write the assigned arm and every segment variable onto each response as it is collected, which is exactly what the analysis needs and what the export does not otherwise carry. The limitation is that attributes target rather than quota. Targeting on an attribute does not hold its distribution constant across arms, so stamping the variable makes the balance check possible without making the balance itself likely. Run the check.
Fielding and delivery
Field every cell of a message test simultaneously through the same channel, and close all cells at the same time. Staggering cells introduces a time confound that no analysis can remove.
Choosing the channel
Choose the channel by which population you need, and decide it before writing the screener, because the channel determines which attributes you already hold and which you have to ask for. Research panels reach the category population rather than your current users, which is usually the right frame for acquisition messaging.
In-product delivery on web apps and websites reaches people already using the product, which is the right frame for expansion messaging. Email reaches a known list, which is the right frame when the message targets a segment you can identify by name.
Panel fielding is the default for most message tests and it carries a planning risk that is easy to discover late.
Confirm the minimum and maximum sample a provider will field, the turnaround, and the country coverage before you commit to a timeline, because none of those are typically published.
Fielding mechanics that protect the comparison
Field all arms in one window, ideally a single week, so news and seasonal effects land on every cell equally.
Do not top up one cell that is filling slowly. A topped-up cell is a different sample from a different period.
Field through one channel for the whole study. Mixing in-product and panel respondents across arms introduces a channel confound that looks exactly like a message effect in the output.
Monitor completion rate by arm during fielding, not after. A cell dropping out at a materially higher rate is typically a signal about the message, and catching it live lets you decide whether to keep fielding it.
In-product delivery buys context and costs representativeness. Showing the message inside the product puts it in front of people in the actual usage context, which is closer to real exposure than a panel respondent working through a survey queue. The limitation runs in three directions. You are testing on people who already arrived, which is the population least in need of persuading. The panel alternative that reaches the category population is Enterprise-only and link-surveys-only on this platform, with no published minimum or maximum sample, turnaround or country coverage. And the forced-exposure problem from the critique chapter applies to an in-product prompt exactly as it applies to a panel.
Timing and cadence
Run a message test when a messaging decision is pending and not on a schedule. This method is decision-triggered and not continuous, and the tracking instruments belong to a different study type.
How long to field
Field on until every cell hits its target, and then stop. Do not stop at a date if cells are uneven, and do not keep fielding a cell that already hit target.
Most panel-fielded message tests typically complete in days rather than weeks, which makes fielding time a poor reason to cut the sample.
In-product tests take as long as the traffic takes, which is why the per-cell target is set by the power table and not by the calendar.
Avoid fielding into a known seasonal peak or into a category news event. Neither invalidates a between-arm comparison on its own, because every arm is exposed to the same conditions, but both change the absolute level enough to break comparability with your norm base.
Repeating the study
Re-testing the same message is worth doing when the audience or the category has changed, and it is worth doing on a schedule only if you are tracking, which is a different method with a different design.
Recontact settings matter here.
The recontact waiting period on this platform is an account-wide default applying to everyone shown a study, whether or not they responded, so a message test fielded in-product spends part of the account's recontact budget on people who never answered.
Quality control and data hygiene
Clean a message test on four checks before any analysis runs, and run them in this order: speeders, straightliners, arm balance and control stability.
The four checks
Remove respondents who completed faster than roughly a third of the median completion time, which is a working convention rather than a published standard. A respondent who did not read the exposure screen did not take the test you fielded.
Remove straightliners who gave the identical answer to every diagnostic item, which is a response pattern, not an opinion. Bot detection catches automated traffic and it does not catch a bored human, which is why the speeder check runs first in the sequence.
Verify arm balance by cross-tabulating each arm against the attributes you stamped at collection. Unequal distributions generally mean the comparison is confounded and the result needs weighting or a caveat.
Verify the control cell against the previous study. A control that moved materially means the sample changed, and the message differences you are about to report are sitting on top of that shift.
Document every exclusion before you analyze anything, recording how many respondents were removed by each check and from which arm.
Uneven exclusions across arms are themselves a finding, and they are frequently the first sign that one message confused people enough to drive them out of the study.
The pre-launch checklist
Run this in order before fielding. The last four are the ones nobody checks.
- Freeze the primary measure wording and scale
- Confirm the diagnostic battery is unchanged from the last study
- Set per-cell targets from the power table
- Randomize question and answer order where appropriate
- Stamp the arm variable onto every response
- Confirm arm assignment produces balanced cells on key attributes
- Confirm the control message is byte-identical to the last study
- Confirm no scale label was edited since the last study
- Confirm the per-cell target came from the table and not from habit
Work through that list with the study open in front of you rather than from memory, because the four failures it catches are all invisible in the finished data.
Analysis: the core calculation
The core calculation in a message testing survey is a two-proportion test comparing the top-two-box rate in one arm against another, using a pooled standard error, with a correction applied for the number of pairwise comparisons in the study. Everything else in the analysis is context around that one test, and a report that omits the correction is describing the result of an untracked number of tests.
The test, step by step
Compute the top-two-box proportion for each arm. Count the respondents in that arm who chose one of the top two scale points, divide by the number who answered the question, and exclude blanks and skips from the denominator.
Compute the pooled proportion for the two arms compared: add the two top-two-box counts, add the two denominators, and divide.
Compute the pooled standard error as the square root of [ pooled proportion times (1 minus pooled proportion) times (1 divided by n1 plus 1 divided by n2) ].
Divide the difference between the two arm proportions by that standard error to get a z statistic, then convert it to a two-sided p-value.
The mistake this replaces
The wrong comparison, and the one most reports commonly make, is putting a margin of error beside each arm and declaring a winner when those intervals do not overlap.
A single-arm margin of error describes the precision of one estimate. The question you are asking is about the difference between two estimates, which has its own standard error and its own interval.
Report the difference and its interval rather than two separate intervals. The difference is the quantity the decision rests on.
Applying the multiple-comparisons correction
Count the comparisons before interpreting any of them. A study with k messages including the control contains k times (k minus 1), divided by 2, pairwise comparisons, and every one gets tested whether or not it gets reported.
Apply a stated correction and name it in the report.
Bonferroni divides the threshold by the number of comparisons and is conservative, while Holm and Benjamini-Hochberg are less punitive and are fine choices as long as the report says which one was used.
What matters more than the choice is that an uncorrected report of fifteen comparisons will always have a winner, and a reader who does not know how many tests were run cannot discount it.
Reading the diagnostic battery
Analyze the diagnostic battery only after the primary measure, and treat it as explanation and not as a second verdict.
A message that wins on relevance and believability but not on the primary measure is telling you the battery and the primary measure are measuring different things, which is a finding about your instrument, not about the message.
A message that wins on the primary measure while scoring low on clarity is typically a message that is intriguing rather than understood, and that pattern frequently does not survive contact with a real channel.
Analyzing the open text
Code the open-text probe into themes, count the themes by arm, and read the raw verbatims for the losing arms, not only the winners.
The losing arms carry the objections. A message that lost by eight points and produced a consistent objection you had not anticipated is often more valuable than the winner.
Theme counts are diagnostic and they are not a scored metric. Comparing theme counts across studies is only valid if the theme list was applied instead of regenerated, which the artificial-intelligence chapter below returns to.
Automatic theming with response-level traceability is the part of this that genuinely saves time. Open-text theming that regenerates in real time and lets you trace a theme back to the responses that produced it is what makes the probe worth fielding, because an untraceable theme is an assertion. The limitation is the one that decides how you may use it. No minimum response count is documented, no accuracy or reproducibility figure is published, and themes regenerate, so a regenerated theme list is not the same instrument as last period's. Treat theming as diagnostic, and never as a scored metric you trend.
Benchmarks and what a good result looks like
There is no publicly retrievable cross-industry benchmark for message-testing scores that discloses its sample size and method. The norm databases that exist, including Nielsen BASES, Zappi, Ipsos and Kantar, are commercial assets held behind client relationships, and every freely published norm table found for this guide cites no study, no sample and no date. A good result is therefore defined against your own prior studies and your own in-study control.
What was checked, and what came back
ESOMAR publishes codes and guidelines, not norms. The Advertising Research Foundation hosts research reports and publishes no open norms database.
The American Association for Public Opinion Research publishes Standard Definitions and endorses ISO 20252, and no score norms of any kind.
That is a verified absence rather than a failure to look hard enough, and stating it plainly is more useful to a reader than a fabricated range would be.
A worked example of what to refuse
One widely shared reference page publishes top-two-box purchase-intent bands by category and says of its own figures only that they come from aggregated industry data and vary significantly by source.
No study, no sample size, no date and no method sit behind them.
The numbers are not reproduced here, and to that page's credit it argues that its own figures should not be used as thresholds.
Treat every norm table you encounter the same way. If it does not disclose the sample size, the method and the date, it is not a benchmark, and reproducing it transfers someone else's unsourced number onto your decision.
One more absence is worth naming. No trade body publishes a norm for the diagnostic battery either, so relevance and believability scores have no external reference point at all and are only interpretable against your own prior studies.
What to use instead
Two comparisons are worth building, and one of them is available immediately.
Build an internal norm base. Every message you test goes into a record of your own historical distribution on your own frozen instrument, in your own category, with your own audience.
And field a control message inside every study.
An in-study control is the only comparison with a matched sample frame, a matched questionnaire and a matched fielding window, which makes it the only benchmark in this method that is actually comparable.
The in-study control is the load-bearing half of that recommendation and most teams skip it.
A historical distribution without a control tells you this month's number is higher than last month's, and it cannot tell you whether the message or the sample moved.
No minimum score, threshold or floor is published in this guide. No source supports one, and a guide that invented one would be doing the thing this section just told you to refuse.
Interpreting and acting on the result
Interpret a message test as a screen. The output is which messages survive to a live test and which stop here, and it is not a prediction of what the survivor will do in market.
The three readings of a result
A difference larger than your minimum detectable difference, surviving the correction, is a result you can act on. Move the winner forward to a live test.
A difference smaller than your minimum detectable difference is not a result. It is a number, and reporting it as a finding is how a message test becomes folklore inside a company.
A tie between two messages tested at adequate power is itself genuinely useful information.
It means the difference between them is smaller than the difference you said you cared about, so pick on another basis and stop spending on the question.
What to do with the winner
Take the winner to a randomized test in the channel where it will live. That is the decision step, and the survey was the screen that made it affordable.
A message test that narrows eight candidates to two has done its job completely. Running the live test on two messages rather than eight is the value the survey produced, and it is a real saving.
What to do with the losers
Read the open text from the losing arms before discarding them. A message can lose on execution while carrying an idea worth rewriting, and the probe is where that shows up.
Record every tested message and its score in the norm base, including the losers. A norm base built only from winners has a truncated distribution and it will typically mislead you within two cycles.
What to tell stakeholders who want one number
Give them the winning message, the size of its lead in percentage points, and the smallest lead the study was capable of detecting. Three numbers, in that order.
The third number is the one that stops the conversation going wrong later.
A stakeholder who knows the study could only resolve differences above 10 points will not ask why a 4-point lead failed to show up in revenue.
Refuse the request for a score out of 100 or a pass mark. No published source supports a threshold in this method, and inventing one converts a defensible ranking into an indefensible grade.
If the decision genuinely needs a single go or no-go, the honest answer is that this study cannot supply it and the in-market test can.
Reporting the result honestly
State the per-cell n, the minimum detectable difference, the number of comparisons run and the correction applied, in the report rather than in an appendix.
A message-test report without those four numbers cannot be evaluated by the person reading it, and a stakeholder who cannot evaluate it will typically either over-trust it or dismiss it. Both of those outcomes cost more than the single sentence it takes to prevent them.
Running the analysis with Claude or ChatGPT
Running a message test analysis in a language model generally works well for the arithmetic and the coding, and it fails in predictable ways that you can design around.
The two prompts below cover it: the between-arm test that names the winner, and the open-text coding that explains it.
Part one: what the analysis produces
The analysis produces four artifacts: a table of top-two-box proportions by arm with per-cell n, a table of pairwise differences with p-values and a stated correction, a coded theme table broken out by arm, and a stated minimum detectable difference.
Coding open-ended responses into a taxonomy is one task where current models perform at the human ceiling, and that is measured, not assumed.
Mellon and colleagues (2024), in Research and Politics, tested a novel 50-category coding task with no prior training data and found Claude reached 93.9 percent accuracy against a human coder's 94.7 percent, while human-to-human agreement was 86.6 percent.
The honest comparison is human-to-human agreement, not perfection. On that comparison the model is at the ceiling, and the pitfalls further down are about everything except this one task.
Part two: prompt one, the between-arm test
Paste this with the export attached. The bracketed values are the only things you change.
What this prompt does: runs the between-arm significance test for a monadic
message test and applies a multiple-comparisons correction.
What it returns: a proportions table, a pairwise difference table with
corrected p-values, and a stated minimum detectable difference.
Data: the attached export. Each row is one respondent.
Arm column: [ARM_COLUMN, default: the piped attribute named message_arm]
Primary measure column: [PRIMARY_COLUMN, default: Q1_Response]
Top-two-box values: [TOP_TWO_LABELS, default: "Likely" and "Very likely"]
Step 1. Derive the arm for every respondent. If no arm column exists, derive
it from which message question the respondent answered. Report how many rows
you could not assign and stop if that number is above [UNASSIGNED_LIMIT,
default: 2 percent of rows].
Step 2. Exclude blanks, skips, and any non-response value from the denominator
of every proportion. Do not treat a blank as a negative response. Report the
excluded count per arm.
Step 3. Any arm with fewer than [MIN_N, default: 30] valid responses is
reported as insufficient and not compared. An arm with zero valid responses
is under this threshold and is reported as zero, not omitted.
Step 4. Use code to calculate this, not estimation. Compute the top-two-box
proportion per arm, then every pairwise two-proportion z test using the
pooled standard error.
Step 5. Count the pairwise comparisons as k times (k minus 1) divided by 2 and
state the number. Apply [CORRECTION, default: Bonferroni] and report both the
raw and the corrected p-value for every pair.
Step 6. Independently recompute the result by a different method: run a
2 by 2 chi-square test of independence on the raw counts for each pair and
compare it against the z test. These are different calculations, not a
re-read of the first one.
Step 7. If the two methods disagree on any pair at the stated threshold,
report the disagreement and which pairs it affects. Do not silently reconcile
them and do not pick the one that gives a cleaner story.
Step 8. Compute the minimum detectable difference at the observed per-cell n
and the observed baseline, at 80 percent power, and print it above the tables.
Output artifact: a file named message-test-results.md containing the three
tables and a one-paragraph summary naming which differences clear the
corrected threshold.
Part three: prompt two, coding the open-text probe
What this prompt does: codes the open-text probe into themes and attaches the
primary measure to each theme.
What it returns: a theme table by arm with counts, percentages, and the
top-two-box rate of the respondents in each theme.
Data: the attached export.
Open-text column: [PROBE_COLUMN, default: Q6_Response]
Arm column: [ARM_COLUMN, default: message_arm]
Primary measure column: [PRIMARY_COLUMN, default: Q1_Response]
Step 1. Exclude blank responses, single-character responses, and responses
that only repeat the message text. Report the excluded count.
Step 2. Apply [THEME_LIST, default: none, discover inductively on the first
run only]. On every run after the first, apply the frozen list rather than
discovering a new one.
Step 3. Any theme with fewer than [MIN_THEME_N, default: 5] respondents is
reported in an Other or unclear bucket instead of as its own theme. A theme
with zero respondents in an arm is reported as zero, not omitted.
Step 4. Use code to calculate this, not estimation. Count respondents per
theme per arm and compute the top-two-box rate within each theme.
Step 5. Independently recompute the counts by a different method: take a
random sample of [CHECK_N, default: 40] coded responses, re-read the raw text
of each, and report the proportion where your own code was wrong. This is a
manual re-read against the source, not a recount of your own table.
Step 6. If the error rate on that check is above [ERROR_LIMIT, default:
10 percent], say so at the top of the output and recommend rebuilding the
theme list rather than reporting the table as final.
Output artifact: a file named message-test-themes.md containing the theme
table and three verbatim quotes per theme.
Part four: setup and getting your data in
The documented export carries survey and respondent identifiers, a timestamp, a user identifier, one column per question and per response, a themes column and numbered attribute columns. It does not carry an arm column.
Deriving the arm field before any analysis runs is therefore a real step and not a formality.
Recover it from the attribute you piped at collection, and confirm the derived counts match the cells you fielded.
Put the data above the instructions rather than below them, and split a large export rather than pasting all of it.
A connected model reading study data directly removes the export step and adds a governance question. Sprig MCP connects Claude, ChatGPT, Gemini, Copilot and Cursor to study data with published governance: access scoped to the authenticated user's role, response data not used to train models, agents unable to launch or modify a live study, and an org-wide admin kill switch. Two limitations apply directly to this analysis. Calls are capped at 1,000 responses, so a well-powered multi-arm test typically needs several. And a connection hands the model data to reason over instead of a statistical test, which is exactly why the prompt above specifies the test rather than asking which message won.
Part five: reading the output
Read the minimum detectable difference first, because a reported difference smaller than it means the table is describing noise.
Read the comparison count second, because one significant result out of fifteen uncorrected comparisons is the expected outcome when every message is the same.
Read the control cell third, because a control that moved materially against the last study means the sample changed and the message differences sit on top of that shift.
Part six: what to verify before reporting
Recompute one arm's top-two-box proportion by hand from the raw counts. It catches the most common failure, which is a denominator that silently included blanks.
Confirm the arm counts match what you fielded, that the excluded-response counts are plausible, and that the correction named in the output is the one you asked for.
Run the analysis a second time in a new session. Neither chat client exposes a random seed, so a reportable number should be produced twice.
You remain responsible for the result the model produces. The prompts above are written so the arithmetic is checkable, and the check is yours to run before the number reaches anyone else.
Part seven: the pitfalls
Discovering themes is harder than applying them.
Hill and colleagues, in PLOS Digital Health in April 2026, found deductive agreement of 93.5 percent against blinded human coders at 92.7 percent, and yet a kappa of only 0.34 for both, with comprehensive error at 12.4 percent.
Never quote the 93.5 percent without the kappa, because agreement looks high only because code prevalence is low at 7.8 percent. Theme inductively once, rebuild the list yourself, then have the model apply it every period after.
Rare themes get over-predicted, and rare themes are the ones that get escalated. Ashwin, Chhabra and Rao, in Sociological Methods and Research in May 2025, found non-random bias in 10 of 19 codes.
Non-determinism is real and undisclosed, and Thinking Machines Lab reported in September 2025 that 1,000 completions at temperature zero produced 80 unique outputs.
Long contexts lose the middle, which Liu and colleagues documented in TACL in 2024. Put data above instructions and split large sets.
Models agree with you. Sharma and colleagues at ICLR 2024 documented sycophancy, and OpenAI withdrew a model update in April 2025 for being overly agreeable. Never ask a model to confirm the message you already believe won.
Verbatims contain personal information nobody asked for, and the training-data question turns on which tier you use, not which vendor.
One pitfall belongs to this method alone: asking a model which message won. It will name one, every time, because the file contains k times (k minus 1) divided by 2 comparisons and nothing is counting them.
That is the multiple-comparisons problem wearing a language-model costume, and it is the reason the prompt above specifies the test rather than asking the question.
What this chapter does not cover
MaxDiff analysis and conjoint analysis each have their own dedicated guide, and the escalation path leads directly to both. Use those rather than adapting the prompts above, because the estimation is genuinely different.
Theme creation mechanics, the Other or unclear bucket and the recode loop belong to the advanced cross-tab guide, while cross-tabbing a measure across customer segments belongs to the simple cross-tab guide.
The critique you should know
The strongest published critique of message testing is that the evidence connecting a survey score to in-market performance is thin, old and partly failed to replicate. No modern published validation study was found demonstrating that survey-based message testing specifically predicts in-market performance, and searches around the Advertising Research Foundation and the Ehrenberg-Bass Institute returned nothing retrievable. A message test is therefore defensible as a screen and not as a forecast.
The case against
Stated intent is a weak proxy and the size of the gap is measured.
Webb and Sheeran (2006) found that a medium-to-large change in intention of d = 0.66 produced only a small-to-medium change in behavior of d = 0.36, across 47 experimental tests.
Survey-based estimates of advertising effect are frequently wrong by a large multiple.
Gordon and colleagues (2019) found that in half of their 15 field experiments the estimated percentage increase in purchase outcomes was off by a factor of three across all methods.
And survey exposure is not real exposure. Schmalzle and colleagues (2023) recorded recall falling from 6.45 of 20 billboards under free viewing to 2.95 under added attentional demand.
A message a respondent is forced to read is not the message a customer will choose whether to read.
The counter-critique, and why it is thinner than the industry pretends
The canonical defense of copy testing is the Advertising Research Foundation's Copy Research Validity Project, reported by Haley and Baldinger in the Journal of Advertising Research in 1991, with conclusions published in the same journal in 1994 and reprinted in 2000.
The project related 46 measures of advertising effectiveness to five pairs of commercial executions aired in split-cable markets, and found that commercial liking and brand-name recall from a product category cue had evaluative power with regard to split-cable outcomes.
That is a real study and it is worth knowing.
Its primary statistics are behind a paywall and the findings above come from a secondary journal record, so this guide cites the project's shape and does not quote a correlation from it.
The counter-counter, which a research-literate reader already knows
Research Systems Corporation replicated the question across more than 20 split-cable cases and did not support the use of either commercial liking or brand-name recall as a valid criterion measure of advertising copy quality.
Be blunt about what that leaves. The canonical industry defense of copy testing rests on five pairs of commercials, and its central finding failed to replicate on a larger set of cases.
That is a thin evidentiary base for the claim that copy testing predicts in-market performance.
The second counter-critique, which is the honest one
The method survives this critique on a narrower claim than the industry makes for it.
Message testing is well supported as a discrimination instrument, meaning it reliably separates messages that differ, provided the cells are powered for the difference you care about.
The arithmetic in the sample size chapter is not a criticism of the method.
It is the method working as designed, and the failures attributed to message testing are typically failures to run it at adequate power, not failures of the instrument.
There is also a straightforward economic argument the critics do not answer.
Narrowing eight candidate messages to two before a live test is generally cheaper than running eight live tests, and nothing above suggests a powered survey is worse than a coin flip at that job.
The stated position
A message test is a well-powered screening instrument rather than a prediction of in-market performance. It is very good at telling you which of five messages to stop working on.
It is not evidence that the survivor will sell anything, and the only thing that establishes that is an experiment.
Screen with the survey, and decide with the test. Every source in this chapter is consistent with that division of labor, and a guide claiming more than it would be claiming something no retrievable study supports.
Common mistakes
The six mistakes below account for most message tests that produce a confident answer and a disappointing outcome.
Five of them are design errors and one is a reporting error, and the design errors are all cheaper to prevent than to detect.
Sizing the cell from habit
Fielding 100 respondents per cell because that is the figure everyone uses produces a study that can only detect a 19-point difference, on a method where real differences typically run 5 to 10 points.
The fix is one line of arithmetic run before fielding, not after. Decide the smallest difference that would change your decision, read the per-cell figure off the table, and field that.
Testing too many messages at once
Adding messages raises cost twice, once through the extra cells and once through the multiple-comparisons correction, and teams typically absorb both by shrinking the cells.
The fix is to cut the message list before fielding instead of cutting the cell size. Four messages tested properly beats eight tested badly, and the two cost about the same.
Omitting the in-study control
Without a control message fielded in the same study, this period's numbers are being compared against last period's across two different samples, and any sample shift is attributed to the message.
The fix here costs exactly one cell. Field the current live message in every study, unchanged, and check it before reading anything else.
Editing the instrument between studies
Improving a scale label, adding a diagnostic item or rewording the primary measure generally breaks comparability with every prior study, and no adjustment recovers it afterward.
The fix is to version the instrument and treat changes as a deliberate restart of the norm base, not as an improvement. Write down the date the instrument was frozen.
Testing messages that are not comparable
Messages that differ in length, format or specificity are not being compared on their content, and the longer or more concrete message typically wins on that difference alone.
The fix is a pre-field read of the message set side by side, checking that every message describes the same offer at the same level of abstraction in roughly the same number of words.
Where a message genuinely needs more words to work, that is a finding about the message and it belongs in the report. It is not a reason to field that message against shorter alternatives and then call the result a preference.
Reporting a ranking without a detection floor
A ranked table with no minimum detectable difference beside it reads as a result to everyone who sees it, including people who would have discounted it if they knew the study could not resolve differences that small.
The fix is a single line above the table stating the per-cell n, the minimum detectable difference and the number of comparisons run.
Common survey mistakes commonly come from omitted context, not from wrong arithmetic, and this is the clearest case of it in the method.
Synthetic respondents
Synthetic respondents are not usable to decide a message test today, and the reason is specific, not general: the evidence shows that model-generated means look reasonable while differences between cells do not.
A message test is nothing but a difference between cells.
The direct evidence
Bisbee and colleagues (2024), in Political Analysis, generated 3,614,400 synthetic responses from personas based on 7,530 human respondents.
Their finding, verbatim, is that the average scores generated by ChatGPT correspond closely to the averages in the baseline survey, and that sampling by ChatGPT is nevertheless not reliable for statistical inference, because there is less variation in responses than in the real surveys and regression coefficients often differ significantly from equivalent estimates.
Forty-eight percent of coefficients differed significantly from their benchmark counterparts, and among those the sign flipped 32 percent of the time.
The variance collapse is the part that matters here. It produces false precision, which is the worst failure mode available in a method whose output is a between-cell difference.
The strongest pro-synthetic result, handled honestly
Maier and colleagues (2025) report semantic similarity rating, which elicits textual responses and maps them to scale distributions using embedding similarity rather than asking a model for a number, and which reached 90 percent of human test-retest reliability across 57 personal care product surveys covering 9,300 human responses.
Four caveats belong in the same sentence as that number. It is a preprint, not a peer-reviewed paper. It is co-authored by a vendor selling synthetic-consumer services with a category partner supplying the data.
It covers one product category, and the authors flag that cross-category applicability is unvalidated and that performance likely benefited from abundant personal care product discussions in training data. And 90 percent is 90 percent of the human test-retest ceiling, not 90 percent accuracy.
The authors describe it as a proof of concept.
The position
Do not use synthetic respondents to decide a message test.
Watch the direction, because eliciting text and mapping it to a distribution is a real advance over asking a model for a rating, and the validation so far is one category and one vendor.
How Sprig supports a message testing survey
Sprig supports the design, fielding and collection stages of a message test from one study definition, and it does not run the test that names the winner.
Capabilities described here were verified against published product documentation and marketing pages on September 11, 2026.
What the platform covers
The question types cover the whole design fork: rating scale, matrix, rank order, MaxDiff and conjoint, with randomization documented at three levels and display logic for routing.
Distribution runs through in-product web and native mobile, native email with a custom sending domain, shareable links, QR codes and external research panels.
Attribute piping and attributes passed through the study URL let you stamp the assigned arm and every segment variable onto each response at collection, which makes the analysis in this guide possible.
More than 300 targeting attributes are available for defining who sees a study.
Open-text theming with response-level traceability handles the probe, and a connected model reading data through the published Model Context Protocol server handles the arithmetic the platform does not.
What it does not cover, stated plainly
No significance testing is documented anywhere in the product, which for this method is the structural gap and not a missing convenience.
No trend view or cross-study comparison returning numbers is documented either, so an internal norm base is an export plus your own tooling.
No weighting approach, sample-size guidance or statistical power method is published. For a guide whose central asset is a derived power table, that is worth saying plainly rather than softening.
Quotas are response-based and cannot be set on pre-existing attributes. In-product audience sampling is documented as a throttle that spreads delivery over time rather than as a probability sample, and it does not state that selection is random.
Export is comma-separated values only, capped at 950,000 rows, with a link that expires after a week and no warehouse connector. Hosting is United States only on Amazon Web Services with no published residency options.
Plan gating matters more in this method than in most.
MaxDiff is Enterprise and Surveys only, conjoint is Enterprise and link surveys only, and Rank Order and Display Logic are Enterprise, so the escalation path above eight messages is not available on every plan.
Third-party standing
Sprig holds a G2 rating of 4.3 out of 5 across 199 reviews as of August 17, 2026, on the product listing and not the seller listing.
TrustRadius shows 8.5 out of 10 across 10 reviews as of August 13, 2026, a base small enough that it should not be compared against a platform carrying hundreds.
Capterra is a not-found and not a zero, because the product is still filed under a legacy slug with no reviews.
Gartner Peer Insights returned an access error to every tool available and is likewise reported as not found.
Alternatives and adjacent methods
Choose the method by what varies and by how many things vary at once. Message testing varies expression while holding the offer constant, and four adjacent methods answer questions it cannot.
When to run something else
Concept testing varies the product idea rather than its expression, and it is the right method when the offer itself is what you are unsure about.
Concept validation with your own users is the lighter version, and it is generally faster to field.
MaxDiff, also called best-worst scaling, recovers a stable ranking across many items from short rotating tasks, and it is the right escalation above roughly eight messages when you need an order rather than a level.
Conjoint analysis estimates tradeoffs across combinations of attributes, while MaxDiff measures relative preference among individual items.
When your messages differ on benefit, proof point and tone at the same time, the unit of analysis is the component, not the message.
Brand perception and brand awareness studies measure standing rather than message reaction, and they are tracking instruments run on a cadence, not decision-triggered studies.
Marketing site clarity testing measures whether a live page communicates, which is a usability question, not a preference question.
What sits outside the survey category entirely
Moderated and artificial-intelligence-moderated depth interviews answer why a message failed, from few participants, through follow-up questioning. They are qualitative instruments in a different category and they are not survey platforms.
Message testing and brand tracking are frequently confused, and the difference between them is the decision each one is built to support. Randomized in-market tests are the only instruments that establish what a message does to behavior.
They sit downstream of every method on this list, and routing to them is the point of running a screen in the first place.
Frequently asked questions
What sample size do I need for a message testing survey?
Size each cell from the smallest difference you need to detect. At a 30 percent baseline top-two-box rate, detecting a 10-point difference at 80 percent power needs about 356 respondents per cell, a 5-point difference needs about 1,377, and 100 per cell detects only about a 19-point difference.
Is 100 respondents per cell enough for message testing?
No, 100 respondents per cell is not enough for most message tests. At a 30 percent baseline it detects a difference of roughly 19 percentage points and nothing smaller, while real message-test differences typically run 5 to 10 points. No professional body publishes a per-cell rule, and the figure circulates as platform capacity.
Should message testing be monadic or comparative?
Monadic is the conventional default because it more closely mirrors typical exposure, but the published predictive-validity evidence favors comparative collection. Morwitz, Steckel and Gupta (2007) found intentions correlated more strongly with purchase when collected in a comparative mode than monadically. The counterargument that comparative exaggerates differences has no peer-reviewed support.
What is the difference between monadic and sequential monadic testing?
Monadic testing shows each respondent exactly one message, so each message gets an independent cell. Sequential monadic shows each respondent several messages in rotation, reusing respondents across cells to save sample, and it carries a documented second-position penalty. Rotate order and analyze by position instead of pooling.
What should a message testing survey measure?
A message testing survey measures one primary scale that decides the winner, a short diagnostic battery that explains the result, and one open-text probe. Relevance, believability, uniqueness and clarity are generally the four diagnostics worth the space. Report the primary measure as a top-two-box proportion.
How do you analyze message testing results?
Compute the top-two-box proportion for each arm, excluding blanks from the denominator, then run a two-proportion z test using a pooled standard error for each pair. Count the pairwise comparisons as k times k minus 1 divided by 2, apply a named correction, and report the minimum detectable difference beside the table.
Does message testing predict how a message will perform in market?
No, message testing does not predict in-market performance. No modern published validation study demonstrating that link could be found, and the canonical industry defense rests on five pairs of commercials whose central finding failed to replicate. Treat a message test as a screen and route the decision to a randomized in-market test.
Do I need a control message in a message test?
Yes, field your current live message as a control cell inside every study. A control is the only comparison with a matched sample frame, questionnaire and fielding window, which makes it the only benchmark this method reliably has. A control that moves between studies typically tells you the sample changed, not the message.
What is a good top-two-box score for a message test?
There is no publicly retrievable cross-industry benchmark for message-testing scores that discloses its sample size and method. The norm databases that exist are commercial assets held behind client relationships. Build an internal norm base on a frozen instrument and compare every message against an in-study control rather than against a published range.
Can I use synthetic respondents or AI-generated answers for message testing?
No, synthetic respondents are not usable to decide a message test today. Model-generated averages track human averages reasonably well while between-cell differences do not, and a message test is nothing but a difference between cells. Bisbee and colleagues (2024) found 48 percent of coefficients differed significantly from human benchmarks.
How many messages can I test in one survey?
Two to four messages typically suit a monadic design at affordable sample. Five to eight push monadic cost up sharply and push a sequential design into order effects. Above roughly eight, MaxDiff recovers a ranking more efficiently, and that threshold is reasoning rather than a sourced rule because no professional body publishes one.
What if I do not have enough traffic or budget for a properly powered test?
Cut the number of messages before cutting the cell size, because four messages at 550 respondents each can find a 10-point difference and eight at 275 each cannot, for the same total cost. If the sample is still out of reach, switch to a comparative or MaxDiff design, which recovers an ordering from far fewer respondents, or publish a directional read with its detection floor stated.
Who should I survey for a message test?
Survey the population the message is meant to persuade, which is frequently not your existing customers. Screen on category behavior rather than self-reported interest, keep the screener to two or three items, and avoid naming your brand in it. Existing customers are the right audience only when the message targets people already using the product.
Bottom line
Run a message testing survey to narrow a list of candidate messages before you spend money putting one in market, and size it so the narrowing is real rather than apparent.
Two decisions carry the study. Size each cell from the difference you need to detect rather than from the conventional figure, because the conventional figure typically detects nothing smaller than about 19 points.
And field an unchanged control message in every study, because it is the only benchmark this method will ever have that is matched on sample, instrument and time.
Then hand the winner to an experiment. Screen with the survey, decide with the test, and report the per-cell n, the minimum detectable difference and the number of comparisons you ran.
Your next study is one correctly powered monadic test with an in-study control, and the monadic concept test template is the starting structure.
Ignore any sample-size figure printed on that template or on any other, and size the cells from the power table above, which is the one number in this guide that is derived and not inherited.