Executive answer
A brand tracking survey measures the same brand metrics, on the same instrument, at a fixed interval, so a marketing team can tell whether awareness, consideration and mental availability moved. Running one well takes four decisions: what to measure, how often, who to measure, and how large each wave has to be. The fourth decision is the one most trackers commonly get wrong, and it is the one that decides whether the first three matter.
Most published brand trackers report sampling noise as movement. At the common default of 1,000 completes per wave, a quarterly tracker can detect a true six-point change in aided awareness at a 50 percent baseline. Published trackers frequently report and act on movements of one to three points, which typically sit inside the noise band.
The design this guide recommends is prompted recognition as the core awareness metric, a category-entry-point battery in place of a generic attribute battery, discrete waves with independent samples and fixed quotas, fielded to the category population.
Go continuous only when your annual sample budget can give a rolling window at least as large as the wave it replaces.
What a brand tracking survey is
A brand tracking survey is a repeated cross-sectional study that fields a frozen questionnaire to a fresh sample from the same population at a fixed interval, producing a time series of brand metrics.
Each wave is an independent sample rather than a repeat interview with the same people, which is what makes wave-over-wave comparison a two-sample problem and not a paired one.
What the instrument measures
The instrument typically carries four families of measure: awareness in prompted and unprompted forms, consideration and preference, brand associations as an adjective battery or a category-entry-point grid, and usage, which anchors the rest.
How a tracker differs from brand measurement
Brand tracking is not brand measurement in general.
A tracker is defined by repetition under constant conditions, so the constraint that matters most is question stability, not question quality.
Aided, unaided and top-of-mind awareness
Three terms get used interchangeably and should not be. Aided awareness, also called prompted recognition, asks whether a respondent recognises a brand from a shown list. Unaided awareness, also called spontaneous recall, asks which brands come to mind in a category with no list shown. Top-of-mind awareness counts only the first brand named.
Those three are frequently reported as though a brand could be separately good and bad at each.
Laurent, Kapferer and Roussel showed they are one latent trait measured at three levels of difficulty, which changes how all three should be read.
Mental availability and category entry points
Mental availability is a distinct construct again. It is the propensity of a brand to be noticed or come to mind in buying situations, reflecting the quantity and quality of memory structures, and it is measured against category entry points.
When to use a brand tracking survey
Use a brand tracking survey when you need to know whether a brand metric changed over time, and you are willing to field the same instrument repeatedly to find out.
The method answers change questions, not causal ones, and not diagnostic ones.
Four situations that justify a tracker
Four situations where a tracker earns its cost:
- Sustained category investment
- Competitive share shifts
- Post-launch brand building
- Multi-market brand comparison
A tracker is generally most valuable where the metric is slow-moving and the investment behind it is large enough that nobody can otherwise say whether it worked.
Rather than treating the tracker as a general brand dashboard, treat it as an instrument aimed at one question over time.
It is also commonly the right instrument for a competitive read under identical conditions.
Measuring your brand and four competitors in the same wave, on the same list and frame, gives a comparison that transfers in a way cross-industry averages never do.
What to commission a tracker for
Rather than commissioning a tracker to prove a campaign worked, commission it to detect drift, then use a designed experiment to attribute the movement. A tracker only does the first job honestly.
The decision rule: run a brand tracker when you need a repeated, comparable measure of a slow-moving brand metric across a category population, and you can fund a wave large enough to detect the change you would act on.
When not to use a brand tracking survey
A brand tracking survey cannot tell you six things, and for each one there is a better instrument. These are limitations of the method and not of any vendor.
The three the method loses outright
A tracker cannot tell you whether your marketing caused the movement. A tracker is an observational time series with no counterfactual. It can tell you the number changed, and it cannot separate your campaign from a competitor's price move, a category shock, or seasonality. Use a geo-holdout or matched-market test, or a market-response model with the tracker's metrics entered as mediators, which is the design Srinivasan, Vanhuele and Pauwels (2010) and Hanssens et al. (2014) both use.
A tracker cannot tell you what your brand associations mean, if you use a generic attribute battery. Romaniuk, Bogomolova and Dall'Olmo Riley (2012) examined 45 datasets across categories, countries and methods and found brand association responses strongly and systematically linked to past brand usage, qualitatively and largely quantitatively. An adjective battery is substantially a usage measure in a perception costume, so a trust score that rises after a good sales quarter has often told you nothing new. Use a category-entry-point battery, which asks which brands come to mind in a situation instead of which adjectives attach to a brand. For the differentiation question specifically, use distinctive-asset testing.
A tracker cannot tell you which message or creative moved the number. A tracker is an outcome measure with no manipulation in it. Use monadic message testing, or MaxDiff across claims, which is a first-party survey question type with a published analysis method.
The three that depend on how the tracker is resourced
A tracker cannot tell you whether a small quarter-on-quarter movement is real. At the sample sizes typically fielded, most reported movements sit inside the noise band. This is typically a limitation of the instrument as commonly resourced rather than of the method in principle. Use the same tracker with pooled waves and an annual read, or a properly powered experiment when the decision genuinely turns on a small effect.
A tracker cannot tell you why non-buyers do not buy. Awareness and consideration batteries generally measure incidence, not reasoning, and open text gets closer without getting there. Use moderated or AI-moderated depth interviews, which are qualitative instruments in a different category from survey platforms and answer a different question from a small number of participants. A tracker generalises from many people at low depth and depth interviews do the opposite, so the two are generally complements.
A tracker cannot tell you what a movement is worth. A tracker produces percentage points, not currency. Use a market-response or marketing-mix model that connects the attitudinal metric to sales, for which Hanssens et al. (2014) is the template. For the price question specifically, use Gabor-Granger or Van Westendorp.
Three of the six are dimensions the method loses outright, and a readout that does not name them will be asked about them anyway. No counterfactual, no manipulation, and no reasoning from the people who did not buy.
Two of those six are load-bearing and should not be dropped from any tracker readout: the absence of a counterfactual, and the usage contamination in an adjective battery.
Rather than discovering them in a board meeting, state them in the tracker's own documentation.
Study design, and the three forks that decide everything else
Brand tracker design comes down to three forks that interact: what you measure, how often you measure it, and who you measure.
Rather than deciding them separately, decide them together against one number, which is your annual sample budget.
Fork A: the funnel battery or the mental availability battery
The traditional battery typically walks a funnel. Awareness, then consideration, then preference, then purchase, plus a grid of attitude statements. The mental availability battery replaces the attitude grid with category entry points, which are the situations and occasions that trigger a category purchase, and asks which brands come to mind in each.
The research favours the second. Romaniuk and Sharp (2004) redefined salience as a brand's propensity to be noticed or come to mind in buying situations, and distinguished it explicitly from awareness and from attitude.
The objection to switching batteries is usually cost. Vaughan, Corsi, Beal and Sharp (2021) removes it. Their finding, verbatim: "From a practical market research perspective, adding MA metrics to existing brand health tracking will have no data collection costs where brand perceptions are already being measured."
Consideration is the exception, earning its place on different evidence. Srinivasan, Vanhuele and Pauwels (2010) found advertising awareness, brand consideration and brand liking account for almost one-third of explained sales variance across more than 60 brands of four consumer goods. Keep consideration, and stop drawing arrows between the stages.
The recommendation: keep prompted recognition and consideration, replace the adjective battery with a category-entry-point grid, and drop the funnel arrows from the reporting.
Fork B: discrete waves or continuous fielding
Discrete waves field a fixed sample at a fixed interval. Continuous fielding collects a few completes every day and typically reports a rolling window.
Continuous wins two real things. It timestamps a movement against a dated event, which a quarterly wave cannot. And it supports the early-warning use Srinivasan, Vanhuele and Pauwels evidence with their wear-in analysis, where mind-set metrics move before sales do.
Continuous loses three. Overlapping rolling windows are not independent samples, so a two-sample significance test between them is invalid. The number of comparisons explodes, which is a problem the next section quantifies. And at realistic budgets each window is generally too small to detect anything.
The arithmetic is checkable, and one widely circulated piece of vendor guidance fails it. Quali-Fi's always-on brand tracking guide recommends a Standard tier of 20 to 30 completes per day and states that a four-week rolling average "provides the same statistical reliability as a quarterly wave but updates every week." At 20 to 30 per day, four weeks is 560 to 840 completes and a quarter is roughly 1,820 to 2,730. The margin of error on the rolling window is about 1.8 times wider, not the same. No study is cited for the tiers or for the equivalence claim.
The recommendation: run discrete waves unless your annual budget can give a rolling window at least as large as the wave it replaces. At the volumes vendors publish, it usually cannot.
Fork C: the category population or your own users
A tracker fielded to your own users measures perception among people who already chose you.
That is a legitimate study and not a brand tracker, because the metric a tracker exists to move is awareness outside the franchise.
Romaniuk (2006) is the sourced reason this matters. Unprompted approaches were less likely to elicit associations from non-brand users and smaller share brands, which is exactly the group a growing brand needs to hear from.
The recommendation: field to the category population, defined by a category buyership screener, and measure your own users as a segment inside that frame rather than as the frame itself.
The tracker design tree
Work the branches in order, because each answer constrains the next.
| Decision point | If yes | If no |
|:-------------------------------------------------------------:|:--------------------------------------------:|:---------------------------------------------------------------------:|
| Can you fund at least 1,000 completes per wave? | Continue to cadence | Reduce to two waves a year at full size instead of four at half size |
| Do you need to attribute movement to a dated event? | Continuous is worth considering | Discrete waves |
| Can a rolling window match the wave it replaces in size? | Continuous is viable | Discrete waves |
| Is the category population reachable outside your product? | Panel or link fielding to the category | Your study is a customer perception study, and should be labelled one |
| Can you hold more than three quota cells constant? | Interlocking age by gender by category frame | Coarse frame only, and report age drift as a caveat |
| Does your platform return a wave-over-wave significance test? | Report in platform | Export, concatenate and test outside the platform |
hree of those branches are foreclosed by platform limits rather than by methodology. The honest thing is to mark them, not discover them at wave two.
Assumption flagged for review (A1): wave structure. This guide assumes each wave is its own study, with the wave identifier carried in the study name plus an attribute stamped on every response at collection. Whether a tracker is better built as one study run repeatedly is unresolved against published Sprig documentation. The one-study-per-wave design works under either answer, and it makes the concatenation step in the analysis chapter correct either way.
Sprig, and the constraint. Sprig documents response-based quotas built on screener questions, with the changelog for July 1, 2026 describing up to three screener questions used to define user segments and a target set for each. Three answer-based cells is enough for a coarse age and gender frame and is not enough for anything finer. No interlocking or nested quota cells are documented, quotas cannot be set on pre-existing attributes, and filtering responses by quota on the Data or Responses tab shows a maximum of 3,000 responses. That ceiling is the structural constraint this method runs into first.
The platform constraint belongs in the design chapter rather than the product chapter, because it changes the design. Sprig can field this method and does not document the analysis half of it. There is no documented trend view, no documented cross-study comparison that returns numbers, and no documented significance testing anywhere in the product.
A brand tracker is a time series by definition. So the comparison happens outside the platform, in the prompt shown later or in a spreadsheet, and the export-and-concatenate step is part of the method and not a workaround.
Writing the instrument, and freezing it
A brand tracker instrument is written once and then frozen, because a tracker measures change and an edited question measures something else.
The most common way a tracker silently breaks its own time series is typically an improvement made to a question between waves.
Question order is a design decision rather than a formatting one. Unprompted questions always precede prompted ones, because a shown brand list contaminates every recall measure after it, and demographics generally go last.
Romaniuk (2006) gives a second ordering reason. Unprompted methods were more prone to priming and inhibition effects over the duration of the attribute battery, so the longer an unprompted battery runs, the less it measures what it claims.
Rather than randomising the brand list itself, randomise its display order and hold the list membership constant.
A brand added at wave three changes the aided awareness denominator for every brand on the list.
The frozen instrument spec
Lift this without modification, and keep the order, because the order is part of the spec.
- Category buyership screener
- Unaided recall, open text
- Aided recognition, randomized list
- Category entry point grid
- Consideration
- Brand and category usage
- Distinctive asset check
- Demographics
The wording that goes with each step, in launch order:
- Screen on category purchase or intent in a stated window, before any brand is named
- Ask which brands come to mind when thinking about the category, as open text, with no list
- Show the frozen brand list in randomized order and ask which brands are recognised
- Present a pick-any grid of brands against buying situations, asking which brands come to mind in each
- Ask which listed brands would be considered for the next purchase
- Ask which brands were bought in a stated recent window
- Show a logo, colour or character stripped of the brand name and ask which brand it belongs to
- Collect age, gender and region last, matching your quota frame categories
The category entry point grid is where most of the design effort should go.
Build the situations from the category and not from the brand, which is Romaniuk's own design principle: "Design for the category, analyse for the buyer, report for the brand."
Six to ten entry points is a working default and not a published rule. Fewer typically under-covers the category, and more adds length the tracker would rather give to sample size.
Length discipline matters more here than in a one-off study, because a tracker pays its length cost every wave. Add a question only if its movement would change a decision.
One practical rule that nobody writes down: version-lock the instrument at wave one and keep the wave-one questionnaire as a file.
An edited question label is a new question, even when the text reads identically to a human, because the export column name changes and the concatenation step silently produces two columns.
How many competitors to put on the brand list
A working default, and not a published rule: your brand, the category leader, the three to five brands a buyer would consider, and one growth challenger.
Rather than listing every brand in the category, list the set you would act on.
A longer list is often not free. Aided recognition rises mechanically with list length for small brands, so a list that grows between waves inflates the very metric the tracker reports.
Include an anchor brand you expect to be stable. When your number moves and the anchor moves with it, the movement is usually fieldwork rather than brand.
Include a fictional or dormant brand as a fraud check if the list can carry it.
A respondent who recognises a brand that does not exist is telling you something useful about your sample.
Sprig, and the constraint. The Design Agent generates a fully programmed study from an uploaded document, with response options, logic and randomization already in place, flags unclear questions and estimates completion time. That is genuinely useful for building wave one. The constraint is one nobody expects: for a tracker the instrument must be frozen, so regenerating the questionnaire between waves is the fastest way to break the time series the tracker exists to produce. Generate wave one, then stop generating.
Scale and scoring choices
Formats and scoring rules
Use pick-any binary formats for awareness, consideration and category entry points, and reserve rating scales for the small number of attitude measures you keep.
Binary formats are what the brand tracking literature measures, and they are what makes a two-proportion test the right analysis.
Aided awareness is generally scored as the proportion of the sample selecting the brand from the shown list.
Unaided awareness is scored as the proportion naming the brand in open text, which requires coding and is typically where most measurement error enters.
Top-of-mind is scored as the proportion naming the brand first. It is the least stable of the three and the most commonly reported.
Romaniuk, interviewed on Better Brand Health in 2023, states the case against it plainly: "Top of mind awareness is particularly bad because it has all the bad biases.
It biases against the audience that we want to know most about (non-users) and small brands."
Mental market share is the scoring choice specific to the category entry point grid. Count the brand's mentions across all entry points, divide by total mentions for all brands across all entry points, and report the share. Rather than reporting an average per entry point, report the share, because the share is what competes.
Why raw percentage gaps are not comparable
The nonlinearity problem applies to every one of these scores. Laurent, Kapferer and Roussel (1995) showed that the relationships between aided, spontaneous and top-of-mind awareness are close but highly nonlinear, and can be linearized by a logistic transformation, which amounts to describing the answering process with a Rasch model where brand salience is equivalent to a student's competence and question type is equivalent to test difficulty.
Two consequences follow and neither is commonly applied. A raw percentage-point gap is not comparable across levels of the scale, so a five-point gap at 20 percent and one at 80 percent are not the same distance. And a small brand's top-of-mind score is typically unstable for a structural reason rather than a fieldwork one.
Rating scales and open-text coding
Where you keep rating scales, keep them identical across waves and report top-two-box instead of means.
A mean on a five-point scale moves when the shape of the distribution changes without the top of it changing at all, which is a poor property for a tracked metric.
Five-point and seven-point scales both work. What does not work is changing between them, or relabelling a midpoint, either of which converts a trend into two unrelated series.
Code unaided recall against a fixed codebook built at wave one, not against a list regenerated each wave.
Spelling variants, sub-brands and parent companies each commonly need a stated rule, written down once, applied identically every wave.
If you intend to build norms across categories, transform first. If you intend only to track your own brand against its own history on one instrument, raw proportions are generally fine, which is the case for most readers of this guide.
Sample size and precision, derived
No professional body publishes a defensible rule for brand tracking wave size or for the minimum detectable wave-over-wave change. ESOMAR, the ESOMAR and GRBN online sample quality guideline, and the ARF were all checked, and none publishes brand tracking wave sizing, cadence, or wave-over-wave testing standards. Kantar's own sample-size explainer is general research guidance, giving a typical range of 300 to 1,000 respondents for general population studies with no tracking-specific rule attached.
So this guide publishes the arithmetic instead of a borrowed number. Every figure below was computed for this guide and independently recomputed.
The load-bearing distinction, which almost nothing on this topic makes: the margin of error on a single wave is not the standard against which a wave-over-wave change is judged. A difference between two independent samples carries the error of both, so the critical difference is roughly 1.41 times the single-wave margin, and the sample needed to detect a change with adequate power is larger again.
Everything below assumes 95 percent two-sided confidence, 80 percent power, two independent samples of equal size, no design effect and no weighting.
Apply a design effect for weighting and clustering, which will typically make every number worse.
Table A, derived: aided awareness at a 50 percent baseline
| Completes per wave | Single-wave margin of error | Observed change reaching p < 0.05 | True change detectable at 80% power |
|:------------------:|:---------------------------:|:---------------------------------:|:-----------------------------------:|
| 200 | plus or minus 6.9 points | 9.8 points | 13.8 points |
| 300 | plus or minus 5.7 points | 8.0 points | 11.3 points |
| 500 | plus or minus 4.4 points | 6.2 points | 8.8 points |
| 1,000 | plus or minus 3.1 points | 4.4 points | 6.2 points |
| 1,500 | plus or minus 2.5 points | 3.6 points | 5.1 points |
| 2,000 | plus or minus 2.2 points | 3.1 points | 4.4 points |
Table B, derived: unaided awareness at a 20 percent baseline
Readers frequently assume smaller numbers need less sample. They need less in absolute points and more in relative terms.
| Completes per wave | Single-wave margin of error | Observed change reaching p < 0.05 | True change detectable at 80% power |
|:------------------:|:---------------------------:|:---------------------------------:|:-----------------------------------:|
| 200 | plus or minus 5.5 points | 7.8 points | 12.3 points |
| 500 | plus or minus 3.5 points | 5.0 points | 7.5 points |
| 1,000 | plus or minus 2.5 points | 3.5 points | 5.2 points |
| 2,000 | plus or minus 1.8 points | 2.5 points | 3.7 points |
Table C, derived: completes per wave needed to detect a given true change at 80 percent power
| Baseline | 2 points | 3 points | 5 points | 8 points | 10 points |
|:--------:|:--------:|:--------:|:--------:|:--------:|:---------:|
| 20% | 6,510 | 2,943 | 1,094 | 447 | 294 |
| 35% | 9,041 | 4,042 | 1,471 | 583 | 376 |
| 50% | 9,806 | 4,356 | 1,565 | 609 | 388 |
The formulae, so the tables are checkable instead of trusted. The single-wave margin of error is 1.96 times the square root of p times one minus p divided by n. The critical difference between two independent proportions of equal size is 1.96 times the square root of two times p times one minus p divided by n, using the pooled proportion under the null. The detectable true change at 80 percent power adds 0.84 standard errors, computed against the two separate proportions and not the pooled one.
The headline, stated in prose rather than left in a table. At the common default of 1,000 completes per wave, a quarterly tracker can detect a true six-point change in aided awareness at a 50 percent baseline. Published brand trackers routinely report and act on movements of one to three points. Detecting a true three-point change at a 50 percent baseline needs about 4,400 completes per wave, which is roughly four and a half times the default. Most brand trackers are reporting noise as movement, and this is the arithmetic that shows it.
The multiple-comparisons multiplier
Twelve metrics reported quarterly is 48 wave-over-wave comparisons a year. At an alpha of 0.05, with no true change anywhere in the data, the expected number of significant movements is 2.4 and the probability of at least one is 91.5 percent.
That is often why a tracker seems to have news. Weekly reporting of the same twelve metrics is roughly 624 comparisons a year, with an expected count of about 31 spurious findings.
The practical rule: nominate two or three primary metrics before the wave fields, apply a correction across the rest, and report the rest as directional and not as findings.
Rolling windows are not independent samples
A three-month rolling window shifted forward by one month shares two thirds of its fielding period with the prior window.
Comparing the two with a two-sample test is invalid, because the test assumes independence the data does not have.
Compare non-overlapping windows, or model the series properly. Running the test anyway and caveating it is not a substitute for either.
This guide publishes no minimum sample floor and no minimum number of waves before a trend is readable. No source supports either, and a threshold with no source behind it is worse than an absence, because it gets quoted.
Verification note (A3). The Sprig brand awareness template was opened and checked on September 11, 2026. It states no sample size, no participant count and no cadence recommendation, so there is nothing on it for this derivation to contradict. If a figure is added to that template later, the fix belongs on the template rather than in this derivation.
Audience, targeting and screening
Defining and freezing the sample frame
Define the tracker's sample frame as the category population, screened on category buyership or purchase intent in a stated recent window, and hold that definition byte-identical across every wave.
The frame is the thing being held constant, and a frame that drifts turns a tracker into a sequence of unrelated studies.
Screening comes before any brand is named, because a screener that mentions your brand selects for people who know it, which inflates every awareness measure downstream.
Quotas, not targeting, hold composition constant. Targeting controls who is invited and quotas control who ends up in the data, so only the second protects a comparison.
Mecredy, Wright, Feetham and Stern (2023) supply the sourced reason age quotas matter most. Across 1,862 respondents in five markets and four age groups, recognition, recall and consideration showed an inverse-U shape with peak cognitive performance at age 56. A tracker whose age composition drifts between waves will move when nothing about the brand has.
Set quotas on the smallest frame that generally protects the read. Age band and gender is typically the minimum. Category buyership is frequently the more important third cell, because a wave that accidentally recruits more heavy category buyers will show more of everything.
Rather than weighting after the fact, quota before the fact where you can.
Weighting repairs composition at the cost of effective sample size, which the tables in the previous section do not account for.
Interlocking cells are generally better than marginal ones. A marginal quota on age and a separate marginal quota on gender can still produce a wave that is mostly young men and mostly older women, which two marginal quotas will report as balanced.
Screening business-to-business categories
Business-to-business categories need a role screener as well as a category one, because awareness among people who cannot buy is a different metric. Screen on purchase involvement, then quota on it.
Where the category population is unreachable, say so and relabel the study. A customer perception tracker is useful and it is not a brand tracker.
Record each wave's realised composition, not only its intended one.
Composition is the first thing anyone asks about when a number moves, and reconstructing it later is typically impossible.
Sprig, and the constraint. Sprig documents more than 300 targeting attributes across demographic, professional, behavioral and firmographic categories, plus attribute piping and attributes passed through a URL. That means a wave identifier and every segment variable can be stamped onto each response at collection, which is precisely what the wave-over-wave analysis needs and what most trackers assemble by hand afterwards. The constraint is the distinction above: attributes target, they do not quota. Targeting on an attribute does not hold its distribution constant between waves.
Fielding and delivery
Holding delivery constant across waves
Field every wave through the same channel, in the same field window length, covering the same days of the week.
Delivery mode is a confound, and changing it between waves produces a movement that has nothing to do with the brand.
The choice of channel follows from fork C. A category population is reached through an external research panel or a link distributed outside your product. Your own users are reached in product or by email.
In-product delivery is the wrong instrument for a brand tracker and the reason is generally structural rather than technical.
An in-product tracker measures whoever was in your product during the field window, which is close to the opposite of the population the method requires.
Panel and in-product hazards
Panel fielding frequently introduces its own hazards. Panel composition shifts over time, incidence rates change frequently, and a provider's sourcing mix is rarely disclosed. Rather than assuming stability, ask your provider to hold sourcing constant and to say when it changes.
Field window length is a design parameter treated as an operational detail.
A two-week window and a six-week window typically reach different people, because those who answer on day one differ from those who answer on day 30.
Keep the incentive constant, because an incentive change is generally a composition change.
Device and mode effects
Mode effects are frequently large enough to swamp a brand movement. A survey taken on a phone produces different open-text recall from the same survey on a desktop, because typing effort differs and unaided recall is scored on what people bother to type.
So hold device mix roughly constant, or measure it and report it beside the awareness numbers. Device is part of the instrument, not a demographic.
Sprig, and the constraint. One study definition can be delivered through in-product surveys on web and mobile, email with a custom sending domain, shareable links, QR codes and external research panels, which matters when a tracker's category population sits outside the product. The constraints are two. There is no native SMS delivery, though a link can be sent through your own tool. And Sprig Panels is documented as Enterprise and link surveys only, with no published minimum or maximum sample size, fielding turnaround time, incidence rate or country coverage, which is a real risk for a method whose value rests on identical fielding conditions wave after wave.
Timing and cadence
Choosing the interval
Set cadence from category purchase velocity and from your annual sample budget, not from a calendar convention.
The right interval is the one where a decision could plausibly change between waves and where each wave is large enough to detect a change you would act on.
Quarterly is the common default and it is typically a budget artifact rather than a methodological finding. Two waves a year at 2,000 completes will detect a 4.4-point change. Four waves at 1,000 will detect a 6.2-point change and will cost the same.
The trade-off is stated plainly: more waves buy resolution in time and lose resolution in measurement. Fewer waves do the reverse, so rather than defaulting to quarterly, decide which axis your question lives on.
Romaniuk, Wight and Faulkner (2017) is the one longitudinal design study retrievable on this question, with sample sizes of approximately 300 whisky consumers per wave in three countries.
Their conclusion is a caution against a single default: the choice of brand awareness measure depends on the brand's market share and on whether the team wants higher sensitivity or higher stability.
Hold the field window to the same weeks of the year where seasonality is real in your category.
A tracker fielded in November and then in February has measured a season as well as a brand.
Changing cadence, and pooling waves
Change cadence only between years, and treat the change as a break in the series.
A tracker that moves from quarterly to monthly has not gained resolution, it has started a second tracker.
Pooling is generally the underused move. Two adjacent waves pooled give roughly the precision of one double-sized wave, and the pooled estimate is frequently the only number in the deck with enough power to carry a decision.
Report the pooled annual number as the headline and the individual waves as the operational read.
This guide states no minimum number of waves before a trend is readable. No source publishes one. A slope fitted to three points is a slope fitted to three points, and the honest report generally gives the uncertainty rather than a rule about when the line becomes trustworthy.
Quality control and data hygiene
In-field hygiene checks
Run the same hygiene checks every wave and record the results beside the numbers, because a movement in data quality looks exactly like a movement in brand.
Speeders, straight-liners and duplicate respondents commonly shift awareness upward, since selecting more boxes is faster than reading them.
Check the anchor brand first, because a stable competitor that moves with your brand is typically evidence about your fieldwork and not about either brand.
The wave-integrity pre-launch checklist
Run this in launch order, and note that the last four are the ones nobody typically checks.
- Confirm the screener wording and the qualifying window are unchanged from the prior wave
- Confirm the quota cells and their targets are unchanged, including the category definitions inside them
- Confirm the incentive, the field window length and the days of the week covered all match
- Confirm the brand list is byte-identical to last wave, including order of options and spelling
- Confirm no question label was edited, since an edited label produces a new export column
- Confirm the wave identifier is stamped on every response at collection rather than added later
- Confirm the unaided recall codebook is the wave-one codebook and not a regenerated list
- Confirm the prior wave's raw export is archived somewhere that is not a link with an expiry
Apply bot and fraud filtering identically across waves and record the filter version, because a provider that tightens its screening between waves has changed your sample without telling you.
Code open text against the frozen codebook and hold back a sample for a coding check. Rather than trusting a single pass, recode a random subset and compare.
Rather than cleaning aggressively, clean identically. A cleaning rule applied at wave three and not at wave two creates a movement in the series by itself.
Analysis: the core calculation
The core calculation in a brand tracker is a two-proportion test between two independent waves, run on each primary metric, with the comparisons counted and corrected.
Everything else in the analysis is generally preparation for that test or interpretation of it.
Start by building the wave field. Concatenate the raw exports into one file and derive a wave column, because no survey platform export contains one.
Then compute each metric as a proportion within each wave. Aided awareness is the count selecting the brand divided by respondents who reached that question, not by total responses, which differ once anyone drops out mid-survey.
State the denominator explicitly for every metric and keep it constant. A metric whose denominator changes between waves is not a metric, it is two.
The two-proportion test, step by step
Run this for each primary metric, comparing the current wave against the prior one.
- Compute the proportion and the base size for the metric in each of the two waves
- Compute the pooled proportion as the total successes across both waves divided by the total base
- Compute the standard error as the square root of the pooled proportion times its complement, times the sum of one over each base
- Divide the difference between the two proportions by that standard error to get the z statistic
- Compare the absolute z statistic against 1.96 for a two-sided test at 95 percent confidence
- Report the difference, its confidence interval, and the base sizes together
Report the interval, not only the verdict. A three-point movement whose interval runs from minus one to plus seven is a different story from one whose interval runs from plus one to plus five, and a significance flag hides the difference.
Counting and correcting the comparisons
Count every comparison the report makes, including those made by looking at a chart, then apply a correction across the secondary metric set.
A Holm or Benjamini-Hochberg correction is appropriate and easy to apply.
Rather than reporting every metric that crossed 0.05, report the primary metrics against an uncorrected threshold nominated in advance and the secondary ones against a corrected threshold.
The correction is not a statistical nicety here, because the earlier arithmetic shows a 91.5 percent chance of at least one spurious significant movement per year in a twelve-metric quarterly tracker with no true change anywhere.
Mental market share and the competitive read
Compute mental market share per wave as the brand's total mentions across all category entry points divided by all brands' total mentions. Test the share the same way, as a proportion.
The competitive read is the one comparison that transfers. Your brand against four competitors, measured in the same wave under identical conditions, holds the confounds constant, which no comparison against an external norm does.
Sprig, and the constraint. Theming applies AI open-text analysis to unaided recall and to any free-response probe, with real-time regeneration and traceability back to individual responses, which removes most of the manual coding load on the recall question. The constraint is the important one and it is easy to miss: no minimum response count is documented for Theming, no accuracy or reproducibility claim is published for theme counts, and themes regenerate. A regenerated theme list is not the same measuring instrument as last wave's, so a theme count is not a trackable metric. Use Theming to explore the verbatims, and code recall against your frozen codebook for anything you intend to trend.
Benchmarks, and why there is no useful one
The benchmark answer is an absence
There is no publicly retrievable, sourced cross-industry benchmark for brand awareness. Every norm table we could find cites no study, no sample and no database, and the one large commercial dataset that discloses its scale, Tracksuit's 6,075 data points from 781,615 survey responses across 1,440 brands, does not publish the figures themselves.
That is a verified finding and not a hedge, checked across vendor guides, agency blogs and the professional bodies, and re-confirmed on September 11, 2026.
Representative of what circulates: one widely shared statistics roundup lists consumer goods at 70 to 90 percent recognition and technology brands at 60 to 80 percent, alongside a dozen other figures, and not one of them carries a source.
This guide reproduces no uncited figure, not even to illustrate a range, because an uncited number that appears inside a guide about measurement error acquires a credibility it has not earned.
Three sourced reasons norms do not transfer
Three sourced reasons cross-industry norms cannot transfer, which is the part a vendor benchmark page will not print.
Awareness scales with brand size and category penetration. Ehrenberg, Goodhardt and Barwise (1990) state it directly: "In any given time period, a small brand typically has far fewer buyers than a larger brand. In addition, its buyers tend to buy it less often." A cross-industry average is an average across brand sizes, so it is not a target for any particular brand.
Awareness scores are nonlinear in the underlying trait. Laurent, Kapferer and Roussel (1995) showed the three measures linearize only under a logistic transformation, so averaging raw percentages across categories with different question difficulty is not a meaningful operation.
The attitude-to-sales relationship itself varies across brands and categories. Hanssens et al. (2014), verbatim: marketing-attitude and attitude-sales relationships "are predominantly stable over time but differ substantially across brands and product categories."
Your benchmark is your own prior waves, on the same instrument, in the same category, fielded to the same frame, plus the competitor set measured in the same wave. Rather than asking whether a 42 percent aided awareness score is good, ask whether it is higher than last wave by more than the critical difference, and how it sits against the competitors measured beside it.
Two circulating thresholds this guide refuses
Two thresholds in circulation deserve to be named and refused. Quali-Fi publishes an alert threshold of "typically 3-5 points for awareness metrics" with no citation, and at the sample sizes recommended on the same page that threshold sits below the detection limit. Tracksuit publishes the advice that "unaided awareness is a metric best tracked when you've achieved at least 40% brand awareness," also with no citation. Both are commonly repeated and neither has a source behind it.
Interpreting and acting on the result
Read a tracker in three passes: fieldwork first, competitors second, your own brand last.
Reading your own number first is typically how teams talk themselves into a story about a composition shift.
Pass one asks whether the wave was fielded the same way. Compare realised composition, completion rate, median duration and device mix against the prior wave before looking at any brand metric.
Pass two asks what the competitors did, because if every brand on the list moved in the same direction, the category or the sample moved, not your brand.
Pass three asks whether your metric moved by more than the critical difference for its base and baseline.
If it did not, the correct report is that the metric did not move, stated plainly rather than softened into a direction.
A tracker is an instrument for detecting change, so the most common correct finding is no change. A readout that reports movement every quarter is often describing its own noise.
The funnel-leak trap
Do not diagnose a funnel leak from a cross-brand comparison. This is the most common analytical error in the method.
Ehrenberg, Goodhardt and Barwise (1990) established double jeopardy: a small brand has fewer buyers and its buyers buy it less often.
The consequence for trackers is that funnel-stage scores and the conversion ratios between them scale with brand size.
So a small brand's apparent conversion problem between awareness and consideration is usually arithmetic and not a problem. Compare your conversion ratio to your own prior waves, not to the market leader's.
Turning a movement into an action
A tracker movement is not an action on its own, and the honest workflow generally routes it into a second study, which is what the tracker is for.
- Confirm the movement exceeds the critical difference before escalating it to anyone
- Check whether the competitor set moved in the same direction in the same wave
- Identify which category entry points drove a mental market share change
- Design a message test or a geo-holdout to attribute the movement to a cause
- Re-read the metric at the next wave before committing budget against a single wave
What a tracker can support on its own is a resource-allocation argument across entry points.
If your brand is absent from three of eight buying situations, that is a brief for creative work and it needs no causal claim to be actionable.
Report a tracker as a standing document and not a quarterly surprise.
The version that earns trust states the instrument, the frame, the base sizes, the critical difference, the comparison count and the correction.
Running the wave-over-wave analysis with Claude or ChatGPT
What the analysis produces
The analysis produces one table per metric showing each wave's proportion, base size, the difference against the prior wave, the critical difference at that base and baseline, and a verdict of moved or did not move.
It also produces a count of comparisons run and a corrected threshold applied to the secondary metric set.
That is the whole deliverable, and it is small, repetitive and arithmetic, which is generally the kind of work a language model does well when told to compute rather than estimate.
What it does not produce is an explanation. Rather than asking a model why the number moved, ask whether it moved, and route the why into a designed study.
Setup and getting your data in
Export each wave as CSV, since the documented Sprig export columns are surveyId, surveyName, visitorId, href, createdAt, completedAt, userId, partnerAnonymousId, os, browser, userAgent, captchaScore, customMetadata, triggeringEvent, responseGroupUid, eventProperties, Q#_Question_Text, Q#_Response, Themes and Attributes_#.
There is no wave column. This is the step no other guide documents and it is the step every tracker analysis depends on.
The wave identifier is surveyName when each wave is its own study, or an Attributes_# column when you stamped a wave label at collection.
Either way you concatenate the exports and derive a wave field before analysis runs.
Assumption flagged for review (A1): wave structure. This chapter assumes one study per wave, with the wave identifier in surveyName and a stamped attribute as a redundant second copy. That design makes the concatenation step below correct whether a tracker is built as one study per wave or as one study run repeatedly, which is why it is the design this guide recommends.
Stamp the wave attribute even when the study name already carries it, because a study renamed once destroys the only copy of the wave label in the file.
Connector setup and theme creation are covered in the documentation and in the cross-tab analysis guides, and this chapter does not rebuild them.
Prompt one: per-wave metrics and the wave-over-wave test
What this prompt does: computes each brand metric per wave from a concatenated
brand tracker export and tests every wave-over-wave change with a two-proportion
z test.
What it returns: one table per metric with per-wave proportions, base sizes,
differences, critical differences and a moved / did not move verdict, plus a
count of comparisons run.
Use code to calculate this, not estimation. Every proportion, standard error
and z statistic must come from executed code, not from reading the file.
DATA: the attached concatenated export. Wave field: [WAVE_COLUMN, default:
surveyName]. Metrics: [METRIC_COLUMNS, default: aided awareness, unaided
awareness, consideration, category entry point mentions]. Primary metrics,
nominated before fielding: [PRIMARY_METRICS, default: aided awareness and
consideration].
EXCLUSION RULE: exclude blanks, nulls, whitespace-only cells, and any response
coded as a non-response or "prefer not to say" from both the numerator and the
denominator. Do not treat a blank as a negative. Report how many rows you
excluded per metric per wave.
BASE SIZE RULE: compute the base per metric per wave as the number of
respondents who reached that question, after exclusions, and report it beside
every proportion. Derive the underpowered threshold from my own wave size as
[THRESHOLD, default: the base below which the critical difference exceeds
[ACTIONABLE_CHANGE, default: 5] points] and flag any metric under it while
still reporting it. If a base is zero, report zero explicitly and compute no
proportion. Zero is under the threshold and must be named as zero rather than
omitted.
FOR EACH METRIC AND EACH ADJACENT WAVE PAIR:
1. Compute p1, n1, p2, n2, the pooled proportion and the z statistic.
2. Compute the critical difference at 95 percent two-sided confidence.
3. Report the observed difference, the critical difference, the 95 percent
confidence interval on the difference, and a verdict of moved or did not
move.
INDEPENDENT RECOMPUTE: recompute every confidence interval on the difference by
a second, genuinely different method. Use a bootstrap resample of the two waves
with [BOOTSTRAP_ITERATIONS, default: 10000] iterations. Do not recompute by
repeating the same formula.
HONESTY RULE: if the z test verdict and the bootstrap interval disagree on any
metric, report it in a section headed "Mismatches" with both numbers shown. Do
not silently reconcile them, drop the metric, or pick the more decisive one.
OUTPUT: a file named brand-tracker-wave-tests.csv with one row per metric per
wave pair, plus the Mismatches section printed in the chat.
Prompt two: multiple comparisons and the corrected read
What this prompt does: counts every wave-over-wave comparison made in the
tracker and applies a multiple-comparisons correction to the secondary metric
set.
What it returns: the comparison count, the corrected threshold, and a revised
verdict table separating primary from secondary metrics.
Use code to calculate this, not estimation. Count comparisons programmatically
from the results file and not by reading the table.
DATA: brand-tracker-wave-tests.csv from the previous prompt. Primary metrics
are [PRIMARY_METRICS, default: aided awareness and consideration]. Correction
method is [CORRECTION, default: Benjamini-Hochberg].
EXCLUSION RULE: exclude rows whose base size is zero from the comparison count
and from the correction, and list them separately. A cell with zero respondents
is not a comparison, and zero is under the threshold. Do not exclude rows
merely because they were not significant.
COUNT RULE: report total comparisons as metrics multiplied by adjacent wave
pairs, the expected false positives as that count times 0.05, and the
probability of at least one as 1 minus 0.95 raised to that count.
THEN apply the correction across the secondary metrics only, report primary
metrics against the uncorrected threshold labelled as nominated in advance, and
report secondary metrics against the corrected threshold labelled as
corrected.
INDEPENDENT RECOMPUTE: verify the corrected threshold by a genuinely different
method. Run a permutation test that shuffles wave labels within the
concatenated file [PERMUTATIONS, default: 5000] times and reports the empirical
rate at which a difference of the observed size or larger appears under no true
change.
HONESTY RULE: if the correction and the permutation disagree on whether any
metric survives, report both under a heading "Mismatches" and name the affected
metrics. Report the disagreement rather than resolving it.
OUTPUT: a file named brand-tracker-corrected-read.md containing the comparison
count, both thresholds, the two verdict tables, and the Mismatches section.
Reading the output
Read the base sizes before the verdicts, since a metric flagged as moved on a base of 180 is typically telling you about your sample and not your brand.
Read the Mismatches section second, because a mismatch typically means a small base or a proportion near zero or one, and both are reasons to report the metric as directional.
What to verify before reporting
Verify four things by hand before any number leaves the analysis. Researchers remain responsible for the arithmetic the model executed.
- Recompute one proportion and one critical difference in a spreadsheet and confirm they match
- Confirm the base sizes in the output match the row counts in the source file per wave
- Confirm the exclusion counts are plausible and roughly stable across waves
- Confirm the wave labels in the output map to the field windows you actually ran
Run the analysis twice from a fresh session and compare. Any difference between runs is a reason to distrust the output, not a rounding artifact.
Pitfalls
Seven failure modes, six of them documented in the literature and one specific to this method.
Asking a model what changed. Give a multi-wave file to a model and ask what changed, and it will find something every time, because there are dozens of comparisons in the file and it is not counting them. That is the multiple-comparisons problem wearing a language-model costume, and it is why the prompt specifies the test rather than the question.
Discovering themes is harder than applying them. Hill et al., PLOS Digital Health, April 2026, found models matched human analysts on deductive coding against an existing codebook, at 93.5 percent against 92.7, and were materially worse at inductive discovery, with strict fabrication at 1.2 percent and comprehensive error at 12.4 percent. The rule: theme inductively once, rebuild the list yourself, then have the model apply your list every wave after.
Rare themes get over-predicted, and rare themes are what get escalated. Ashwin, Chhabra and Rao, Sociological Methods & Research, May 2025, found non-random bias in 10 of 19 codes, correlated with respondent characteristics, with sparse codes systematically over-predicted.
Non-determinism. Thinking Machines Lab reported in September 2025 that 1,000 completions at temperature zero produced 80 unique outputs. Neither chat client exposes a seed, so a reportable number gets produced twice.
Lost in the middle. Liu et al., TACL 2024. Put the data below the instructions and split large files.
Sycophancy. Sharma et al., ICLR 2024, plus OpenAI publicly withdrawing a model update in April 2025 for being overly agreeable. Never ask a model to confirm the movement you already suspect.
Verbatims contain personal information nobody asked for. Free-text recall collects names, employers and grievances, and the training-data question turns on which service tier you are on and not on which vendor you chose.
Sprig, and the constraint. Sprig MCP publishes its governance rather than leaving it to a trust page: access is scoped to the authenticated user's role, calls are capped at 1,000 responses, response data is not used to train models, agents cannot launch or modify a live study, and there is an org-wide admin kill switch. That makes the export-and-analyse loop above tolerable to run repeatedly. The constraints are two. The 1,000-response cap means a multi-wave tracker cannot be pulled in a single call, so the concatenation happens across several. And MCP hands the model data to reason over, not a significance test, so the arithmetic above still has to be specified in the prompt. Researchers remain responsible for validating the test, the base sizes and the correction before any number reaches a deck.
The critique you should know
The measurement design this guide recommends sits on one side of a live academic disagreement, and both sides are peer-reviewed.
A reader who finds the opposition somewhere else and not here will discount everything above it.
The case against the traditional tracker
Romaniuk and Sharp (2004), in Marketing Theory, redefined salience away from top-of-mind and toward a brand's propensity to be noticed or come to mind in buying situations, reflecting both the quantity and the quality of memory structures, and distinguished it explicitly from awareness and from attitude. That paper is where the measurement fight starts.
Romaniuk (2023), Better Brand Health, Oxford University Press, extends the argument to the instrument itself. The publisher's framing is the case: brand health tracking is among the largest and costliest sources of brand performance insight marketers buy, and most existing trackers were designed before the growth framework they now serve. This guide cites the position and attaches no number to it, because the book was not read.
Ehrenberg, Goodhardt and Barwise (1990) supply the arithmetic consequence. Funnel-stage scores and the conversion ratios between them scale with brand size, so diagnosing a funnel leak from a cross-brand comparison is typically diagnosing double jeopardy.
Mecredy, Wright, Feetham and Stern (2023), in Marketing Letters, tested whether the funnel explains what it is most often used to explain.
Aggregated survey data across 1,862 respondents in five markets and four age groups showed an inverse-U shape for recognition and in some cases for recall and consideration, with peak cognitive performance at age 56, and found that age-related differences in brand awareness and consideration do not greatly impact age-related increases in loyalty.
The counter-critique
Srinivasan, Vanhuele and Pauwels (2010), in the Journal of Marketing Research, is the strongest single defence of the traditional battery.
Vector autoregressive modelling of the metrics for more than 60 brands of four consumer goods showed that advertising awareness, brand consideration and brand liking account for almost one-third of explained sales variance, with wear-in times revealing that mind-set metrics can be used as advance warning signals allowing time for managerial action before market performance is affected.
That finding does not rescue the funnel as a model of behaviour. It rescues the metrics as leading indicators inside a sales model, which is a narrower and more defensible claim.
Hanssens, Pauwels, Srinivasan, Vanhuele and Yildirim (2014), in Marketing Science, is a separate study with a separate sample: four-weekly marketing, attitude and sales data for 24 brands across four categories over seven years.
Combining marketing and attitudinal criteria improved prediction of brand sales performance, often substantially, with mean absolute percentage error of 12.0 percent for the combined model against 15.7 percent for the attitude model and 17.7 percent for the marketing-mix-only model.
Felipe Thomaz of Saïd Business School, University of Oxford, makes the sharpest public criticism of the framework this guide leans on.
Interviewed by Contagious in September 2022, he argues that the Ehrenberg-Bass model requires that brand market shares do not change and that brands are undifferentiated, and asks: "How can you use results of a model that requires no growth in order to explain and offer guidance on growth?" He also argues that in removing differentiation "he's broken branding," and that "There's not a clear measurement of availability."
Thomaz's stochastic frontier analysis, reported by AMI in October 2024, covers more than 1,000 campaigns and roughly one million customer journeys across 72 touchpoints and 11 media channels using Wavemaker and Kantar data. It reports that about 1 percent of campaigns achieved double-digit business lift, that average lift was below 2 percent, and that over 80 percent achieved 95 to 96 percent reach without converting it into business outcomes. Both accounts are trade interviews, so treat this as a named senior academic's public critique with its dataset described, not a published paper.
One further challenge, weaker and worth one sentence: Murray and Tenzer, writing in Marketing Week in January 2026 on an electric vehicle study whose framework is proprietary and whose sample sizes are unpublished, found only half of the ten most mentally available attributes were prototypical category entry points, and the largest brand was disadvantaged on product entry points while advantaged on fitness and social signals.
The position this guide takes
The funnel is a bad model of how buyers behave and a defensible set of leading indicators inside a sales model.
Those two statements are not in conflict, and almost every page on this topic treats them as though they were.
So measure recognition rather than top-of-mind, replace the adjective battery with category entry points, keep consideration because the market-response literature shows it earns its place, and stop drawing arrows between the stages.
Common mistakes
Nine mistakes commonly account for most broken trackers, and seven of them are design decisions made once at wave one.
The seven design mistakes
Editing the instrument between waves. A reworded question, a reordered option list or a new brand on the list all break comparability, and the edit is usually made as an improvement. Freeze at wave one and keep the original.
Reporting a movement inside the critical difference. This is the single most common error in the method, and the tables above are the correction. A movement smaller than the critical difference is not a small movement, it is no movement.
Treating the single-wave margin of error as the test. A difference between two independent samples carries the error of both, so the bar is roughly 1.41 times higher than a single-wave margin suggests.
Those three happen for one reason: nobody between the person writing the survey and the person reading the dashboard owns the arithmetic.
Naming an owner for the critical difference fixes more of these than training does.
Running the significance test across overlapping rolling windows. Adjacent rolling windows share most of their fielding period, so the independence the test assumes is absent.
Diagnosing a funnel leak from a cross-brand comparison. Double jeopardy means conversion ratios scale with brand size, so the small brand always looks like it has a conversion problem.
Letting sample composition drift. Age composition alone will move an awareness number, for the reason Mecredy et al. (2023) document, and a tracker without age quotas will report that drift as brand movement.
Reporting every metric that crossed 0.05. Twelve metrics reported quarterly gives a 91.5 percent chance of at least one spurious significant movement a year with no true change anywhere.
Trending a theme count. Themes regenerate, and a regenerated list is generally a different instrument, so a theme count is not a trackable metric even when the underlying verbatims are unchanged.
Fielding an in-product tracker and calling it a brand tracker. An in-product study typically measures whoever was in the product during the field window, which is not the category population the method requires.
The two reporting mistakes that survive review
Changing delivery mode or field window between waves. A wave fielded over two weeks and a wave fielded over six reach different people, and a wave that moved from email to panel has changed the instrument as surely as an edited question would.
Benchmarking against a published industry norm. Cross-industry norms are averages across brand sizes, on instruments with different question difficulty, and the ones in circulation typically cite nothing at all.
Two of these survive a competent review. Composition drift passes because the wave hit its quotas and the quotas were too coarse to catch it. Theme-count trending passes because the number is real, reproducible within a wave, and measuring a different thing each time.
Rather than treating these as a review checklist, treat the first six as design decisions and settle them before wave one fields.
Synthetic respondents in a brand tracker
Why a memory measure cannot be simulated
Synthetic respondents are not usable for a brand tracker, at any wave, for any metric.
The objection specific to this method is harder than the general one and it is worth leading with.
Awareness is fundamentally a memory measure. A synthetic respondent has no memory of your category to fail to retrieve, so an unaided recall question put to a language model is not measuring the thing the metric is defined as.
Aided recognition generally fares no better. A model's recognition of a brand is drawn from its training corpus, which is generally a measure of your brand's presence on the internet and not its presence in your buyers' heads. Those two things typically diverge most for exactly the brands a tracker is most often commissioned to watch.
The general evidence on synthetic respondents
The general evidence points the same way. Bisbee et al. (2024), in Political Analysis, compared 7,530 real respondents against 3.6 million synthetic responses and found that averages corresponded closely while 48 percent of regression coefficients differed significantly and 32 percent of those flipped sign, with artificially deflated variance that breaks power calculations outright.
Deflated variance is the part that matters most here, because every table in the sample size chapter above depends on variance being real.
One narrower use survives. Use synthetic responses for instrument piloting, to catch a broken skip pattern or an ambiguous question. Do not let a synthetic number into the series.
How Sprig supports a brand tracking survey
Sprig supports the design, fielding and collection half of this method and does not document the analysis half.
Rather than presenting that as a workflow choice, this guide has stated it as a constraint in the design chapter, because it changes where the wave-over-wave comparison happens.
What Sprig covers well
What the platform covers well is the part that has to be identical every wave.
One study definition delivers through in-product surveys on web and mobile, email with a custom sending domain, shareable links, QR codes and external research panels, with more than 300 targeting attributes available to stamp segment variables onto each response at collection.
The documented gaps
The gaps are documented absences and not inferences, and they are worth stating before anyone builds a tracker on the assumption they are absent.
- No documented trend view and no documented cross-study comparison returning numbers
- No documented significance testing in reports, dashboards, or the Synthesize Agent
- Response-based quotas capped at three screener questions, with no interlocking cells documented
- In-product audience sampling documented as a delivery throttle and not a random sample
- An account-wide recontact waiting period that a continuously fielded tracker will spend entirely
- CSV export only, with a download link that expires after a week and no warehouse connector documented
- No published weighting approach, sample size guidance, or statistical power method
The weighting absence matters most. A tracker that cannot weight to a category population cannot claim representativeness, and that is the claim a tracker is usually bought to make.
Third-party ratings, for a reader assessing the platform and not the method: G2 rates Sprig 4.3 out of 5 across 199 reviews on the product listing, retrieved August 17, 2026, and TrustRadius 8.5 out of 10 across 10 reviews, retrieved August 13, 2026. Those counts are 199 and 10, which is not a comparison of equals. Capterra and Gartner Peer Insights were checked and not found.
Alternatives and adjacent methods
Choosing a different instrument
Use a brand tracker when the question is whether a brand metric changed. Use one of the instruments below when it is not.
| Question | Method | Why it fits better |
|:-----------------------------------:|:-------------------------------------------:|:-------------------------------------------------------:|
| Did our campaign cause the change? | Geo-holdout or matched-market test | Supplies the counterfactual a tracker has none of |
| Which message moved people? | Monadic message testing | Contains a manipulation, which a tracker does not |
| Which claims matter most? | MaxDiff | Measures relative preference among individual items |
| What will people pay? | Gabor-Granger or Van Westendorp | Produces a price response curve, not points |
| Why do non-buyers not buy? | Depth interviews, moderated or AI-moderated | A qualitative instrument answering a different question |
| How do we compare on features? | Competitive analysis survey | Attribute comparison, not memory measurement |
| Is in-product experience improving? | Experience measurement over time | Measures the product, not the brand in the category |
Depth interviews belong in a different category from survey platforms. They generate depth from a small number of participants where a tracker generates incidence from many, and a team that needs both should generally run both rather than choosing.
Rather than treating the tracker as the whole measurement program, treat it as the standing measure other instruments are commissioned against.
A tracker that never triggers another study is generally reporting nothing worth acting on.
Brand perception studies are the closest adjacent instrument and are commonly confused with brand tracking. A perception study typically asks what people associate with a brand at one point in time. A tracker asks whether a frozen measure moved between two points, and the freezing is the whole method.
Frequently asked questions
What is a brand tracking survey?
A brand tracking survey is a repeated cross-sectional study that fields a frozen questionnaire to a fresh sample from the same population at a fixed interval, producing a time series of brand metrics such as awareness, consideration and mental availability.
Each wave is an independent sample rather than a repeat interview, which makes wave-over-wave comparison a two-sample problem and not a paired one.
How often should you run a brand tracking survey?
Set cadence from category purchase velocity and from your annual sample budget rather than from a calendar convention. Quarterly is the common default and it is typically a budget artifact, not a methodological finding. Two waves a year at 2,000 completes detect a 4.4-point change, while four at 1,000 detect only a 6.2-point change for the same spend.
What sample size do you need for a brand tracking survey?
Size the wave against the change you would act on, not against a single-wave margin of error. At a 50 percent aided awareness baseline, 1,000 completes per wave detect a true six-point change at 80 percent power, and a three-point change needs about 4,400. No professional body publishes a defensible minimum, so this guide publishes the derivation instead.
What is the difference between aided and unaided brand awareness?
Aided awareness, also called prompted recognition, asks whether a respondent recognises a brand from a shown list. Unaided awareness, also called spontaneous recall, asks which brands come to mind in a category with no list shown. Laurent, Kapferer and Roussel (1995) showed these are one latent trait measured at three levels of difficulty, not separate constructs.
What is a good brand awareness score?
There is no publicly retrievable, sourced cross-industry benchmark for brand awareness, so there is no external score to be good against. Every norm table available cites no study, no sample and no database. Your benchmark is your own prior waves on the same instrument, plus the competitor set measured in the same wave under identical conditions.
Does brand awareness predict sales?
Srinivasan, Vanhuele and Pauwels (2010) found that advertising awareness, brand consideration and brand liking account for almost one-third of explained sales variance across more than 60 brands of four consumer goods, with wear-in times letting the metrics work as advance warning signals.
That defends the metrics as leading indicators inside a sales model, not the funnel as a model of behaviour.
How many competitors should you include in a brand tracker?
Include your brand, the category leader, the three to five brands a buyer would consider, and one growth challenger.
A longer list is not free, because aided recognition rises mechanically with list length for small brands.
Should you run continuous or wave-based brand tracking?
Run discrete waves unless your annual sample budget can give a rolling window at least as large as the wave it would replace.
Continuous fielding timestamps movements against dated events, but overlapping rolling windows are not independent, so a two-sample test between adjacent windows is invalid however large they are.
Why is my brand tracker not moving?
A tracker that is not moving is frequently reporting correctly, because brand metrics are slow and most fielded wave sizes cannot detect small changes.
Check whether the movement exceeded the critical difference for your base size and baseline before treating flatness as a finding, and check whether the competitor set moved.
Is the marketing funnel dead?
The funnel is a poor model of how buyers behave and a defensible set of leading indicators inside a sales model, and those statements are not in conflict.
Funnel-stage scores scale with brand size under double jeopardy, so cross-brand conversion comparisons mislead, while the individual metrics still predict inside a market-response model.
Can AI respondents replace a brand tracker?
Synthetic respondents are not usable for a brand tracker, at any wave, for any metric. Awareness is a memory measure, and a language model has no memory of your category to fail to retrieve. Bisbee et al. (2024) found 48 percent of regression coefficients differed significantly against real respondents, with deflated variance that breaks power calculations.
How long should a brand tracking survey be?
Length is a budget decision more than a design one, because at a fixed budget every added question competes with sample size, and sample size decides whether the tracker can detect anything. No professional body publishes a length rule. The defensible test is whether a question's movement would change a decision, applied before wave one freezes.
What questions should a brand tracking survey include?
Order the instrument as a category buyership screener, unaided recall as open text, aided recognition against a randomized brand list, a category entry point grid, consideration, category and brand usage, a distinctive asset check, and demographics last.
Unprompted questions always precede prompted ones, because a shown brand list contaminates every recall measure after it.
Bottom line
A brand tracking survey is worth running when you can field a wave large enough to detect the change you would act on, and it is worth very little when you cannot.
That is the decision this guide exists to make explicit, and the tables above are the arithmetic behind it.
The design to build: prompted recognition as the core awareness metric, a category entry point battery in place of a generic attribute battery, and discrete waves with independent samples and fixed quotas, fielded to the category population and frozen at wave one.
If your annual budget supports two well-powered waves, run two. If it supports four underpowered ones, you will spend the same money measuring your own noise four times instead of your brand twice.
The next step is not a demo, so freeze the instrument, size wave one from the tables above, and field it. The brand awareness template is a working starting artifact, and the market and consumer insights page covers the wider program this sits inside.