Solutions
Experience measurement
Track sentiment and KPIs with AI-driven gap analysis
Strategic & foundational discovery
Uncover market whitespace with AI-led foundational studies
Journey & behavioral research
Connect user actions to motivations across the lifecycle
Market & consumer Insights
Understanding markets, audiences, & opportunity
Concept & prototype testing
Test designs and prototypes with rapid feedback
Agents
Design
Structure rigorous studies
Field
Run adaptive studies at scale
Analyze
Get statistical analysis you can trust
Synthesize
Turn results into research reports
Deploy
Email
Reach external audiences with native deliverability
Panels
Recruit from 300K+ verified participants
Web apps and websites
Embed studies in web experiences
Mobile apps
Run studies in iOS and Android apps
Customers
Community
Events
Join curated gatherings shaping the future of research
Blog
Insights on integrating AI into research craft
Book icon
Guides
Ultimate playbooks for enterprise survey research
Pricing
Sign in
Book a demo
Sign in
Book a demo
Guide

How to Run: Usage and Attitudes Study

October 2, 2026

By The Sprig Team

Example H2
Example H3
Example H4
Example H5
Example H6

Introduction

A usage and attitudes (U&A) study is a survey of a whole category that measures who uses it, how often, on which occasions, with which brands, and what those people believe about the brands they know. Running one well comes down to four choices: the category frame, bounded usage questions, a sample sized from the smallest subgroup you will report, and reading attitudes against each brand's user base.

Run a U&A on the category population rather than on your own customers.

A study of your customers is a customer study, and it excludes the non-users and competitor-only users who make up the growth opportunity.

Measure usage with a bounded reference period and the most recent occasion rather than with a typical-frequency scale.

In digital media use, self-reports correlate with logged use at only r = 0.38 across 106 effect sizes, so treat stated frequency as a rank order and not a level.

Read every brand attitude against the size of that brand's user base. Attitudes largely follow usage, so a raw attribute ranking often measures which brand is biggest.

Size the sample from the smallest subgroup you intend to report rather than from the total.

Comparing two subgroups on a 50 percent metric and detecting a ten-point gap needs about 400 respondents in each.

This guide covers:

  • Choose the category frame, and add a customer boost only when you need customer depth
  • Bound every recall question and ask about the last occasion instead of a typical one
  • Keep the core module short and rotate the rest as randomly assigned modules
  • Size the study from the smallest subgroup you will report, using the derived tables
  • Read brand attitudes against each brand's user base instead of as a raw ranking

Platform capabilities and documentation cited here were verified against sprig.com and docs.sprig.com on September 24 and 25, 2026. No prices appear anywhere in this guide.

What a usage and attitudes study is

A U&A study is foundational category research. Ipsos defines it as research which aims to "understand a market" and identify growth opportunities by answering questions on "whom to target, with what and how."

The same source compresses it to one line that is worth keeping: "a U&A is about what people do and what people feel regarding a category." Those two halves, doing and feeling, are the two halves of the instrument.

The usage half covers who uses the category, how often, how much, when, where, with whom, and with which brands.

The attitude half covers awareness, brand associations, needs, and the reasons people give for using or not using.

How a U&A differs from its neighbors

A U&A study maps a category once, at the population level. A brand tracker watches a handful of those numbers move over time.

A needs assessment ranks unmet outcomes. Persona research discovers the segments.

Each of those three is its own study with its own instrument, and each has its own guide.

This guide owns the map and hands the other three off at the points where the U&A stops being the right tool.

| Study | Question it answers | Frame | Run | |:-------------------:|:-------------------------------------------------:|:---------------------:|:----------------------:| | Usage and attitudes | Who uses the category, how, and what they believe | Category population | Once, then refreshed | | Brand tracker | Did our key numbers move | Category population | Repeated waves | | Needs assessment | Which unmet outcomes matter most | Category or customers | Once per roadmap cycle | | Persona research | Which natural groups exist | Category or customers | Once, then validated |

A U&A is the widest of the four and the shallowest on any one question, which is its strength and its main risk.

When to run a U&A study

The clearest case is entering a category you do not yet understand.

A team launching into a new market, a new segment, or a new use case needs the map before it needs a tracker or a segmentation.

A U&A fits when the decision is strategic, not tactical.

Choosing which occasion to build for, which competitor to position against, or which non-user group to pursue are all questions the category map generally answers well.

It fits when nobody on the team can say, with a number attached, how many people use the category, how often, and with which brands.

It fits before segmentation. A U&A generally collects the variables a later segmentation will run on, so running it first means the segmentation starts from category data rather than from a customer list.

And it fits when an old map has gone stale.

A category that has changed shape since the last study, through a new entrant, a new channel, or a shift in how people use it, needs a fresh map, not a patched one.

Where a U&A does not fit is as a way to answer one narrow question cheaply.

If the decision turns on a single number, a focused study on that number will often answer it with less sample and a shorter instrument.

When not to run a U&A study

Seven situations, each naming the instrument that is actually right. The first two are the ones most likely to change what you do.

You need to know what people actually do. Self-report is a weak measure of behavior. Parry and colleagues, in a pre-registered meta-analysis in Nature Human Behaviour in 2021, found that self-reported media use correlated with logged use at r = 0.38 across 106 effect sizes, and that only 6.12 percent of studies had mean self-reports within 5 percent of the logged values. That evidence covers digital media use, which is the closest base for software categories and should not be stretched to every category. The better instrument is product analytics or logged event data for your own users, and consumer purchase panel data for the category.

You need to know whether an attitude drives usage. Attitude responses are strongly tied to past usage, which is covered in the critique section below, so the direction of any link cannot be read from the survey. A correlation between "trust" and usage inside a U&A cannot be read as a driver. The better instrument is an experiment: message testing for claims, a concept test for propositions, and conjoint for attribute trade-offs.

You need to know what people will buy next. Stated intent predicts behavior imperfectly. Sheeran's 2002 review of 422 studies found an average intention-behavior correlation of r+ = 0.53, about 28 percent of variance, and a later meta-analysis of experiments by Webb and Sheeran found that changing intentions produced only a small-to-medium change in behavior, d+ = .36. The better instruments are Gabor-Granger or conjoint for price and choice, and an in-market test for the real answer.

You need to know whether anything changed. A U&A is a snapshot, and two snapshots built on different instruments are not a trend. The better instrument is a brand tracker built on a frozen subset of the U&A's metrics.

You need to know which unmet need to prioritize. A U&A captures needs broadly and ranks them weakly. The better instrument is a customer needs assessment, which scores importance and satisfaction on customer-derived outcomes, or MaxDiff when the list is long.

You need to know whether natural segments exist. A U&A collects the variables but does not validate the clusters. Dolnicar and Leisch, in Marketing Letters in 2010, propose bootstrapping to test whether a cluster solution reproduces. The better instrument is a segmentation study with a stability check, covered in the persona research guide.

You need to know why, in depth. Closed questions measure incidence, not reasoning, and a follow-up probe is still self-report. The better instrument is depth interviews, moderated or AI-moderated. Those are qualitative tools that answer a different question from fewer people, and they sit outside the survey category.

Hold the first two limitations hardest when the study is presented. They are the two that stakeholders most want the U&A to answer, because a usage number and a driver of usage are the easiest things to put on a slide.

State both limitations on the readout's first page. That protects the study later, when someone quotes a usage level out of context.

The pattern across all seven is the same.

A U&A is the right tool for structure and the wrong tool for levels, causes, forecasts and trends, and the sections below are built around that split.

Designing the study, and the three forks

Every U&A takes three design decisions, and they interact. Listing question categories without settling these three is how a U&A ends up expensive and unusable at the same time.

The three forks are who you sample, how you measure usage, and how you read attitudes. The committed default is stated first, and the evidence follows.

Sample the category population, with a customer boost only if you need customer depth, and weight the boost back. Measure usage with a bounded reference period and the most recent occasion, never with a range-anchored frequency scale. Link logged behavior wherever you own it. Read every brand attitude against that brand's user base before comparing brands.

Fork A: who you sample

A U&A is defined by its frame. Rather than surveying the people you can reach most cheaply, survey the people who make up the category, including non-users and people who only use competitors.

A frame of current customers typically answers a different question.

It tells you how your customers use the category, which is useful, and it tells you nothing about the people you do not have yet.

The practical answer when you need both is a dual frame.

Field the category sample through a panel, add a boost of your own customers through email or in-product delivery, and weight the boost back to its true share of the category before reporting any category-level number.

Weighting is not optional in a dual frame. An unweighted dual-frame study frequently overstates your own brand, because your customers are over-represented by design.

Fork B: how you measure usage

Measure usage with a bounded period and the most recent occasion.

Rather than asking how often someone typically does something, ask whether they did it in the last seven days, and then ask about the last time.

The evidence against the typical-frequency question is old and consistent.

Schwarz and colleagues, in Public Opinion Quarterly in 1985, gave respondents television-viewing scales with different ranges and found that people shown the low-range scale reported less viewing.

The authors concluded that "subjects inferred the average amount of television watching from the response alternatives provided them and used it as a standard of comparison." The scale is not a neutral container for the answer. It is part of the question.

So ask frequency as an open number over a stated period, never on pre-banded ranges.

"How many times in the last 30 days" with a number box produces a different and more defensible answer than five bands running from never to daily.

Bound the period because recall leaks outside it.

Neter and Waksberg documented telescoping in household expenditure interviews in 1964, meaning events from before the reference period get reported inside it, and a later Bureau of Labor Statistics review notes they "found higher levels of telescoping for larger home repairs."

The last-occasion question follows from those two findings, not from a study of its own, and it is a practitioner judgment.

A specific recent event is generally easier to recall accurately than an average, and it gives you a clean record of where, when, with whom and with which brand.

Where you own behavioral data, link it. Rather than choosing between self-report and logs, collect both for your own customers and use the comparison to calibrate how far stated frequency drifts from actual frequency in your category.

For your own customers, attach known usage attributes to each response so the comparison is possible. Panel respondents cannot be linked this way, because they are not your users.

Fork C: how you read attitudes

Read every brand attitude against the level expected from that brand's user base.

Rather than ranking brands on raw attribute scores, compare each score with what a brand of that size would normally get, and report the deviation.

The reason is covered in the critique section: attitudes largely follow usage.

A brand used by 40 percent of the category will typically outscore a brand used by 10 percent on nearly every positive attribute, simply because more respondents know it and use it.

The worked example in the core-calculation section shows the method end to end.

The short version is that a raw ranking tells you which brand is biggest, and the deviation tells you which brand stands out for something.

The design tree

Work down the questions in order until a branch fires.

Do you need to know about people who do not use your product? Yes, field to the category population through a panel. No, and you only need customers, then you are running a customer study, and the title of the deliverable should say so.

Do you also need depth on your own customers? Yes, add a customer boost from your own list and weight it back to its category share. Weighting happens outside the platform, so budget analyst time for it.

Do you own logged usage data for your customers? Yes, pipe it onto the customer boost responses and compare stated frequency against logged frequency. No, report stated frequency as a rank order and say so in the readout.

Does the attitude battery need both piped brand lists and randomized brand order on the same question? Yes, you have to pick one on that question, because response piping cannot be combined with randomization. Keeping randomization and showing the full brand list is the safer default, and that is a judgment rather than a sourced rule.

Will the instrument run past about 20 minutes? Yes, split it into a core module everyone answers and rotating modules assigned at random. No, keep it as one instrument and hold the length.

Is the smallest subgroup you plan to report below about 10 percent of category users? The 10 percent line is a working convention. Yes, read the sample-size tables before anything else, because the total you need may be several times what you planned.

Writing the instrument

A U&A instrument has a fixed order, and the order does as much work as the wording. Unprompted questions come before prompted ones, behavior comes before attitude, and profiling comes last.

The instrument spec

This is the instrument in fielding order, ready to lift. Replace the category wording and the brand list, and keep the order.

  1. Screener. Category use within a bounded period: "In the last 30 days, have you [CATEGORY ACTIVITY]?" Keep past-12-month users as a separate lapsed group instead of screening them out.
  2. Most recent occasion. "Think about the last time you [CATEGORY ACTIVITY]." Then when it was, where, with whom, what for, and which brand or product was used.
  3. Frequency. "How many times in the last 30 days did you [CATEGORY ACTIVITY]?" as an open number, with a sanity ceiling set at the build stage.
  4. Repertoire. Which brands the respondent has used in the last 12 months, as a multi-select from a randomized list.
  5. Awareness. Unaided brand recall as open text first, then aided recognition against the randomized brand list.
  6. Brand associations. A pick-any grid of brands by statements, asked of everyone for every brand they recognize.
  7. Needs. A short list of what matters when choosing, kept under ten items.
  8. Barriers. Why the respondent does not use, or stopped using, specific brands.
  9. Profiling. Demographics, firmographics for business categories, and any variable you want to describe segments with later.

Four rules sit under the list:

  • Never use a pre-banded frequency scale for any usage question
  • Bound every recall question to a period stated in the question text
  • Randomize brand order in every list while keeping the list itself fixed
  • Keep profiling variables at the end to protect attention for the core blocks

The fourth rule is standard practice rather than a sourced finding, and the other three follow from the evidence in the design section.

Why awareness comes before associations

Unaided recall has to come before any question that names a brand. Once a respondent has seen your brand list, their unaided answer is often no longer unaided.

The awareness block has its own evidence base, covered in depth in the brand tracking guide, including why aided recognition is generally the more stable core measure.

The U&A borrows that ordering rule instead of rebuilding it.

Why associations are pick-any, not rated

A pick-any grid asks which brands fit each statement, and the respondent ticks any that apply.

Rather than rating each brand on each statement on a scale, which multiplies length, the grid collects associations for every brand in one pass.

Pick-any also makes the usage-conditional reading in the core-calculation section straightforward, because the output is a simple share of respondents linking each brand to each statement.

Keeping the needs list short

The needs block in a U&A is a scan, not a ranking.

Rather than building a full importance-and-satisfaction battery here, keep the list short and hand the prioritization question to a needs assessment, which is built for it.

A needs list with 30 items inside a U&A generally produces flat, uniform answers late in the survey, which is exactly the position effect the length section describes.

In Sprig: the Design Agent can draft the full instrument from a written brief, including logic and randomization. It will not know your category definition or your reference periods unless the brief states them, so check both before fielding. Response piping cannot be combined with randomization on the same question, so a brand battery has to choose between piping each respondent's own brands and randomizing the order. Loop and merge is not documented, so a brand-by-statement battery is built as a matrix or as a pick-any multi-select, not as a loop.

Question formats and length

Three format decisions shape the data more than any wording choice: how frequency is asked, how the brand grid is built, and how long the whole instrument runs.

Frequency as an open number

Ask frequency as a number over a stated period.

The Schwarz finding in the design section is the reason, and the practical rule is to show a number box with a stated period instead of bands.

Set a ceiling during the build, such as 60 occasions in 30 days for a daily-use category, and review answers above it rather than deleting them.

A small number of respondents will often type implausible values, and the ceiling makes them visible.

The brand grid

Build the association grid as a matrix of brands by statements with multi-select cells, or as one multi-select question per statement listing all brands.

Rather than asking every statement about every brand separately, keep it to one screen per statement where possible.

Keep the statement list to what the decision needs. Eight to twelve statements is a common working range, and that is a convention, not a sourced figure.

How long a U&A can run

Length is the binding constraint on a U&A, and it degrades the back half of the instrument first.

Galesic and Bosnjak, in Public Opinion Quarterly in 2009, manipulated stated length at 10, 20 and 30 minutes.

They found that "the longer the stated length, the fewer respondents started and completed the questionnaire," and that answers to later questions "were faster, shorter, and more uniform than answers to questions positioned near the beginning."

Revilla and Ochoa, in the International Journal of Market Research in 2017, asked web panelists directly and found a median ideal length of 10 minutes and a maximum of 20.

That was one panel in one country, Mexico in 2016, so treat it as a reference point, not a universal limit.

The split questionnaire

The fix for length is a split design rather than a longer survey.

Adigüzel and Wedel, in the Journal of Marketing Research in 2008, open by noting that "massive questionnaires are pervasive in marketing practice," and show that splitting a long questionnaire across respondents and imputing the missing blocks improves data quality by reducing burden.

The practical version is simpler than their optimal design. Put the screener, the occasion block, frequency, repertoire and awareness in a core module everyone answers.

Split the association grid, needs and barriers into modules, and assign each respondent one at random.

Every module then typically reports on a smaller base, which feeds directly into the sample-size tables below.

A module seen by a third of respondents needs a total three times as large to reach the same subgroup precision.

The length budget worksheet

Budget the instrument in minutes before writing a single question. Rather than discovering the length at the pilot, assign each block a time and hold the core under the ceiling.

| Block | Module | Planned minutes | Who answers | |:------------------------:|:--------:|:---------------:|:--------------------:| | Screener | Core | 1 | Everyone screened | | Most recent occasion | Core | 3 | Category users | | Frequency and repertoire | Core | 2 | Category users | | Awareness | Core | 2 | Category users | | Brand association grid | Module A | 5 | Random half of users | | Needs and barriers | Module B | 4 | Random half of users | | Profiling | Core | 2 | Everyone qualified |

In this example the core runs 10 minutes and each respondent adds one module, so no respondent sees more than 15. The planned minutes are illustrative, and your own pilot timings replace them.

The worksheet also tells you the base for every block. The association grid in Module A reaches half the category users, so every brand-level number inside it rests on half the sample.

Rotate module order as well as module assignment. Galesic and Bosnjak's position effect means whichever module comes last gets the thinnest answers, so rotating order spreads that cost across modules instead of concentrating it in one.

In Sprig: random assignment of respondents to question blocks is supported inside one study, so a split design fields as a single survey, not as parallel ones. Rating Scale, Multiple Choice, Matrix and Rank Order cover the formats above, and the open-number frequency item is a text-input question with validation. The limitation is analysis, not fielding: there is no documented imputation for the missing blocks, so each module is reported on its own base instead of being modeled back to the full sample.

Sample size and precision

Size a U&A from the smallest subgroup you intend to report, not from the total. That single rule changes the answer more than any other design decision in this guide.

No methodology source publishes a U&A sample-size rule. Published vendor and agency guidance gives either no figure or an unsourced range.

Kadence writes that "a typical usage and attitudes study will involve a sample of participants, usually between 100 and 500, depending on the size of the target market," and cites no basis for it.

At 500 total, a subgroup making up 10 percent of the sample is 50 people, and its margin of error on a single proportion is about plus or minus 13.9 points.

That is too wide to report as a finding, and a U&A exists to report subgroups.

The derived tables

The tables below are computed, not copied, and every figure was recomputed a second way with a closed-form sample-size formula.

They assume 95 percent two-sided confidence, 80 percent power for detection, simple random sampling, no design effect and no weighting.

Weighting widens every figure. A dual-frame study weighted back to category proportions will need more sample than these tables show, and the size of the increase depends on how uneven the weights are.

Table A. Margin of error for one subgroup, and the smallest true gap detectable between two equal subgroups

| n per subgroup | Margin at 50% | Margin at 20% | Detectable gap between two subgroups, 50% base | |:--------------:|:------------------------:|:-----------------:|:----------------------------------------------:| | 100 | plus or minus 9.8 points | plus or minus 7.8 | 19.8 points | | 150 | plus or minus 8.0 | plus or minus 6.4 | 16.2 | | 200 | plus or minus 6.9 | plus or minus 5.5 | 14.0 | | 300 | plus or minus 5.7 | plus or minus 4.5 | 11.4 | | 400 | plus or minus 4.9 | plus or minus 3.9 | 9.9 | | 600 | plus or minus 4.0 | plus or minus 3.2 | 8.1 |

Table B. Total completes needed so the smallest subgroup reaches a chosen n

| Smallest subgroup's share | To reach 100 | To reach 200 | To reach 300 | |:-------------------------:|:------------:|:------------:|:------------:| | 50% | 200 | 400 | 600 | | 25% | 400 | 800 | 1,200 | | 15% | 667 | 1,333 | 2,000 | | 10% | 1,000 | 2,000 | 3,000 | | 5% | 2,000 | 4,000 | 6,000 |

Table C. People screened per 1,000 qualified completes, by category incidence

| Incidence | 80% | 50% | 30% | 15% | 5% | |:---------:|:-----:|:-----:|:-----:|:-----:|:------:| | Screened | 1,250 | 2,000 | 3,333 | 6,667 | 20,000 |

The formulas

The margin of error is 1.96 times the square root of p times one minus p divided by n, with the whole fraction under the root.

The detectable gap between two equal subgroups is the sum of 1.96 and 0.84, times the square root of two times p times one minus p divided by n, again with the whole fraction under the root.

Table B is the subgroup n divided by the subgroup's share of the sample. Table C is the target completes divided by incidence.

What the tables mean in practice

To compare two subgroups on a 50 percent metric and detect a ten-point gap, you need about 400 in each.

If the smaller subgroup is 10 percent of category users, that is about 4,000 completes.

A split module seen by half the sample doubles that again for any question inside it.

The combined effect is why a U&A generally costs more than teams expect, and why the 100-to-500 range in circulation cannot support the subgroup reads a U&A is commissioned for.

The brand tracking guide's margin table agrees with Table A at the shared points, including plus or minus 6.9 at 200, and its plus or minus 3.1 at 1,000 follows from the same formula, so the two guides can be read together.

A worked sizing example

A team wants to compare lapsed users with current users on a barrier question, and expects the gap between them to be about ten points on a metric near 50 percent.

Table A says that needs about 400 respondents in each group. Lapsed users are expected to be 15 percent of category users, so the Table B formula puts the total at roughly 2,667 category users to reach 400 lapsed ones.

The barrier question sits in a split module shown to half of users. That doubles the requirement to about 5,333 qualified category users.

Category incidence is 30 percent, so Table C puts the number screened at about 17,800. That is several times what the team planned, and it is the arithmetic rather than the vendor that makes it so.

The honest options are to move the barrier question into the core, accept a wider gap as the smallest one the study can detect, or drop the lapsed comparison. Fielding 1,000 completes and reporting the lapsed group anyway is not one of them.

No minimum floor

This guide publishes no minimum sample. Rather than borrowing a floor, pick the n in Table A that matches the smallest gap your decision needs to see, then read the total off Table B.

Audience, frames and screening

Screen on category behavior, not on brand. The screener decides who is a category user, and that definition becomes the denominator for every number in the study.

Write the definition down before building anything. "Used any project-management software for work in the last 30 days" and "is responsible for choosing project-management software" produce two different categories with two different sizes.

Business categories need one more decision: whether the respondent is the user, the buyer or both. A U&A of project-management software fielded to people who use it daily and to people who approve the budget will produce two different category pictures.

Screen for the role the decision cares about, and record the other role as a profiling variable rather than mixing both into one denominator.

Keep lapsed users instead of screening them out.

People who used the category in the last 12 months but not the last 30 are often the easiest non-users to win back, and they generally answer the barriers block better than anyone else.

Keep the screener behavioral rather than attitudinal. Asking whether someone is interested in the category selects for people who will answer favorably, which commonly inflates every attitude number that follows.

Set quotas on the variables that would distort category estimates if they drifted, typically age band, gender and region for consumer categories, and company size and role for business ones.

Record incidence as the study fields. The share of screened respondents who qualify is itself a category estimate, and it is often the most reliable number in the study because it rests on the largest base.

In Sprig: one study can field to a panel for the category sample and to your own list by email or in-product for the customer boost, which is what makes the dual frame practical. Response quotas are available on every plan, using single-select multiple choice or rating scale screener questions. The limitations: the two frames still have to be weighted together outside the platform, Panels documents no country coverage or fielding turnaround, and a change of delivery mode between frames is itself a measurement difference to note in the readout.

Fielding

Field the whole study in one window. A U&A fielded over several months commonly mixes seasons and events into the category picture, and nothing in the analysis can separate them afterwards.

Field the category sample and the customer boost at the same time. Rather than running the customer boost first because it is cheaper, start both together so the two frames describe the same moment.

Pilot with a small soft launch before the full field. Twenty to fifty completes, a practitioner convention, will show you broken logic, confusing screener wording and a completion time that runs longer than planned.

Watch the completion time against your length target. If the median runs past the planned length, cut or split modules before the full launch rather than accepting a degraded back half.

Hold the brand list and the statement list fixed once fielding starts. An edit mid-field typically splits the data into two instruments, and the two halves cannot be pooled.

The pre-launch checklist

Work through this list in order before the full launch. Most items cannot be fixed once responses start arriving.

  1. Write the category definition and the user definition in one sentence each, and confirm the screener measures exactly those.
  2. Confirm every recall question states its reference period in the question text.
  3. Confirm the frequency item is an open number with a ceiling, not a set of bands.
  4. Confirm unaided recall sits before any question that names a brand.
  5. Confirm brand order randomizes in every list while the list itself stays fixed.
  6. Confirm no question uses response piping and randomization together.
  7. Stamp a frame identifier on every response if the study runs a dual frame.
  8. Confirm module assignment is random and recoverable from the export.
  9. Read the smallest planned subgroup against Table A, and raise the total if the margin is too wide for the decision.
  10. Soft launch, record the median completion time, and cut or split any block that pushes the core past its budget.

Items 1, 7 and 9 are the ones teams most often skip. A missing definition, an unstamped frame or an undersized subgroup commonly surfaces only at analysis, when the fix means fielding again.

In Sprig: AI Follow-Ups (Field Agent) can add a short "why" probe after the barrier and last-occasion questions, which gets closer to reasoning than a closed list does. A follow-up is still self-report and not a depth interview, and follow-up text adds to the theming load instead of replacing it.

Timing and cadence

Run a U&A when a strategic decision is open and the category map is missing or stale.

No source publishes a repeat cadence for a U&A, and this guide does not invent one.

Refresh the map when something structural changes, such as a new entrant, a new channel, a regulation, or a visible shift in how the category is used.

Watch seasonality when choosing the field window. A category with strong seasonal peaks, such as travel, tax software or fitness, should be fielded in a window that represents the period the decision cares about, and the window should be stated in the report.

Keep the field window as short as the sample allows, and state its dates in the report. A window of two or three weeks is a common working choice, and that is a practitioner convention rather than a sourced rule.

If a brand tracker will follow the U&A, launch the tracker's first wave close to the U&A field window. The U&A then becomes the tracker's baseline, and the two studies describe the same moment.

When you do refresh, keep the category definition, the frame and the core questions identical.

A refreshed U&A that changes any of the three is a new study, and comparisons with the old one are not valid.

Quality control

Three checks matter most, and the first is specific to U&A studies.

Implausible frequencies. Review every frequency answer above the ceiling you set at the build stage. Rather than deleting outliers automatically, check whether they cluster, because a cluster can mean a confusing question, not a careless respondent.

Straight-lining in the grid. Flag respondents who tick every brand for every statement, or none for any. In a pick-any grid, ticking everything or nothing is rare enough to be worth reviewing respondent by respondent.

Speeders. Take a third of the pilot median completion time as a working floor, and review before deleting anything. That fraction is a convention, so check what it actually excludes before applying it.

Consistency checks specific to a U&A

Occasion and frequency that disagree. Cross-check the last-occasion date against stated frequency. A respondent whose last occasion was three weeks ago but who reports 20 occasions in the last 30 days has given two answers that cannot both be true.

Flag those respondents instead of guessing which answer is right. A high rate of disagreement usually points to a question problem, and it is one of the few internal consistency checks a U&A offers for free.

Frame drift. Compare the demographic mix of each frame against its targets as fielding runs. A panel sample that skews young mid-field will distort penetration, and correcting it with heavy weights later widens every margin in the study.

Report exclusions with the results. State the starting count, the number removed, the rule used, and the final count by frame, because unequal exclusions across frames change the weighting.

In Sprig: bot detection covers the automated end of quality control, which matters most for panel and link fielding. It does not catch a human respondent answering carelessly, so the three checks above still apply.

The core calculation

A U&A produces four numbers that no other study on this list produces together: weighted penetration, the frequency distribution, the share of volume from heavy users, and an estimate of total category occasions.

A fifth, usage-conditional brand association, is how the attitude half gets read.

Weighted penetration

Penetration is the share of the category population that used the category in the reference period.

It is the count of qualified category users divided by the count of everyone screened, weighted to the population.

Penetration rests on the whole screened sample rather than on the qualified users, which is why it is generally the most precise number in the study.

At 2,000 screened respondents and 30 percent penetration, the margin of error is about plus or minus 2.0 points.

The frequency distribution

Report frequency as a distribution, not as a mean.

Rather than a single average, show how many users fall into each band of stated occasions, because the shape tells you more than the center does.

Category frequency distributions are typically skewed, with many light users and a small number of heavy ones.

A mean pulled up by a few heavy users will often overstate how often a typical user engages.

Share of volume from heavy users

Sort users by stated frequency, take the heaviest 20 percent, and compute their share of all stated occasions.

That single number often corrects the most common strategic mistake a U&A gets used to justify, which is focusing only on heavy users.

Sharp, Romaniuk and Graham, in a 2019 working paper titled "Marketing's 60/20 Pareto Law," report that "a brand's heaviest 20% of buyers generally contribute not much more than half of a brand's sales, and these same buyers will contribute less in the following time period."

That paper uses purchase panel data, not self-report, and it is a working paper, not a peer-reviewed article.

A U&A's stated frequencies will not reproduce its figures cleanly, but the direction is the useful part: if the pattern carries over to category occasions, which the paper does not test, light users carry a large share of volume and a strategy aimed only at heavy users misses it.

Total category occasions

Category occasions equal the population size, times penetration, times mean occasions per user in the period. Every input should be printed next to the output, because each one carries its own uncertainty.

A worked example with illustrative numbers shows how that uncertainty compounds.

Take a population of 10 million adults, weighted penetration of 30 percent from 2,000 screened respondents, and a mean of 6.0 stated occasions per user per 30 days.

The point estimate is 10 million times 0.30 times 6.0, or 18 million occasions per 30 days. The penetration margin alone, 28 to 32 percent, moves that estimate from 16.8 million to 19.2 million.

Then add the self-report problem. If stated frequency in your category runs 20 percent above actual frequency, actual mean occasions are 5.0 and the estimate falls to 15.0 million.

That 20 percent is an assumption for the example, not a sourced figure, and the linked-data comparison in the design section is how you would measure it for your own customers.

The honest way to report a category size from a U&A is as a range with its inputs shown. A single number with no range often invites a precision the method does not have.

Usage-conditional brand association

Read each brand's association scores against the level expected for a brand with that many users.

Fit a line across all brands, relating each statement's association share to each brand's usage share, and report how far each brand sits above or below the line.

Here is a worked example with four brands and the statement "good value."

| Brand | Used by | Observed "good value" | Expected from usage | Deviation | |:-------:|:-------:|:---------------------:|:-------------------:|:---------:| | Brand A | 40% | 30% | 30.5% | minus 0.5 | | Brand B | 25% | 22% | 19.8% | plus 2.2 | | Brand C | 15% | 9% | 12.6% | minus 3.6 | | Brand D | 10% | 11% | 9.1% | plus 1.9 |

The raw ranking puts Brand A first on value by a wide margin, and that ranking mostly reflects Brand A's size.

The deviation column tells a different story: Brand C sits 3.6 points below what a brand of its size would normally get on value. That is the candidate finding, and at a base of 300 it is still inside the plus or minus 4.5-point margin in Table A, so it needs a larger base before anyone acts on it.

The expected values come from an ordinary least-squares line with a slope of 0.714 and an intercept of 1.93, recomputed from the closed-form slope formula with the same result.

Check the deviations against the subgroup margins in Table A before acting on them, because a deviation of two or three points on a small base can be noise.

The hand-off for everything else is clean. Cutting any of these numbers by segment belongs to cross-tab analysis, and coding open-text barriers into themes belongs to the theming workflow, covered in the simple and advanced cross-tab guides, with the U&A-specific steps in the Claude chapter below.

Benchmarks and what a good result looks like

There is no published cross-category benchmark for a U&A study's core outputs.

Penetration, usage frequency and brand attitude scores all depend on how your study defines the category and the user, so a number from someone else's study is not comparable to yours.

That absence is structural, not a gap waiting to be filled. A benchmark needs a stable definition, a defined population and a disclosed sample, and U&A studies share none of the three across organizations.

Two studies of "streaming video" can define a user as anyone who streamed in the last week or anyone who pays for a subscription.

The first will typically show higher penetration and lower frequency than the second, and neither is wrong.

The one regularity that transfers

Heavy users generally matter less than the 80/20 rule suggests. The 60/20 working paper cited in the core-calculation section is the evidence, and it reports brand-level purchase data, not category self-report.

That regularity is a sanity check, not a benchmark.

If your U&A shows the heaviest 20 percent carrying 90 percent of stated occasions, the likelier explanation is generally a frequency question that heavy users over-answered, not an unusually concentrated category.

Brand attitudes have a benchmark of sorts

The expected line in the usage-conditional reading is itself a within-study benchmark.

Every brand is compared with what a brand of its size gets in the same category, on the same statements, in the same field window.

That comparison transfers across brands because it is measured under identical conditions. It is the closest thing a U&A has to a norm, and it is generally more useful than any external table.

Build your own

Your benchmark is your own previous U&A on the same category definition, the same frame and the same questions.

Rather than comparing against another company's published figures, archive your study with the instrument version recorded so the next refresh has something valid to compare against.

This guide does not publish a minimum number of prior studies before a comparison is meaningful. No source supports one.

Interpreting and acting on the result

Work through the output in a fixed order, because the order is what stops a U&A from being read as proof of whatever the team already believed.

Start with the category definition and incidence. Confirm the definition matches the decision, and read incidence first, because every other number is a share of it.

Then penetration and frequency together. A category with high penetration and low frequency needs a different strategy from one with low penetration and high frequency. The first is generally a habit problem and the second a reach problem.

Then the occasion map. The last-occasion block shows where, when and with whom the category gets used. Occasions where your brand under-indexes against its overall share are typically the most concrete growth targets a U&A produces.

Then repertoire. Most category users use more than one brand, a pattern Ehrenberg's duplication-of-purchase work describes across many categories. The overlap between your users and each competitor's users shows who you actually compete with, which is often different from who the team assumes.

Then attitudes, against the expected line. Read deviations, not raw scores. A brand that over-indexes on a statement often owns something, and a brand that under-indexes has a specific weakness.

Then barriers and needs. These explain the patterns above rather than standing alone. Pair each barrier with the group that gave it, because a barrier cited by lapsed users means something different from one cited by never-users.

Here is how one finding travels through that order. A team sees that its brand under-indexes on weekday evening occasions, the repertoire block shows those occasions go mostly to one competitor, and the barriers block shows lapsed users citing price for exactly those occasions.

That chain supports a specific hypothesis worth testing, which is a much stronger output than any single number in the study. The test itself belongs to a concept test or message test, not to the U&A.

What the result licenses

A U&A licenses decisions about where to play: which occasions, which groups, which competitors, which positioning territory.

A U&A does not license a demand forecast, a claim that one attitude causes usage, or a statement that a number has moved. Those questions belong to the instruments named in the when-not-to-run section.

And a U&A does not license a segmentation on its own. It collects the variables, and a stability-tested segmentation study turns them into segments.

In Sprig: AI open-text theming (Synthesize Agent) codes the barrier and last-occasion verbatims into themes with response-level traceability, and AI Study Reports summarize the study once it holds enough responses. Themes regenerate and no reproducibility claim is published, so theme counts are not comparable between two U&A studies, and a refreshed map should be re-themed against the original theme list and not regenerated. Researchers remain responsible for the final theme list.

Running the analysis with Claude or ChatGPT

The U&A calculations are typically simple arithmetic applied to a large, messy file, which is exactly where a model helps and exactly where it fails quietly.

Weighting and the sizing arithmetic happen outside the platform either way, because Sprig documents neither.

This chapter covers only what is specific to a U&A.

Theme creation belongs to the advanced cross-tab guide, cutting any number by segment belongs to the simple cross-tab guide, and connector setup belongs to the documentation.

What the analysis produces

The analysis produces five outputs in one pass: weighted penetration with its margin of error, the frequency distribution, the heavy-user share of volume, a category occasion estimate as a range, and usage-conditional brand associations with deviations.

Each output is saved as its own named file. Rather than one long summary, separate files make each number checkable on its own.

Prompt one: weight, size and describe the category

You are analyzing a usage and attitudes study exported as CSV.

WHAT THIS PROMPT DOES: weights the sample to population targets, then computes
penetration, the frequency distribution, the heavy-user share of volume and a
category occasion estimate.
WHAT IT RETURNS: a weights file, a results table and a short method note.

Files: [PATH TO EXPORT FILE OR FILES]
Frame column: [FRAME ATTRIBUTE NAME, default Attributes_1]
Population targets: [PASTE TARGET SHARES BY WEIGHTING VARIABLE]
Maximum weight: [WEIGHT CAP, default 5, a common convention]
Category user question: [SCREENER QUESTION COLUMN]
Frequency question: [FREQUENCY QUESTION COLUMN]
Reference period in days: [PERIOD, default 30]
Population size: [POPULATION, no default, I must supply it]
Frequency ceiling: [CEILING, default 60]

Use code to calculate this, not estimation. Do not compute any value by
reasoning about it in prose.

Rules:
1. Exclude a respondent from any calculation where the relevant answer is
   blank, null or a non-response code. Blank is not zero. Zero occasions is a
   real answer and counts as a value.
2. Compute raking weights to the population targets. Cap weights at the
   maximum and report how many were capped and the design effect.
3. Penetration is weighted qualified users divided by weighted screened
   respondents. Report it with a 95 percent margin of error that accounts for
   the design effect.
4. Report the frequency distribution in bands and the weighted mean. Flag
   every answer above the ceiling and report results with and without them.
5. Sort users by stated frequency and report the share of stated occasions
   from the heaviest 20 percent.
6. Category occasions equal population times penetration times mean
   occasions. Report the point estimate and the range implied by the
   penetration margin, and print every input next to the output.
7. Any subgroup with fewer than [MINIMUM BASE FROM TABLE A, default 100] respondents,
   including zero, is reported as below threshold and not interpreted.

Verify the penetration and the heavy-user share a second way by bootstrapping
[ITERATIONS, default 2000] resamples of respondent rows with their weights, and
compare the bootstrap intervals with the analytic ones. These are different
methods and should broadly agree.

If any result disagrees between the two methods by more than [TOLERANCE,
default 2 points], say that you cannot reconcile them and show me the inputs.
Do not correct a mismatch by choosing the answer that looks more reasonable.

Save as ua_weights.csv, ua_category_results.csv and ua_method_note.md.

Prompt two: read brand attitudes against usage

WHAT THIS PROMPT DOES: reads each brand's association scores against the level
expected from its usage share.
WHAT IT RETURNS: one table per statement with observed, expected and deviation
for every brand.

Files: ua_weights.csv and [PATH TO EXPORT FILE]
Brand usage question: [REPERTOIRE QUESTION COLUMN]
Association grid columns: [LIST THE GRID COLUMNS]
Brands: [LIST OF BRANDS]
Minimum base per brand: [MINIMUM BASE FROM TABLE A, default 100]

Use code to calculate this, not estimation.

Rules:
1. A respondent counts toward a brand's association base only if they
   recognized that brand. Blank grid cells for recognized brands mean not
   associated. Blank cells for unrecognized brands are excluded.
2. Compute weighted usage share and weighted association share per brand per
   statement.
3. For each statement, fit an ordinary least-squares line of association share
   on usage share across brands. Report slope, intercept, expected value and
   deviation per brand.
4. Any brand with a base under the minimum, including zero, is reported as
   below threshold and left out of the fitted line.
5. Report the 95 percent margin on each observed share next to its deviation,
   and mark any deviation smaller than its margin as not distinguishable from
   the line.

Recompute every slope with the closed-form formula, covariance of usage and
association divided by the variance of usage, and compare it with the fitted
slope. If any pair differs by more than [TOLERANCE, default 0.01], say so and
show the inputs rather than choosing one.

Do not describe any brand as strong or weak on a statement from the raw share.
Describe it only from the deviation.

Save as ua_brand_associations.csv.

Setup and getting your data in

Export the responses as CSV, or pull them through the MCP connector. The documented export columns include surveyId, visitorId, userId, createdAt, completedAt, Q#_Question_Text, Q#_Response, Themes and Attributes_#.

There is no frame column by default. In a dual-frame design, stamp a frame identifier as an attribute at collection, panel or customer, so the weighting step can tell the two apart.

A split design generally leaves blanks in the modules a respondent was not assigned.

The prompts treat those as excluded, not as zero, which is correct, but confirm the module assignment is recoverable from the populated columns or from an attribute before running anything.

The MCP connector returns at most 1,000 responses per call, so a U&A of several thousand completes needs several calls or a CSV export.

Put the data above the instructions in the prompt, and split very large files by frame or module.

Reading the output

Read the method note before the results. It states the design effect, the number of capped weights and the exclusions, and a large design effect means every margin in the results is wider than a simple-random calculation would suggest.

Read the category occasion estimate as its range, not its point. The prompt prints every input next to the estimate, and in the worked example the penetration margin alone moves it by about 7 percent in each direction.

Read the brand table from the deviation column. A brand marked not distinguishable from the line has no finding on that statement, however large its raw share looks.

What to verify before reporting

Check five things before any number leaves the analysis:

  1. Confirm the weighted sample matches the population targets on every weighting variable.
  2. Confirm the bootstrap and analytic intervals agree.
  3. Confirm no subgroup below the minimum base is interpreted.
  4. Confirm the frequency ceiling did not remove a meaningful share of users.
  5. Recompute one brand's expected value by hand from the reported slope and intercept.

If the hand calculation disagrees with the model, stop and rerun rather than reporting either number.

Researchers remain responsible for the weighting targets, the category definition and every number that leaves the analysis. The model does the arithmetic, and it does not decide what the arithmetic means.

Pitfalls

Six documented failure modes apply to any survey analysis run in a language model, and one more applies to U&A files specifically.

Discovering themes is harder than applying them. Hill and colleagues, in PLOS Digital Health in April 2026, found models matched human analysts on coding against an existing codebook, 93.5 percent against 92.7, and were materially worse at inductive discovery. Theme inductively once, rebuild the list yourself, then have the model apply your list.

Rare themes get over-predicted. Ashwin, Chhabra and Rao, in Sociological Methods and Research in May 2025, found non-random bias in 10 of 19 codes, with sparse codes systematically over-predicted. The barriers block is full of rare themes.

Outputs vary between runs. Thinking Machines Lab reported in September 2025 that 1,000 completions at temperature zero produced 80 unique outputs. Neither chat client exposes a seed, so produce any reportable number twice.

The middle of a long input gets lost. Liu and colleagues, in the Transactions of the Association for Computational Linguistics in 2024, documented it. Put data above instructions and split large files.

Models tend to agree with the user. Sharma and colleagues documented sycophancy at ICLR in 2024, and OpenAI withdrew a model update in April 2025 for being overly agreeable. Never ask a model to confirm a finding you already suspect.

Verbatims contain personal information nobody asked for. Barrier and occasion text often names people, places and employers. Whether that data trains a model depends on the vendor's terms and your plan tier.

Asking a model to find the insights in a U&A file. A U&A file holds hundreds of possible cross-cuts, and a model asked what is interesting will find something every time, because it is not counting the comparisons. Specify the calculation rather than the question.

The critique you should know

Three objections apply to U&A studies specifically, and each one aims at a different half of the instrument. The defense is real and it is narrower than the method's advocates usually claim.

The usage half rests on self-report

The usage half of a U&A rests on the weakest measurement in survey research.

Parry and colleagues found self-reported media use correlated with logged use at r = 0.38, and that self-reports "were rarely an accurate reflection of logged media use."

Schwarz and colleagues showed that the response scale itself shifts reported frequency, because respondents read the scale as information about what is normal.

And telescoping pulls events from outside the reference period into it, which inflates counts for exactly the memorable purchases a U&A often cares about.

Put together, those findings mean the usage half of a U&A is typically better at ranking people than at counting what they do.

The attitude half is substantially a second usage measure

Attitudes largely follow usage. Bird, Channon and Ehrenberg, in the Journal of Marketing Research in 1970, found that "the proportion of people who express an attitude about a given brand generally depends on how recently they have used the brand."

Romaniuk, Bogomolova and Dall'Olmo Riley tested that generalization forty years later across 45 datasets, in the Journal of Advertising Research in 2012, and found brand association responses still strongly and systematically linked to past brand usage.

The consequence is uncomfortable for the way U&A results are usually presented.

A brand that leads on "trusted" or "innovative" often leads because more respondents use it, and presenting that lead as a perception advantage tells the team nothing it did not already know from market share.

The instrument's breadth degrades its own back half

Length is the third objection. Galesic and Bosnjak showed that later questions get faster, shorter and more uniform answers, and a U&A is usually longer than the single-topic surveys a team runs.

The questions most exposed are commonly the ones placed last: the attitude grid, needs and barriers. Those are frequently the sections the business commissioned the study to answer.

None of these objections says a U&A produces random output. They say its output means something narrower than the slide titles usually claim.

The fix is almost entirely in how the result is read, which is why the recommendations in this guide concern reading and design rather than whether to run the study at all.

The counter-critique

Three responses, and each one is supported, not asserted.

Self-reports carry real signal. A correlation of 0.38 is moderate, not zero. It supports sorting heavy users from light users far better than it supports estimating how many occasions each one has.

The conditions under which stated measures predict are known. Morwitz, Steckel and Gupta, in the International Journal of Forecasting in 2007, identify when purchase intentions predict sales better: for existing products over new ones, for durables, over shorter horizons, for specific brands over whole categories, for trial over total sales, and when collected comparatively rather than monadically. Those conditions apply to any intent questions a U&A carries, and a U&A that asks about specific brands over short horizons keeps its intent questions on the better side of that list.

The attitude objection has a design answer. Reading each brand against the level expected from its usage share, as the core-calculation section does, separates the part of an attitude score that size explains from the part it does not. That is this guide's judgment about how to respond to Bird and Romaniuk, not a sourced rebuttal of them.

Length is designable around. Adigüzel and Wedel showed that split questionnaires can deliver comparable data quality with lower respondent burden, which turns the length objection from a reason not to run the study into a reason to run it differently.

Where this guide lands

Use a U&A for structure and rank order, not for levels. Trust it for who uses the category, on which occasions, with which repertoire of brands, and in what rough proportions.

Distrust it for how much people use, for why one brand beats another, and for anything that sounds like a forecast.

Link logged behavior wherever you own it, read attitudes against usage, and split the instrument before it degrades.

Common mistakes

Running it on your own customers and calling it a U&A. A customer frame typically excludes non-users and competitor-only users. The result describes your customers accurately and the category not at all.

Asking frequency on banded scales. The bands become part of the answer. Ask for a number over a stated period instead.

Asking about a typical occasion. Typical is an average respondents construct on the spot. Ask about the last occasion, which is a real event with a date.

Ranking brands on raw attribute scores. The biggest brand often wins almost every positive statement. Read deviations from the expected line instead.

Sizing from the total instead of the smallest subgroup. A 1,000-complete study with a 10 percent subgroup reports that subgroup on 100 people, which is plus or minus 9.8 points on a 50 percent metric.

Forgetting that split modules shrink the base. A module seen by a third of respondents has a third of the sample. Size for the module, not for the study.

Reporting a dual-frame study unweighted. Your own customers are generally over-represented by design. Every category-level number is wrong until the boost is weighted back.

Reading a category size as a point. Penetration margin, frequency skew and self-report drift all widen the estimate. Report the range and the inputs.

Treating a U&A as a segmentation. The U&A collects the variables. A stability-tested segmentation study turns them into segments.

Changing the instrument at refresh. A refreshed U&A with a new category definition or new core questions is a new study. Freeze the core so the next map compares with this one.

Mistakes in the instrument and screener

Screening out lapsed users. People who used the category in the last year but not the last month are the cheapest non-users to win back. Keep them as a group and ask them the barriers block.

Asking profiling questions first. Demographics and firmographics asked early prime the answers that follow and spend attention on the least valuable questions. Put them last.

Letting the needs list grow. A thirty-item needs list inside a U&A produces flat, uniform answers late in the survey. Keep it short and hand prioritization to a needs assessment.

Asking unaided recall after naming brands. Once a brand list has appeared, unaided recall is contaminated for the rest of the survey. Order the awareness block before anything that names a brand.

Throwing away incidence. The share of screened people who qualify is a category estimate resting on the largest base in the study. Record it and report it.

Using the customer boost to describe the category. The boost exists to add depth on your own users. Quoting it as though it described everyone in the category repeats the customer-frame mistake with extra steps.

Skipping the soft launch. A U&A is long enough that broken logic or a confusing screener almost always hides somewhere. Twenty to fifty pilot completes find it before it costs the full sample.

Comparing theme counts across two studies. Regenerated themes are a different instrument each time. Apply the original theme list to the refresh instead of generating a new one.

Synthetic respondents

Synthetic respondents are not usable for the usage half of a U&A.

Usage is a behavior, and a language model has never performed it, so a synthetic answer to "how many times in the last 30 days" is not measuring the thing the question defines.

The general evidence points the same way. Bisbee and colleagues, in Political Analysis in 2024, compared 7,530 real respondents against 3.6 million synthetic responses and found averages corresponded closely.

But 48 percent of regression coefficients differed significantly from the human data, and 32 percent of those flipped sign, with artificially deflated variance.

Subgroup reads, usage-conditional attitudes and occasion-by-brand links are all relationships between variables, which is exactly what that study found synthetic data gets wrong.

They are not a source of category estimates, brand associations or subgroup reads, at any sample size.

Pre-testing is a legitimate and cheap use. Running a draft instrument past a model can surface ambiguous wording, a missing answer option, or a screener that admits the wrong people, before any real respondent sees it.

Treat those outputs as editorial notes on the questionnaire, not as data. Nothing a model says about how often it uses a category belongs in a category estimate.

How Sprig supports usage and attitudes research

The platform's strongest fit for a U&A is distribution.

One study can reach a panel for the category sample, your own customers by email, and your product users in-app, which is what makes the dual frame fieldable as a single instrument instead of separate projects.

The second fit is linkage. Attribute piping and URL-passed attributes attach known customer data to each response, which makes the self-report-versus-logged comparison possible for your own customers.

That comparison is the most direct answer this guide offers to the critique of self-reported usage.

The third is the fielding layer: random block assignment for split modules, response quotas on every plan, bot detection, AI Follow-Ups for "why" probes, and theming for the open text.

The Design Agent drafts the instrument, and the MCP connector gives an AI client governed access to the responses.

Attribution matters here. Theming, AI Study Reports and AI Follow-Ups are Sprig agents.

Weighting, significance tests, the sizing arithmetic and cross-tabs are run by the connected AI client, not by the platform.

What happens outside the platform

Weighting happens outside, because no weighting method is published. So do significance testing and the category sizing arithmetic, which is why the Claude chapter carries both.

Response piping cannot be combined with randomization on the same question, and loop and merge is not documented.

Panels documents no country coverage or turnaround, and panel participant figures differ between pages, so confirm reach for your category before planning a multi-country study.

Hosting is on AWS in the United States with no published regional options, which matters for any U&A covering respondents under data-residency rules. Export is CSV, with no warehouse connector documented.

The platform does not offer synthetic respondents and does not claim to replace a researcher's judgment on category definition, weighting targets or interpretation. Researchers remain responsible for all three.

Alternatives and adjacent methods

A brand tracker when the question is whether a known number moved. Freeze a subset of the U&A's awareness and association questions and field it in waves.

A customer needs assessment when the question is which unmet outcome to build for. It scores importance and satisfaction on customer-derived outcome statements, which a U&A needs block cannot do.

Persona research and segmentation when the question is which natural groups exist. The U&A supplies the variables, and the segmentation study tests whether the groups are real.

Conjoint or MaxDiff when the question is which attributes drive choice, or which of a long list of features matters most.

Message testing and concept testing when the question is causal. Both manipulate the stimulus, which a U&A never does.

Product analytics when the question is what your own users actually do. Logged behavior generally beats self-report on levels where it is available, though logs carry their own errors, such as one person using several devices.

Depth interviews when the question is why. They typically answer it from far fewer people and are a qualitative method, outside the survey category.

Frequently asked questions

What is a usage and attitudes study?

A usage and attitudes study is a category-wide survey measuring who uses a category, how often, on which occasions, with which brands, and what they believe about those brands.

It is typically run once as a strategic map rather than repeatedly as a tracker.

How do I run a usage and attitudes study?

Run a usage and attitudes (U&A) study in six steps. Define the category and the user in one sentence each, field to the category population with a customer boost only if needed, and ask usage with a bounded period and the last occasion. Keep each respondent under about 20 minutes with rotating modules, size the sample from the smallest subgroup you will report, and read brand attitudes against each brand's usage share.

What is the difference between a U&A study and brand tracking?

A usage and attitudes (U&A) study maps the whole category once, covering usage, occasions, repertoire and attitudes. A brand tracker repeats a small frozen subset of those measures in waves to see whether they move.

Run the U&A first, then build the tracker from its most decision-relevant metrics.

How many respondents does a U&A study need?

A usage and attitudes (U&A) study needs enough respondents for its smallest reported subgroup, not a fixed total. Comparing two subgroups on a 50 percent metric and detecting a ten-point gap needs about 400 in each.

If the smaller group is 10 percent of category users, that is about 4,000 completes before weighting.

How long should a U&A survey be?

A usage and attitudes (U&A) survey should keep each respondent's total under about 20 minutes, with the core module well inside that.

Web panelists in one study named 10 minutes as ideal and 20 as the maximum, and answers to later questions get faster and more uniform as length grows.

Split anything beyond the core into randomly assigned modules.

Should I survey my own customers or the whole category?

Survey the whole category for a U&A, because a customer sample excludes the non-users and competitor-only users who represent growth.

Add a customer boost when you need customer depth, and weight it back to its true category share before reporting category numbers.

How do I ask about usage frequency?

Ask usage frequency as an open number over a stated period, such as times in the last 30 days, never on banded ranges.

Banded scales shift the answers because respondents read the bands as a signal of what is normal.

Why do big brands score highest on every attribute?

In a usage and attitudes (U&A) study, big brands score highest on most attributes because attitude responses are strongly tied to usage, and more respondents use big brands. Read each brand's score against the level expected from its usage share, fitted across all brands in the study, and report the deviation, not the raw ranking.

Can a U&A study estimate market size?

A U&A study can estimate category occasions as population times penetration times mean frequency, reported as a range with every input shown.

Self-reported frequency typically drifts from actual behavior, so treat the estimate as directional unless you can calibrate it against logged data.

Can a U&A study replace segmentation?

A U&A study cannot replace segmentation, though it generally collects the variables a segmentation will use. Turning those variables into segments needs a clustering method with a stability check, which is its own study.

Is there a benchmark for U&A results?

There is no published cross-category benchmark for U&A results, because penetration, frequency and attitude scores all depend on how each study defines the category and the user.

Your benchmark is your own previous U&A on the same definition, frame and questions.

How often should a U&A study be repeated?

No source publishes a repeat cadence for U&A studies.

Refresh the map when the category changes structurally, through a new entrant, channel or usage pattern, and keep the definition and core questions identical so the two studies compare.

How long does a U&A study take from kickoff to readout?

A usage and attitudes (U&A) study's timeline is set mostly by instrument design, the soft launch, the field window and weighting. The field window alone commonly runs two to three weeks, a practitioner convention, and low-incidence categories take longer to fill. Budget separate time for weighting and analysis, because both happen outside the survey platform.

Where do weighting targets for a U&A study come from?

Weighting targets for a consumer usage and attitudes (U&A) study usually come from official population statistics, such as the US Census Bureau's American Community Survey for age, gender and region. Business categories can use government business counts by company size and industry. A customer boost is weighted back using your own customer count as a share of estimated category users.

What does a U&A study cost to run?

A U&A study's cost is driven mostly by incidence, total sample and length.

Low-incidence categories often need many more people screened per qualified complete, and subgroup reads multiply the total, so the sample-size tables are the best early guide to scale.

The bottom line

A U&A study is the broadest instrument in this guide series, and it is most useful when it is designed around what it cannot do.

It maps who uses a category, on which occasions and with which brands, and it maps that structure well.

It measures usage levels poorly, reads attitudes that largely echo usage, and loses quality in its back half as length grows. Each of those has a design answer.

Field to the category rather than your customers. Bound every recall question and ask about the last occasion.

Link logged data wherever you own it. Read attitudes against the expected line.

Split the instrument before it runs long. And size the study from the smallest subgroup you will report.

Then hand the map to the studies built for the next question, a tracker for movement, a needs assessment for priorities, and a segmentation for groups, rather than asking the U&A to answer them itself.

Back to top
Solutions
Experience measurementStrategic & foundational discoveryJourney & behavioral researchMarket & consumer insightsConcept & prototype testing
Agents
DesignFieldAnalyzeSynthesize
Deploy
EmailPanelsWeb apps and websitesMobile app
Pricing
Community
EventsBlogGuides
CustomersIntegrationsCompare
Company
About usCareersService agreementPrivacy policyData addendumSystem status
Socials
LinkedInX