Solutions
Experience measurement
Track sentiment and KPIs with AI-driven gap analysis
Strategic & foundational discovery
Uncover market whitespace with AI-led foundational studies
Journey & behavioral research
Connect user actions to motivations across the lifecycle
Market & consumer Insights
Understanding markets, audiences, & opportunity
Concept & prototype testing
Test designs and prototypes with rapid feedback
Agents
Design
Structure rigorous studies
Field
Run adaptive studies at scale
Synthesize
Turn results into research reports
Deploy
Email
Reach external audiences with native deliverability
Panels
Recruit from 300K+ verified participants
Web apps and websites
Embed studies in web experiences
Mobile apps
Run studies in iOS and Android apps
Customers
Community
Events
Join curated gatherings shaping the future of research
Blog
Insights on integrating AI into research craft
Book icon
Guides
Ultimate playbooks for enterprise survey research
Pricing
Sign in
Book a demo
Sign in
Book a demo
Guide

How to Run a Voice of Customer (VoC) Survey

September 15, 2026

By James Villacci

Example H2
Example H3
Example H4
Example H5
Example H6

The short answer

Running a voice of customer survey takes four decisions in order. Decide which of the two definitions your question belongs to, choose the instrument that matches it, field it to a population you can name and count, and analyze who did not answer before reporting what the people who did answer said.

Most programs skip the last step. That omission is why a voice of customer number is often a precise measurement of the wrong population.

This guide covers:

  • Choose between the 1993 research method and the 2026 software category before writing a question
  • Size the qualitative leg by saturation and the quantitative leg by the precision you actually need
  • Record the response rate as a reported number rather than an operational detail
  • Analyze the composition of the people who did not answer before reporting any satisfaction figure
  • Build the action layer outside the survey tool, because almost no survey tool publishes one

Platform capabilities and documentation cited here were verified against sprig.com and docs.sprig.com on September 11, 2026. No prices appear anywhere in this guide.

What a voice of customer survey is

Voice of customer means two documented and incompatible things.

One is a product-development research method published in 1993. The other is a software category defined by industry analysts in 2026. A reader searching this term typically does not know which one they are about to get, and most pages answering the query silently pick one without saying so.

The 1993 research method

The founding reference is Griffin, A. and Hauser, J. R. (1993), "The Voice of the Customer," Marketing Science 12(1), 1 to 27, DOI 10.1287/mksc.12.1.1. In that paper voice of customer is a needs-elicitation technique for product development, qualitative first and quantitative second, producing a structured and prioritized hierarchy of customer needs.

The method is built on one-on-one interviews, not on a questionnaire.

From the Wiley Encyclopedia entry by Gaskin, Griffin, Hauser, Katz and Klein, verbatim: "In a typical study between 10 and 30 customers are interviewed for approximately one-hour in a one-on-one setting."

The output is a needs hierarchy, not a score. Nothing in the founding method produces a tracked metric or involves fielding a survey to a customer list.

The 2026 software category

The industry definition is a market category, not a method. Gartner's market definition for voice of the customer platforms, retrieved September 11, 2026, describes products that "collect, analyze, and act on customer feedback across multiple channels."

That definition spans four capability areas: direct feedback such as surveys and polls, indirect feedback such as reviews and recorded interactions, AI-assisted analysis of the text, and an action layer routing findings to owners and tracking resolution.

The action layer is the part most teams generally underestimate. It is one of the mandatory capability categories in the market definition, so a program without one is not a voice of customer program under the industry definition at all.

Which definition your question belongs to

Use the 1993 method to discover what customers need from a product that does not exist yet. Use the 2026 category to measure and act on what existing customers experience now.

The two are not interchangeable and their sample logic is generally opposite. The research method needs a small number of long conversations, and the operations category needs a defined population, a repeatable instrument and a denominator you can state.

Rather than choosing a definition by preference, choose it by what your decision requires. A roadmap prioritization question is typically a needs-elicitation question, and a retention or service-quality question is usually a measurement question.

When to run a voice of customer survey

A voice of customer survey is the right instrument when you need a repeatable, population-level read on how a defined group of customers experiences something you already ship.

It is often a measurement instrument, not a discovery instrument, and it typically earns its place when the same question will be asked again in three months against the same frame.

The four questions this method answers well

The method is well suited to questions with a stable denominator and a stable instrument:

  • Measure how satisfied a defined customer population is with a product or a service interaction
  • Track whether that satisfaction is moving across periods on an instrument that did not change
  • Compare experience across segments you can identify from data you already hold
  • Surface the recurring reasons behind a score through open text attached to the same response

The signal that you are ready

You are ready to field when you can name the population, count it, and state how a respondent will be selected from it.

Teams that cannot do all three are typically not running a measurement program yet. They are collecting comments, which is a legitimate activity with a different name.

One more readiness condition is commonly missed.

You need somewhere for a response to go. A program that generates flagged responses with no owner and no resolution state produces a backlog rather than an improvement, and the backlog is visible to the customers who wrote into it.

When not to run a voice of customer survey

A voice of customer survey cannot establish cause, describe the customers who did not answer, or predict behavior reliably. Six situations call for a different instrument, and the first two are not negotiable.

When you need to know whether a change caused the outcome

A voice of customer survey is observational and has no counterfactual. Satisfaction rising after a release is generally consistent with the release helping, with seasonality, with a composition shift in who responded, and with all three together.

The better instrument is an online controlled experiment.

Kohavi et al. (2013), KDD 2013, DOI 10.1145/2487575.2488217, reports that "Only one third of the ideas tested at Microsoft improved the metric(s) they were designed to improve," and, citing prior work, that "90 percent of large randomized experiments produced results that stood up to replication, as compared to only 20 percent of nonrandomized studies."

When you need to know what your silent customers think

This is the defining limitation of the instrument, and it is measured rather than asserted.

Park, Cha and Rhim (2018), WWW '18 Companion, DOI 10.1145/3184558.3186579, analyzed 173,886 English chat sessions and 5,641,172 speech units over one year at a live chat service.

Verbatim: "This survey, however, had been answered by only 16.2% of the chat customers." And: "the mean satisfaction score of raters is higher (79.7% positive or neutral) than the inferred satisfaction score for non-raters (45.5% positive or neutral)."

The better instrument is behavioral analytics on the non-responding population, paired with the non-response bias analysis below.

When you need to know what customers will pay or choose

Stated willingness to pay is typically inflated, and the size of the inflation is known.

Murphy, J. J. et al. (2005), Environmental and Resource Economics 30, 313 to 325, DOI 10.1007/s10640-004-3332-z, covered 28 studies and 83 observations and found "a median ratio of hypothetical to actual value of only 1.35," adding that "a choice-based elicitation mechanism is important in reducing bias."

The real inflation factor is roughly 1.35 times, not the folklore two to three. The better instruments are discrete choice or conjoint analysis, and for a single price point, Gabor-Granger.

When you need needs for something genuinely novel

Customers reason from experience they have, which frequently fails when the thing you ask about has no precedent. von Hippel, E. (1986), Management Science 32(7), 791 to 805, puts it plainly: "for very novel products or in product categories characterized by rapid change, such as 'high technology' products, most potential users will not have the real-world experience needed to problem solve and provide accurate data to inquiring market researchers."

The better instruments are lead-user analysis and contextual inquiry, which is where a structured customer needs analysis belongs.

When you need behavior rather than reported behavior

A survey measures what people report, and it also changes what they report.

Chandon, P., Morwitz, V. G. and Reinartz, W. J. (2005), Journal of Marketing 69(2), 1 to 14, DOI 10.1509/jmkg.69.2.1.60755, found the intention-behavior correlation substantially higher among surveyed consumers than among comparable consumers who were not surveyed.

The survey often manufactures part of the relationship you believe you are observing. The better instruments are behavioral analytics, cohort retention analysis and session replay.

When you need a customer's own reasoning

An incidence battery commonly measures what, not why. Open text often gets closer without arriving, because a text box collects whatever a respondent typed in the fifteen seconds they gave it.

The better instrument is a moderated or AI-moderated depth interview.

Those are qualitative instruments in a different category from survey platforms, answering a different question from a small number of participants. They do not produce population estimates and do not substitute for a measurement program.

Study design, and the three forks

Voice of customer study design typically turns on three decisions that interact, which is why they are rendered here as one decision rather than three lists.

Fork A. Closed items or open text

Closed items carry the measurement load and open text carries the diagnostic load. The reliability evidence is adjacent-domain and should be read as such.

Wijngaards, I., Burger, M. and van Exel, J. (2019), "The promise of open survey questions: The validation of text-based job satisfaction measures," PLOS ONE 14(12), DOI 10.1371/journal.pone.0226408, studied job satisfaction, not customer satisfaction, so the mechanism transfers and the magnitudes do not.

Semi-open questions outperformed fully open ones there, correlating with human coding at r = .774 against .508, with convergent validity at r = .510 against .457.

But both text-based measures were materially worse than closed measures on test-retest reliability, at r = .255 against r = .543.

The committed default follows from the reliability gap. Do not make an open-text-derived score your tracked metric, and do not let a theme count stand in for a measurement.

Fork B. Solicited survey or unsolicited feedback

Both sources are biased, in opposite directions, and both biases are measured. Solicited feedback skews positive through non-response, as Park, Cha and Rhim documented at a 16.2% response rate.

Unsolicited feedback skews through acquisition and underreporting.

Hu, N., Pavlou, P. A. and Zhang, J. (2017), "On Self-Selection Biases in Online Product Reviews," MIS Quarterly 41(2), 449 to 471, states it directly: "two self-selection biases, acquisition bias and underreporting bias, render the mean rating a biased estimator of product quality, and they result in the well-known J-shaped (positively skewed, asymmetric, bimodal) distribution."

Neither source is representative and neither is generally reportable as what customers think. Triangulate them, and state which population each source actually covers rather than merging them into one number.

Fork C. Relational cadence or transactional trigger

No academic body and no professional standard takes a position on this fork.

The only named-body recommendation retrievable is New York State's NYX program guidance, page metadata November 2024, which states that agencies should "Use a combination of relationship and transactional surveys."

That guidance cites nothing and gives no sizing or frequency rule. It is a government program's stated practice, not evidence, and this guide presents it as exactly that.

The committed default is to run both and report them separately.

A transactional read is conditioned on the customer having just had an interaction, so its numbers are not comparable to a relational baseline and the two should never be blended into one trendline.

How the three forks interact

The three forks are not independent, and treating them as three separate choices is the most common design error in the method.

Fork C determines your denominator, fork B determines whether you have one at all, and fork A determines whether the thing you report can be compared across periods.

A transactional trigger with an open-text-derived score is commonly the worst combination available.

The denominator changes every period because interaction volume changes, and the metric changes every period because the coding scheme regenerates, so nothing in the series is comparable to anything else in it.

Decide all three forks before writing a single question, and write the choices down, because a fork resolved after fielding has begun is typically resolved by whatever was easiest to configure.

A relational cadence with a frozen closed battery and semi-open probes is the combination that survives. The frame is stable, the instrument is stable, and the open text explains movement rather than producing it.

The voice of customer scope decision tree

This is the first of five original assets in this guide. Read it top to bottom and stop at the first branch that matches your question.

| Question you are trying to answer | Definition in play | Instrument | Sizing logic | |:-------------------------------------------------------------:|:------------------:|:----------------------------------------------------------:|:------------------------------------------------------:| | What do customers need from a product that does not exist yet | 1993 method | One-on-one depth interviews, coded into a needs hierarchy | Saturation, 9 to 30 interviews per homogeneous segment | | How satisfied is a defined population right now | 2026 category | Closed battery, relational cadence | Precision, from the margin of error table | | Did satisfaction move between two periods | 2026 category | Same frozen closed battery, same frame | Power, from the detectable-change table | | Why did a specific interaction go badly | 2026 category | Closed battery plus semi-open probe, transactional trigger | Precision on the triggered population only | | Which of several fixes should we build first | Neither on its own | Discrete choice or conjoint, not a satisfaction survey | Design-dependent | | Did our change cause the improvement | Neither | Online controlled experiment | Power on the primary metric |

Three platform constraints frequently sit on these branches.

Response-based quotas are capped at three screener questions and cannot be set on attributes you already hold, in-product audience sampling is a delivery throttle, not a documented probability sample, and a survey platform generally does not publish the closed-loop case management workflow that the dedicated voice of customer platforms are evaluated on.

Where the action layer lives

The action layer lives outside your survey tool, and you design it that way from the start rather than discovering it at launch.

Assigning a flagged response to an owner, tracking a resolution state and closing the loop with the customer are operations functions that typically run in a ticketing or customer relationship management system.

Sprig fields every leg of this method across in-product, email, link and panel delivery, and publishes no case management, no documented assignment of a response to an owner, and no resolution tracking.

Plan the loop in the tool that already owns resolution.

Writing the instrument

A voice of customer instrument has two parts with two different jobs. The structured battery produces the number you report and compare. The semi-open probe set produces the reasons that make the number actionable, and it is never scored.

The structured battery

The battery should be short, frozen and identical across periods. Four to six closed items is typically enough for a relational read, and every additional item generally costs completions where fatigue is already the binding constraint. No source publishes an item-count rule, so treat that range as judgment.

Write one overall satisfaction item, two or three driver items on dimensions you can actually change, and one intent item if a downstream team uses it.

Rather than adding a new item each period, add it once and accept that your trend starts over on the items you touched.

The rule that protects the trend is mechanical.

Once the first period is fielded, the wording, the scale length, the point labels and the item order are frozen, and any change to them is a new instrument with a new baseline.

The semi-open probe set

Use semi-open probes rather than a bare open text box.

The Wijngaards finding above is the reason: semi-open questions correlated with human coding at r = .774 against .508 for fully open ones, so a probe with a stem and a prompt outperforms a blank field.

A semi-open probe names the territory without supplying the answer. "What made you choose that rating?" is semi-open. "Tell us anything else" is not, and it produces the thin text that gives open-ended questions their reputation.

Two probes is typically the practical ceiling, again as judgment rather than as a sourced rule. One tied to the satisfaction rating and one tied to the lowest-rated driver covers most diagnostic needs without adding a second minute to the completion time.

The instrument spec you can lift

This is the second original asset. It is written to be used without modification, and the bracketed fields are the only things you change.

| Item | Type | Wording | Scale and labels | |:----:|:---------------:|:--------------------------------------------------------------------:|:---------------------------------------------------------------------:| | S1 | Closed, primary | Overall, how satisfied are you with [PRODUCT OR SERVICE]? | 5-point, Very dissatisfied to Very satisfied, all five points labeled | | D1 | Closed, driver | How satisfied are you with how easy [PRODUCT] is to use? | Same 5-point scale, same labels | | D2 | Closed, driver | How satisfied are you with the speed of [PRODUCT]? | Same 5-point scale, same labels | | D3 | Closed, driver | How satisfied are you with the support you receive for [PRODUCT]? | Same 5-point scale, same labels | | P1 | Semi-open probe | What made you choose that rating? | Open text, shown to every respondent | | P2 | Semi-open probe | What is the one thing we could change that would matter most to you? | Open text, shown to every respondent |

Two fielding rules travel with the spec.

Show P1 to everyone, not only to detractors, because a diagnostic shown only to unhappy respondents produces a diagnosis that is only about unhappiness. And keep an explicit "Prefer not to answer" option out of the closed items, so that a blank is unambiguous evidence of item non-response rather than a choice.

Wording rules that change the numbers

Three wording decisions typically move the measurement more than most teams expect, and all three are cheap to get right:

  • Ask about a named object rather than about the company in general
  • Put the time frame in the question stem rather than leaving it implied
  • Avoid double-barreled items that ask about two dimensions in one sentence

Each of these is a question-writing rule, not a voice of customer rule, which is why the general guidance on writing effective survey questions applies without modification here.

Sprig callout: AI Follow-Ups. AI Follow-Ups probe an open-text answer conversationally while the respondent is still in the study, which genuinely narrows the gap between what someone rated and why. The limitation is the one this whole guide is built around. AI Follow-Ups deepen a response from someone who already answered. They do nothing about the people who did not answer, who in the best-documented case were 83.8% of those invited, and a richer response from a biased sample is still a biased sample.

Scale and scoring choices

Choose a five-point or seven-point fully labeled scale, report the mean as the primary metric, and report the top-two-box proportion alongside it. The evidence for the mean over a derived loyalty index is stronger than most practitioners expect.

Choosing a scale length

Five points and seven points both work, and the choice generally matters less than the consistency. A five-point scale is often easier to label unambiguously and renders better on mobile, which matters when most in-product responses arrive on a phone.

Rather than switching scale length to chase precision, hold the scale and add sample. A longer scale does not fix a sample that is unrepresentative, and changing scale length between periods destroys the comparison outright.

Labeling every point

Label every scale point, not only the endpoints. Partially labeled scales often invite respondents to interpret the middle differently from each other, and that interpretation varies systematically by how much attention someone is paying.

Numeric-only middles are the commonly seen version of this mistake. A respondent choosing 3 on an unlabeled 1 to 5 scale may mean neutral, may mean adequate, or may mean they did not want to think about it.

Anchoring the scale to a time frame

Anchor the scale to a stated period rather than to a general impression.

"Over the last 30 days" and "during your most recent support conversation" produce different numbers from the same customer, and only one of them is comparable across periods.

Rather than leaving the window implied, put it in the stem and freeze it with the rest of the instrument. A window that drifts between periods is a second silent instrument change layered on top of any wording change.

The mean, the top-two-box, or both

Report both and lead with the mean.

Morgan, N. A. and Rego, L. L. (2006), "The Value of Different Customer Satisfaction and Loyalty Metrics in Predicting Business Performance," Marketing Science 25(5), 426 to 439, found verbatim that "average satisfaction scores have the greatest value in predicting future business performance and that Top 2 Box satisfaction scores also have good predictive value."

That is a top-tier journal result placing the plain satisfaction mean ahead of the derived loyalty indices that dominate practice. It is also the reason this guide recommends a satisfaction battery rather than a single-item loyalty score.

The top-two-box proportion earns its place for a different reason.

It is the form that feeds the precision arithmetic in the next section, because a proportion has a closed-form margin of error and a mean requires you to estimate a standard deviation first.

Sample size and precision

No standards body publishes a sample-size rule for voice of customer surveys, so this section gives you the derivation instead, in two legs, because the method has two legs with different sizing logic.

What was checked: ISO 10004:2018 contains no sizing, sampling or frequency rule, leaving frequency to the organization at clause 6.2.

ESOMAR returned no satisfaction-specific guideline and its catalogue could not be fully enumerated, so it is reported as not found, not as absent. The Advertising Research Foundation surfaced nothing.

The qualitative leg, where 20 to 30 belongs

The qualitative leg has real empirical guidance, and this is typically the one place the famous number legitimately applies.

Griffin and Hauser studied 30 potential customers of portable food-carrying devices and identified 230 needs, reporting that "interviewing 20 customers identifies over 90% of the needs" and that "a single one-on-one interview identified 33% of the 230 needs and two one-on-one interviews identified 51%." Hauser's teaching note adds that "twenty respondents should be sufficient in most product categories."

The finding has been independently replicated three decades later.

Hennink, M. and Kaiser, B. N. (2022), Social Science & Medicine 292, 114523, DOI 10.1016/j.socscimed.2021.114523, reviewed 23 articles and found that "Studies using empirical data reached saturation within a narrow range of interviews (9-17) or focus group discussions (4-8), particularly those with relatively homogenous study populations and narrowly defined objectives."

Why the 20 to 30 figure is misapplied

The 20 to 30 figure measures coverage of a need space through hour-long one-on-one interviews. It is not, and has never been, a precision figure for estimating a population proportion.

A voice of customer survey sized at 20 to 30 responses on that citation is sized on a misreading of its own source.

The number is correct for what it measures and wrong for what it is routinely used to justify, and this is generally the most common sizing error in the query space.

Saturation asks whether you have heard the range of things people say. Precision asks how tightly you have estimated a number.

The quantitative leg, precision on one read and between two

The table below is derived for this guide, and the arithmetic is shown so you can check it.

Every cell assumes 95% two-sided confidence, simple random sampling, no design effect and no weighting. Margin of error on a proportion is 1.96 times the square root of p times one minus p divided by n, and on a mean it is 1.96 times the standard deviation divided by the square root of n. The critical difference for two independent reads of equal size is 1.96 times the square root of the sum of the two variances, and it is the wider test you need whenever you compare two periods.

| Responses | Error at 50% | Error at 70% | Error on a mean | Critical diff. at 50% | Critical diff. at 70% | |:---------:|:------------:|:------------:|:---------------:|:---------------------:|:---------------------:| | 100 | 9.8 points | 9.0 points | 0.20 | 13.9 points | 12.7 points | | 200 | 6.9 points | 6.4 points | 0.14 | 9.8 points | 9.0 points | | 300 | 5.7 points | 5.2 points | 0.11 | 8.0 points | 7.3 points | | 400 | 4.9 points | 4.5 points | 0.10 | 6.9 points | 6.4 points | | 600 | 4.0 points | 3.7 points | 0.08 | 5.7 points | 5.2 points | | 1,000 | 3.1 points | 2.8 points | 0.06 | 4.4 points | 4.0 points | | 2,000 | 2.2 points | 2.0 points | 0.04 | 3.1 points | 2.8 points |

Every figure is plus or minus the value shown.

At 400 completed responses a top-two-box proportion near 70% is estimated to within roughly 4.5 points, and a mean on a five-point item to within a tenth of a scale point. The mean column assumes a standard deviation of 1.0, which you replace with your own after one period.

At 200 responses per period, a top-two-box figure moving from 68% to 74% has not moved in any statistical sense, because the critical difference at that base is roughly 9 points.

Reporting that six-point move as an improvement is a commonly published false positive.

The sample needed to detect a change

The table above tells you whether an observed difference is distinguishable from noise. This one tells you how much sample you need to detect a real difference of a given size, at 80% power and 95% two-sided confidence.

| True change to detect | Responses per period at a 50% baseline | Responses per period at a 70% baseline | |:---------------------:|:--------------------------------------:|:--------------------------------------:| | 3 points | 4,356 | 3,554 | | 5 points | 1,565 | 1,251 | | 7.5 points | 693 | 540 | | 10 points | 388 | 294 | | 15 points | 170 | 121 |

The practical reading is uncomfortable.

Detecting a three-point shift requires several thousand responses per period, and most voice of customer programs are typically sized an order of magnitude below that. A program collecting 200 responses a quarter can detect large changes only, and should report only large changes.

Response rate matters more than response count

For a voice of customer program the response rate matters more than the response count, because the bias Park, Cha and Rhim documented generally does not shrink as n grows.

A 10,000-response survey at a 16% response rate is a precise measurement of the wrong population.

Adding sample narrows the confidence interval around a biased estimate rather than moving the estimate toward the truth. A tighter interval around a wrong number often invites more confidence than a wide one, which makes it the worse outcome.

Rather than targeting a response count, target a response rate and a known denominator, then let the count be whatever the rate produces.

The federal standard nobody cites

The United States Office of Management and Budget publishes the relevant standard, in "Standards and Guidelines for Statistical Surveys," September 2006. Two guidelines apply directly and neither appears in any competing voice of customer page reviewed for this guide.

Guideline 3.2.9, verbatim: "Given a survey with an overall unit response rate of less than 80 percent, conduct an analysis of nonresponse bias using unit response rates as defined above, with an assessment of whether the data are missing completely at random."

Guideline 3.2.10 adds that "If the item response rate is less than 70 percent, conduct an item nonresponse analysis to determine if the data are missing at random at the item level."

Voice of customer surveys rarely clear 80%, and the best-documented rate in a peer-reviewed venue is 16.2%.

The federal standard would therefore require a non-response bias analysis on virtually every voice of customer program in existence, and almost none run one.

Audience, targeting and screening

Define the sampling frame before you define the audience filter. The frame is generally the population you are entitled to generalize to, and every targeting decision after that either preserves it or quietly replaces it with something narrower.

Defining the sampling frame

Write the frame down as a sentence with a count attached.

"All accounts with at least one active seat in the last 30 days, 41,200 accounts as of September 1" is a frame. "Engaged users" is not, because nobody can reproduce it next quarter.

Rather than targeting who is easiest to reach, target who your conclusion will be about.

An in-product study typically reaches people who were in the product during the fielding window, which is a frame that systematically excludes the customers most at risk of leaving.

Screening and quotas

Screening removes people who are out of frame, and quotas typically hold the composition of the people who are in it. The two are commonly confused, and the confusion produces a sample that is correctly filtered and badly composed.

Quotas on most survey platforms are response-based, which means they count people who have already answered rather than shaping who gets invited.

Response-based quotas cannot be set on attributes you already hold about a customer, which is the form a voice of customer program actually needs.

Stamping segment variables at collection

Stamp every segment variable you will want to analyze onto the response at collection. Recovering a segment after the fact requires a join, and an unplanned join is often where segment bias hides.

The variables worth stamping are the ones you can also compute for the whole frame.

A segment variable you hold only for respondents cannot be used to compare respondents against the population, which is exactly the comparison the non-response analysis needs.

Sprig callout: attribute targeting. Sprig publishes 300 or more targeting attributes spanning demographic, professional, behavioral and firmographic fields, and supports attribute targeting plus attribute piping and attributes passed in through a URL, so segment and journey-stage variables can be written onto each response at collection. The limitation is that attributes target, they do not quota. Targeting on an attribute filters who becomes eligible and does nothing to hold that attribute's distribution constant in the responses you get back.

Fielding and delivery

Field on the channels that reach your whole frame, not those easiest to instrument. A program running only in-product typically measures only the customers still showing up.

Choosing channels

Match the channel to the frame.

In-product delivery reaches active users, email delivery reaches the whole contact list including dormant accounts, link and QR code delivery reaches people at an offline touchpoint, and research panels reach people outside your customer base.

Most frames generally need at least two of these. A relational read of all customers fielded only through in-product surveys is a read of active users wearing the label of a customer read.

Invitation and reminder design

Send one invitation and one reminder, and keep both short. Reminders often raise response rates and also change who responds, so record the split between first-wave and reminder respondents and compare their scores before pooling.

A difference between the waves is direct evidence about the direction of your non-response bias, which makes it the cheapest bias diagnostic available.

Recording the denominator

Record the denominator at the moment of fielding, not afterward. You need invitations sent, invitations delivered, studies shown, studies started and studies completed, as five separate counts.

Without them you cannot compute a response rate, and without a response rate you cannot run the analysis the federal standard asks for. This is the most commonly repeated data-collection omission in the method.

Sprig callout: channel coverage. Sprig fields one instrument across in-product web and native mobile, native email with a custom sending domain and domain warming, shareable links, QR codes and external panels, which covers the digital surfaces a voice of customer frame usually spans. The limitation is that the coverage stops at digital. There is no native SMS delivery, though a link can be sent through your own SMS tool, and there is no contact-centre, telephony or offline capture, so the program covers part of the journey rather than the whole of it.

Drafting assumption A9, flagged for resolution. This guide states no participant count for any panel. Two figures are currently published on sprig.com, "2M+ verified participants" on the panels page and "300K+" on the voice of customer analytics blog post, and the second of those is the page this guide is intended to replace. Citing either would put a number in a guide that contradicts a page the guide redirects. The count is omitted entirely until the two surfaces are reconciled, and that reconciliation should happen before the redirect ships rather than before this guide publishes.

Timing and cadence

Run relational reads on a fixed calendar and transactional reads on an event trigger, and never merge the two into one series.

Relational cadence

Quarterly is the common relational cadence and it is a judgment, not a sourced rule. Choose one you can sustain with a frozen instrument for at least four periods, because a trend with two points is not a trend.

The trade-off is mechanical. A shorter cadence generally gives more points and fewer responses per point, which widens the critical difference and makes real changes harder to detect.

Transactional triggers

Trigger transactional reads on a completed interaction, not on a page view. The population is everyone who completed it in the window, and that is the denominator you record.

Transactional reads are conditioned on the interaction having happened, so they cannot be compared to a relational baseline.

The recontact budget

Over-surveying suppresses response, and the effect is documented.

Porter, S. R., Whitcomb, M. E. and Weitzer, W. H. (2004), New Directions for Institutional Research 121, 63 to 73, DOI 10.1002/ir.101, reports "response rates decreasing as the number of prior surveys increases from zero to one to two (68 percent, 58 percent, and 46 percent, respectively)."

That study is of university students, not customers, so the mechanism transfers and the magnitudes do not. Treat those percentages as evidence that fatigue is real and never as customer benchmarks.

A recontact waiting period is typically an account-wide default applying to everyone shown a study, so a continuously fielded transactional program spends the recontact budget of every other study in the account.

Sprig callout: Design Agent. The Design Agent builds a fully programmed study from an uploaded document with response options, logic and randomization already in place, which removes most of the setup cost from standing a program up. The limitation matters specifically for a recurring read. A voice of customer instrument must be frozen across periods, and regenerating the questionnaire between periods silently breaks the comparison even when every regenerated question looks better than the one it replaced.

Drafting assumption A1, flagged for resolution. Whether a single study can be run repeatedly across periods, or whether each period requires its own study, is not resolved in published documentation. This guide assumes one study per period, with the period identifier written into the study name and stamped as an attribute on every response. That design works under either answer and keeps period identification in the export. Confirm the mechanism before standardizing the naming convention.

Quality control and data hygiene

Run quality control in three passes: before launch, during fielding, and after the field closes. Each pass catches a different class of defect and none of them substitutes for the others.

Pre-launch checks

The pre-launch pass is a checklist, not a judgment, and it runs in launch order:

  • Confirm the instrument is byte-identical to the last period on every frozen item
  • Confirm the five denominator counts are being recorded and are queryable
  • Confirm the period identifier is stamped as an attribute on every response
  • Confirm every segment variable you plan to analyze exists for the whole frame
  • Confirm the open-text probes are shown to every respondent, not only to low scorers
  • Confirm bot and fraud filtering is enabled on any link or panel distribution

In-field checks

Watch completion and drop-off while the study is live rather than after it closes.

A drop-off spike at one item is nearly always a wording or rendering defect, and it is recoverable in the first day and unrecoverable in the last.

Watch the composition of responses as they arrive, not only the count. If one segment is overrepresented on day one, the fielding mechanism is selecting, not sampling, and you would rather know before the field closes.

Post-field cleaning

Clean on documented rules rather than on inspection.

Remove straightlined responses where every item carries the identical value and the open text is empty, remove responses completed below a duration floor you set in advance, and keep a count of everything removed.

Report the removed count alongside the analyzed count. A cleaning rule with no reported count is typically indistinguishable from one applied to produce a preferred result.

Analysis, the core calculation

Compute the response rate and the respondent-versus-population composition before computing any satisfaction figure. That ordering is the whole method, and reversing it is how programs frequently end up defending a number they cannot support.

Compute the response rate first

The unit response rate is completed responses over eligible invitations or impressions. The item response rate is answers to an item over the responses that reached it.

Compute both, report both, and check them against the federal thresholds. Below 80% unit response you owe a non-response bias analysis, and below 70% on any item you owe an item non-response analysis.

The non-response bias worksheet

This is the third original asset, built on the Office of Management and Budget standard, not on vendor practice, and it runs in five steps.

| Step | What you compute | What it tells you | |:----:|:--------------------------------------------------------------:|:--------------------------------------------------------------------:| | 1 | Unit response rate, completed over eligible | Whether the federal 80% threshold is cleared, which it will not be | | 2 | Respondent composition on each held attribute | The profile of who answered, on variables you hold for everyone | | 3 | Population composition on the same attributes | The profile of the frame, computed the same way | | 4 | The difference in each attribute, respondents minus population | The direction and size of the composition gap | | 5 | Satisfaction mean within each over- and under-represented cell | Whether the composition gap moves the headline number, and which way |

Step five converts a composition gap into a bias estimate.

If the over-represented segment scores above the overall mean, the headline number is biased upward, and the gap times the difference in cell means is a defensible first approximation of by how much.

Compare early responders against reminder responders as a sixth, cheaper check. Late responders resemble non-responders, so a downward drift between waves is corroborating evidence that the people you never reached score lower still.

The satisfaction mean and the top-two-box

Compute the satisfaction mean as the primary metric and the top-two-box proportion alongside it. Report both with the margin of error from the precision table and the completed response count in the same sentence.

Exclude blanks and non-responses from the denominator of each item rather than treating them as a middle value. Recoding a blank to a neutral score is the most common way an item non-response problem disappears into a satisfaction figure.

Theming the open text

Theme the open text to explain the number rather than to become one.

A theme count is typically a description of how the coder behaved, and it becomes a metric only if the coding scheme is frozen the same way the instrument is.

Automated theming generally does not meet that condition by default, because most implementations regenerate the theme list as new responses arrive.

A regenerated theme list is a different instrument, so a theme count from this period is not comparable to a theme count from the last one.

Sprig callout: open-text theming. Sprig's Theming agent groups open-text responses into themes with real-time regeneration and response-level traceability, so every theme can be opened back to the individual responses behind it, which is what makes automated coding auditable at all. The limitations are real and should be designed around. No minimum response count is documented, no accuracy or reproducibility figure is published, and the themes regenerate, so a regenerated theme list is not the same instrument as last period's and a theme count is not a trackable metric.

Benchmarks and what a good result looks like

There is no publicly retrievable cross-industry benchmark for voice of customer response rates or program outcomes that discloses its sample size and method.

The one large program benchmark that discloses its method is a 2017 self-assessment of 186 large companies, and the best-documented response rate available anywhere is a single company's, at 16.2% across 173,886 chat sessions, published in a peer-reviewed venue.

What was checked, and what came back

The search was specific, and the results are worth stating individually:

| Source | What it publishes | Usable as a benchmark | |:-----------------------------------------------------------:|:-------------------------------------------------------------------------:|:-----------------------------------------------:| | Qualtrics XM Institute, response rates | States that "there is no universal answer to what makes a good rate" | No, it declines to publish one | | Qualtrics XM Institute, State of Voice of the Customer 2017 | Program maturity self-assessment, 186 organizations above 500M in revenue | No response rate is disclosed | | Forrester, State of Feedback Management, 2025 | Published August 8, 2025, gated | No public sample size or method | | ESOMAR | No customer experience or satisfaction guideline surfaced | Reported as not found, catalogue not enumerable | | ISO 10004:2018 | Monitoring and measuring customer satisfaction | No sample-size, sampling or frequency rule | | CustomerGauge | Response-rate tables including a 12.4% figure | No sample size, no date, no method, unusable |

The 2017 self-assessment carries the most quoted finding in the category, that "less than one-quarter of companies consider themselves good at making changes to the business based on the insights."

It is self-assessed, nine years old, and published by an organization that sells the remedy.

Drafting assumption A4, flagged for resolution. No American Customer Satisfaction Index figures appear in this guide. The index is the one methodologically disclosed cross-industry satisfaction benchmark available, but its published terms state verbatim that "No advertising or other promotional use can be made of ACSI data and information without the express prior written consent of ACSI LLC." A lead-generating vendor guide is promotional use on any reasonable reading, so the figures are omitted pending counsel rather than cited with a caveat. The guide does not need them, because the honest benchmark answer here is the absence.

The benchmark that actually works

Your benchmark is your own prior period, on the same instrument, against the same sampling frame, with the response rate stated and a non-response analysis attached.

It is the only comparison where the questionnaire, the population and the fielding mechanism are all held constant.

No minimum response rate, no sample floor and no program-maturity threshold is published in this guide. No source supports one, and a threshold invented to fill the gap would be less useful than the gap.

Interpreting and acting on the result

Read a voice of customer result as three numbers, not one: the satisfaction figure, its precision, and the response rate that produced it. A figure reported without the other two is not interpretable.

Reading a change

Check any period-over-period movement against the critical difference table before calling it a change. At typical program sample sizes most quarterly movement sits inside the noise band and should be reported as flat.

A rising satisfaction mean with a falling response rate is the most common false positive in the method. Fewer people answering typically means the people still answering are more self-selected, and more self-selection generally means more positive.

Rather than reporting a movement alone, report it with its critical difference beside it. A number that names its own noise band is hard to misuse and easy to cite.

The closing-the-loop operating spec

This is the fourth original asset. It exists because no competing methodology page publishes one, and because survey platforms generally leave this layer to the buyer even though the dedicated voice of customer platforms treat it as core.

| Element | What you define | Default worth starting from | |:------------------------:|:---------------------------------------:|:----------------------------------------------------------------------------:| | Trigger | Which responses enter the loop | Bottom-two-box on the primary item, or any open text flagged by a named rule | | Owner | Who receives the response, by role | The account or service owner already responsible for that customer | | Response-time target | How quickly first contact happens | One business day for bottom-box, five for everything else | | Resolution states | The fixed list a case can be in | Open, contacted, resolved, resolved with product change, closed unresolved | | Escalation rule | What leaves the loop and goes to a team | Any case unresolved past the target, and any theme recurring across accounts | | Feedback to the customer | What the customer hears back, and when | An acknowledgement at contact and an outcome at resolution |

Run this in the system that already owns resolution, generally a ticketing or customer relationship management system.

A loop that runs in a spreadsheet often works for one quarter and fails in the second, because the spreadsheet has no owner field anyone is measured on.

The action layer is a category boundary

The gap between a survey platform and a dedicated voice of customer platform is the action layer, and it is a category boundary, not a feature gap.

The industry market definition makes acting on feedback one of three mandatory capability areas, alongside collection and analysis.

Sprig publishes none of it.

There is no documented case management, no documented assignment of a response to an owner and no documented resolution tracking, and the Slack integration notifies on study events rather than on response content. A team expecting to buy a survey platform and receive a closed loop will typically not get one, and that is true of most of the category, not of one product.

The practical consequence is a build decision, not a buy decision. Choose the platform on collection and analysis, then build the loop in the operational system you already run.

Running the analysis with Claude or ChatGPT

Connect your response export to a language model and it can compute the response rate, the composition comparison and the satisfaction read in one pass.

What it cannot do is decide the order, and the order is what makes the analysis defensible.

What the analysis produces

The output is generally four artifacts in sequence: a response rate table with the five denominator counts, a composition table comparing respondents to the population, a satisfaction table by segment with margins of error, and a theme summary labeled diagnostic, not metric.

The prompts below refuse to produce the satisfaction read until the first two are present.

Prompt one, response rate and composition

What this prompt does: computes the unit and item response rates for a voice of
customer study, then compares the profile of respondents against the profile of
the full population on every attribute held for both.
What it returns: a response rate table, a composition table, and a stated
direction of likely bias. It returns no satisfaction figure.

Data: [RESPONSE EXPORT CSV] and [POPULATION FRAME CSV], the frame holding one
row per person eligible to be surveyed with the same attribute columns.

Use code to calculate this, not estimation. Write and run the code, and show it.

1. Compute the unit response rate as completed responses over eligible
   invitations, using [INVITED COUNT COLUMN] and [COMPLETED COLUMN].
2. Compute the item response rate for every question column separately.
3. Exclusion rule: exclude blank cells, whitespace-only cells and the literal
   values "N/A", "null" and "Prefer not to answer" from every item numerator.
   Never recode a blank to a scale midpoint. Report the exclusions per item.
4. Respondent threshold: [MINIMUM CELL SIZE, default the larger of 30 and 2% of
   total completed responses]. Report any cell below it as "below threshold,
   n = X" rather than as a percentage. A cell with zero respondents is below
   the threshold and is reported as "n = 0", never omitted.
5. For every attribute in both files, compute the respondent share, the
   population share, and the difference in percentage points.
6. Independent recompute: derive the unit response rate a second way, as one
   minus the count of frame rows with no matching response over eligible frame
   rows. Do not reuse step 1. Report both figures.
7. If the two differ by more than 0.5 percentage points, stop and report the
   mismatch with both numbers and your best explanation. Do not reconcile them
   silently and do not pick one.
8. Output one markdown file named
   [STUDY NAME]-response-rate-and-composition-[PERIOD].md holding the response
   rate table, the item response table, the composition table, and one
   paragraph naming the over- and under-represented segments.

Prompt two, the satisfaction read by segment

What this prompt does: computes the satisfaction mean and the top-two-box
proportion for a voice of customer study, overall and by segment, with margins
of error.
What it returns: a satisfaction table by segment and a short interpretation
naming its own noise band. It runs only after prompt one has completed.

Data: [RESPONSE EXPORT CSV] joined to [CUSTOMER ATTRIBUTE TABLE] on [JOIN KEY,
default userId].

Use code to calculate this, not estimation. Write and run the code, and show it.

1. Confirm the output of prompt one is present. If it is not, stop and say so.
2. Join the export to the attribute table on [JOIN KEY]. Report how many
   response rows failed to join and what share of the total that is.
3. Exclusion rule: exclude blank, whitespace-only, "N/A", "null" and
   "Prefer not to answer" from every calculation. Never recode a blank to a
   midpoint. Report the excluded count per item.
4. Respondent threshold: [MINIMUM CELL SIZE, default the larger of 30 and 2% of
   total completed responses]. Segments below it are reported as "below
   threshold, n = X". A segment with zero respondents is reported as "n = 0"
   and is never dropped from the table.
5. Compute the mean of [PRIMARY SATISFACTION ITEM] overall and per segment,
   with a 95% margin of error as 1.96 times the standard deviation over the
   square root of n.
6. Compute the top-two-box proportion for the same item, with a 95% margin of
   error as 1.96 times the square root of p times one minus p over n.
7. Independent recompute: rebuild the mean a different way. Take the frequency
   distribution of each scale value, multiply each value by its count, sum, and
   divide by the total count. Do not reuse step 5. Report both means to three
   decimal places.
8. If the two means differ at the third decimal place, stop and report the
   mismatch with both values rather than choosing one.
9. Compare each segment against the overall figure using the critical
   difference rather than the single-sample margin of error, and say so.
10. Output one markdown file named
    [STUDY NAME]-satisfaction-by-segment-[PERIOD].md holding the segment table,
    the excluded counts, both means, and one paragraph naming the differences
    that exceed the critical difference.

Setup and getting your data in

Export your responses as CSV and connect the file directly, or connect the model to your platform through a published connector.

The step nobody documents is the join.

A voice of customer analysis is worth little without the customer attributes the survey did not ask for, so you concatenate the export with your own customer table on a stable identifier. That join is where segment bias becomes visible or stays hidden, which is why the failed-join count is part of the output.

Put the data above the instructions. Liu et al., TACL 2024, documented that models attend least well to the middle of a long context, so large exports are split into passes.

Reading the output

Read the composition table before the satisfaction table, every time. A figure from a sample over-representing your most engaged segment measures engagement as much as satisfaction.

Read every segment difference against the critical difference rather than the single-sample margin of error. Two overlapping single-sample intervals are not a test.

What to verify before reporting

Verify five things before any number leaves the analysis:

  • Confirm the response rate appears in the output and matches the independent recompute
  • Confirm the excluded blank counts are reported per item, not aggregated
  • Confirm every segment cell reports its n, including the cells reporting zero
  • Confirm both computed means agree to three decimal places
  • Confirm the theme summary is labeled diagnostic and carries no period comparison

Pitfalls

Eight failure modes are now documented well enough to design around, and the last two are specific to this method.

Discovering themes is harder than applying them.

Hill et al., PLOS Digital Health, April 3, 2026, DOI 10.1371/journal.pdig.0001189, found deductive agreement of 93.5% against blinded human coders at 92.7%, but kappa was 0.34 for both. Never quote the 93.5% without the kappa, because agreement looks high only because code prevalence is low at 7.8%. Strict hallucination ran at 1.2% and comprehensive error at 12.4%. Theme inductively once, rebuild the list yourself, then have the model apply your frozen list every period after.

Language models are non-deterministic. Thinking Machines Lab, September 2025, produced 80 unique outputs from 1,000 completions at temperature zero, and neither chat client exposes a seed, so any reportable number gets produced twice and compared.

Models are also sycophantic.

Sharma et al., ICLR 2024, documented the effect and OpenAI withdrew a model update in April 2025 for being overly agreeable, so never ask a model to confirm a theme you already suspect. Verbatims also carry personal information nobody asked for, and the governance question turns on which service tier you use.

Automated coding error is not random with respect to who the respondent is.

Ashwin, J., Chhabra, A. and Rao, V. (2025), Sociological Methods & Research, DOI 10.1177/00491241251338246, states verbatim that "the errors that LLMs make in coding interview transcripts are not random with respect to the characteristics of the interview subjects."

In voice of customer terms, automated theming can systematically under-report the concerns of specific segments. That is the same failure the method already has through non-response, arriving a second time through the analysis, and the two compound rather than cancel.

Your respondents may be using a language model too.

Zhang, S., Xu, J. and Alvero, AJ. (2025), Sociological Methods & Research, DOI 10.1177/00491241251327130, found that "34 percent reported using LLMs to help them answer open-ended survey questions" on research panels. That figure comes from a paid panel, where the incentive to automate is stronger than for a customer answering an in-product prompt, so the direction transfers and the magnitude does not.

Theme creation, theme counts and the recode loop belong to the guide on coding open-ended responses into themes, and comparing a metric across segments belongs to the guide on cross-tab analysis. Use those rather than rebuilding either.

Sprig callout: Sprig MCP. Sprig MCP connects study data to Claude, ChatGPT, Gemini and Copilot with published governance that is unusual in the category: access scoped to the authenticated user's role, response data not used to train models, agents unable to launch or modify a live study, and an organization-wide admin kill switch, all documented on the data security page. That governance matters more here than elsewhere, because voice of customer verbatims carry personal information nobody asked for. Two limitations apply. Access is capped at 1,000 responses per call, so a full program pull takes multiple calls, and the connection hands a model data to reason over rather than running a statistical test, which is why the prompts above specify the arithmetic instead of asking for a conclusion.

The critique you should know

The evidence for customer satisfaction as a construct is strong, and the evidence for voice of customer programs as they are run is weak. The gap between those two statements is where most criticism of this method lives.

The founding paper contains the critique

Griffin and Hauser documented the instrument's central flaw in the paper that introduced it.

Verbatim: "Our data demonstrate a self-selection bias in satisfaction measures that are used commonly for QFD and for corporate incentive programs." Thirty-three years later the instrument still carries that flaw.

Over-surveying suppresses response

Porter, Whitcomb and Weitzer (2004) measured response rates falling from 68% to 58% to 46% as prior survey exposure increased from zero to one to two, concluding that "Multiple surveys do appear to suppress response rates."

Those figures come from university students, not customers, so the mechanism transfers and the magnitudes are not customer benchmarks.

A program therefore competes for response with every other study the organization runs, and one that raises its cadence to get more data frequently gets less.

Measuring satisfaction changes the customer

Your voice of customer survey is an intervention, not an observation, and this is the finding almost nobody in the query space knows exists.

Dholakia, U. M. and Morwitz, V. G. (2002), Journal of Consumer Research 29(2), 159 to 167, DOI 10.1086/341568, reports verbatim that "measuring satisfaction (a) changes one-time purchase behavior, (b) changes relational customer behaviors (likelihood of defection, aggregate product use, and profitability), and (c) results in effects that increase for months afterward and persist even a year later."

The same paper states that "These results raise questions concerning the design, interpretation, and ethics in the conduct of applied marketing research studies." A program measuring a population repeatedly is changing its behavior.

Automated open-text analysis is not reliable unvalidated

van Atteveldt, W., van der Velden, M. A. C. G. and Boukes, M. (2021), Communication Methods and Measures 15(2), 121 to 140, DOI 10.1080/19312458.2020.1869198, concluded verbatim that "The best performance is still attained with trained human or crowd coding," that "None of the used dictionaries come close to acceptable levels of validity," and that "machine learning, especially deep learning, substantially outperforms dictionary-based methods but falls short of human performance."

Dictionary correlations with the gold standard ran between 0.12 and 0.33, so a lexicon-based sentiment score is close to noise, and reporting one as a trend is reporting noise as a trend.

Whether programs change anything is not established

No peer-reviewed empirical or experimental evidence was found that closed-loop voice of customer programs improve retention or revenue. Every result returned was vendor marketing, and this is reported as a verified absence, not a negative finding.

The nearest disclosed figure is the 2017 program self-assessment in the benchmark section, which found that "less than one-quarter of companies consider themselves good at making changes to the business based on the insights."

The counter-critique

The construct itself holds up under forty years of evidence.

Mittal, V. et al. (2023), Marketing Letters 34, 171 to 187, DOI 10.1007/s11002-023-09671-w, reports "a meta-analysis based on 535 correlations from 245 articles representing a combined sample size of 1,160,982," finding a positive association between customer satisfaction and both customer-level and firm-level outcomes.

The reported correlations are substantial: retention at r = .60, word of mouth at r = .68, price at r = .39, spending at r = .28, Tobin's q at r = .29 and return on assets at r = .22.

The second counter-critique is sharper.

Morgan and Rego (2006) found the plain satisfaction mean to be the best available predictor of future business performance, which means the instrument works and the derived indices layered on top of it are the weaker part.

Where this leaves the method

The construct is sound and the evidence for it is stronger than for any metric layered on top of it.

What is weak is the program: the sampling frame nobody writes down, the non-response analysis nobody runs, and the action layer that typically does not exist.

Measure satisfaction properly, analyze who did not answer, and build the loop before buying the dashboard. That ordering separates a program that survives scrutiny from one that produces a quarterly number nobody acts on.

Common mistakes

Six mistakes account for most of the failed programs, and five of them are design decisions made before a single response arrives.

Mistakes in design

  • Sizing a survey at 20 to 30 responses on the strength of a qualitative saturation finding
  • Changing the instrument between periods and continuing to report one trendline
  • Blending relational and transactional reads into a single series with one denominator
  • Treating an in-product sample as a customer sample when it only reaches active users

Each of these is generally cheap to prevent and expensive to fix, because the fix restarts the trend.

The sizing mistake and the instrument-drift mistake share a root cause, which is that nobody owns the instrument. A questionnaire with no named owner typically accumulates edits from everyone who reads it, and every edit is defensible alone and destructive in aggregate. Freezing the instrument and writing down the frame before the first period costs an afternoon.

Mistakes in analysis

  • Reporting a period-over-period movement without checking it against the critical difference
  • Recoding blank and skipped items to a scale midpoint instead of excluding them
  • Reporting a theme count as a metric when the theme list regenerates each period
  • Reporting a satisfaction figure without the response rate that produced it

The last one is the most consequential, because it is the omission that lets every other problem stay invisible. A satisfaction figure with no denominator cannot be challenged, which is exactly why it should be.

Recoding blanks to a midpoint deserves its own warning, because it usually happens inside a spreadsheet formula rather than as a decision anyone made.

A blank is evidence that someone chose not to answer, and that choice is typically correlated with the answer they would have given.

Mistakes in acting

The two failures on the action side are opposite and equally common.

Some programs route every response to an owner and create a backlog nobody works, and others route nothing and produce a quarterly report with no owner attached to any finding.

Both failures come from skipping the trigger definition. A trigger that names which responses enter the loop is generally what keeps the volume finite, and a program without one either drowns or does nothing.

A third acting mistake is quieter and often more damaging.

Teams close a case with the customer without recording whether anything changed in the product, so the loop closes on the conversation and stays open on the cause, and the same theme returns next period looking like a new finding.

Rather than routing by volume, route by a named trigger with a response-time target, and publish the count of cases closed alongside the satisfaction figure.

Reporting that pair is also the cheapest defence of the program's own budget, because a satisfaction figure alone reads as a vanity number and a resolution count reads as work delivered.

A program that reports both numbers is measurably harder to ignore than one that reports a score.

Synthetic respondents

Synthetic respondents are not usable for a voice of customer program, at any point, for any metric. The objection here is sharper than the general one and it is worth stating first.

The entire value of a voice of customer program is hearing something you did not already know, from people whose experience you have not modelled.

A synthetic respondent is constructed from what is already in the corpus, so it cannot report an outage you have not published, a checkout flow that broke last Tuesday, or a complaint nobody has written down yet.

The statistical objection is independently documented.

Bisbee, J. et al. (2024), Political Analysis 32(4), 401 to 416, DOI 10.1017/pan.2024.5, generated 3,614,400 synthetic responses from personas based on 7,530 human respondents and found verbatim that "The average scores generated by ChatGPT correspond closely to the averages in our baseline survey. Nevertheless, sampling by ChatGPT is not reliable for statistical inference."

In that study 48% of coefficients differed significantly from their survey counterparts, and among those the sign flipped 32% of the time. Variance collapse produced false precision, the failure mode most likely to survive review unnoticed.

How Sprig supports a voice of customer program

Sprig covers collection and analysis for this method and does not cover the action layer.

What the platform covers

One instrument fields across in-product web and native mobile, native email with a custom sending domain, shareable links, QR codes and external panels, covering the digital surfaces a voice of customer frame usually spans.

Open-text theming with response-level traceability, conversational follow-ups, 300 or more targeting attributes with attribute piping, and a Model Context Protocol connection with published governance cover the diagnostic and analysis legs above.

Third-party ratings, retrieved August 17, 2026: G2 rates the product 4.3 out of 5 across 199 reviews on its product page, and TrustRadius 8.5 out of 10 across 10 reviews.

Capterra and Gartner Peer Insights are both reported as not found, not as zero.

What the platform does not cover

The limitations below are documented absences, not inferences:

  • No published case management, owner assignment or resolution tracking for a response
  • No documented trend view, cross-study comparison returning numbers, or significance testing
  • No published weighting approach and no published sample-size or statistical-power method
  • Response-based quotas capped at three screener questions and not settable on held attributes
  • In-product audience sampling documented as a delivery throttle, not a random selection
  • No native SMS, contact-centre, telephony or offline capture channels
  • Export limited to CSV, capped at 950,000 rows, with a download link that expires after a week

The recontact waiting period is an account-wide default applying to everyone shown a study whether or not they responded. Panels are limited to Enterprise and to link surveys, with no documented minimum sample, turnaround or country coverage.

One further absence belongs in a guide with this title.

Sprig does not appear in the 2026 Gartner® Magic Quadrant™ for Voice of the Customer Platforms, document 7534085, published March 9, 2026, which evaluated twelve vendors: Alchemer, Concentrix, Medallia, Pisano, Press Ganey Forsta, Qualtrics, QuestionPro, Revuze, SMG, Sprinklr, Verint and XEBO.ai. A buying committee will typically check that list, and conceding costs less than omitting.

Gartner® and Magic Quadrant™ are trademarks of Gartner, Inc. and its affiliates.

Gartner does not endorse any vendor, product or service depicted in its research publications and does not advise technology users to select only those vendors with the highest ratings or other designation. Gartner research publications consist of the opinions of Gartner's research organization and should not be construed as statements of fact. Gartner disclaims all warranties with respect to this research.

Drafting assumption A4, second half, flagged for resolution. The Magic Quadrant reference above is included because the concession is the strongest credibility move available to a guide with this title, and the vendor list was verified on September 11, 2026. The exact attribution and disclaimer wording is pending review by counsel, alongside the question of whether American Customer Satisfaction Index figures may be used at all. Both should be answered before publication.

For teams evaluating the category rather than the method, the Medallia comparison covers the buyer-facing version of this decision, and the experience measurement and market and consumer insights pages describe the programs it sits inside.

Alternatives and adjacent methods

Several methods commonly sit next to a voice of customer survey and answer questions it cannot. Choose between them by what the decision requires rather than by what the program already collects.

Customer satisfaction measurement is the closed-item core of this method and has its own template.

Customer effort measurement asks a narrower question about a single interaction and has its own template as well. A needs-focused study belongs to the 1993 branch of the method and maps to meeting customer needs rather than to a satisfaction battery.

For ongoing rather than periodic collection, a continuous product feedback design keeps a channel open between relational reads, and measuring product value covers the perceived-worth question that satisfaction scores do not answer.

Three adjacent methods answer questions the survey cannot. An online controlled experiment establishes cause. Discrete choice, conjoint analysis and MaxDiff establish preference and trade-off. Behavioral analytics and sentiment measurement describe what people did, not what they reported.

Brand tracking sits alongside a voice of customer program rather than inside it.

A brand tracker measures perception among a market population that mostly does not use your product, while a voice of customer program measures experience among customers who do, so the two carry different frames and should never share a chart.

Moderated and AI-moderated depth interviews sit in a different category again.

They are qualitative instruments that produce reasoning from a small number of participants, not estimates for a population, and they belong alongside the 1993 branch of this method rather than alongside a survey platform. Most mature programs increasingly run both, on separate cadences, with separate outputs.

Frequently asked questions

What is a voice of customer survey?

A voice of customer survey is a repeatable instrument for measuring how a defined customer population experiences a product or a service interaction. The term carries two meanings: a 1993 research method built on 10 to 30 one-on-one interviews, and a 2026 software category covering feedback collection, analysis and action across channels. The measurement version is typically what most teams mean.

How many responses does a voice of customer survey need?

A voice of customer survey needs enough responses for the precision the decision requires, typically 300 to 600 for a single read on the precision table in this guide, and several thousand per period to detect a small change. At 400 responses a top-two-box proportion near 70% is estimated to within roughly 4.5 points, and detecting a three-point change requires roughly 3,500 responses per period.

Is 20 to 30 customers enough for a voice of customer survey?

No, 20 to 30 is not a survey sample size. The figure comes from Griffin and Hauser's 1993 interview study, where 20 one-on-one interviews identified over 90% of the needs found across 30. It measures saturation of a need space through hour-long conversations, and Hennink and Kaiser replicated the range in 2022 at 9 to 17 interviews. It is a qualitative coverage number, not a precision figure.

What is a good response rate for a voice of customer survey?

No cross-industry benchmark for voice of customer response rates discloses its sample size and method, so this guide publishes no floor. The best-documented figure in a peer-reviewed venue is 16.2%, from a study of 173,886 chat sessions. The useful standard is the federal one: below an 80% unit response rate, United States Office of Management and Budget guidelines call for a non-response bias analysis.

How is voice of customer different from customer satisfaction measurement?

Customer satisfaction measurement is one instrument inside a voice of customer program, not a synonym for it. A satisfaction battery produces a score for a defined population, and a voice of customer program adds the other feedback sources, the analysis layer that reads open text, and the action layer that assigns a response to an owner and tracks resolution.

Should a voice of customer survey be relational or transactional?

Run both and report them separately. A relational read measures a defined customer population on a fixed calendar, and a transactional read measures the people who just completed a specific interaction. The transactional read is conditioned on that interaction, so its numbers are not comparable to a relational baseline and the two should never be merged into one trendline.

Can artificial intelligence analyze voice of customer open-text responses?

Yes, with a documented limitation. Language models code text into an existing framework about as well as human coders, but Hill et al. (2026) reported 93.5% agreement alongside a kappa of only 0.34, and agreement looks high only because code prevalence is low at 7.8%. Build the theme list once by hand, then have the model apply that frozen list each period.

Do synthetic respondents work for voice of customer research?

No, synthetic respondents are not usable for a voice of customer program. Bisbee et al. (2024) found that while synthetic averages tracked real averages closely, 48% of regression coefficients differed significantly and the sign flipped in 32% of those cases. A model also cannot report an outage you have not published or a complaint nobody has written down.

What tools do you need to run a voice of customer program?

A voice of customer program needs three capability areas: a survey platform to collect responses across the channels your customers use, an analysis layer that reads open text and cross-tabs a metric against segments, and an operational system that assigns a response to an owner and tracks resolution. Most survey platforms generally cover the first two, and the third typically lives either in a dedicated voice of customer platform or in the operational system you already run.

What is the smallest voice of customer program worth running?

The smallest program worth running is one relational read per quarter, on a frozen four-item battery with two open-text probes, fielded to a written-down frame, with the five denominator counts recorded and one named owner for flagged responses. Below that you are collecting comments instead of measuring, which is a legitimate activity with a different name and no trendline attached to it.

How often should you run a voice of customer survey?

Run relational reads on a cadence you can sustain for at least four periods with a frozen instrument, and trigger transactional reads on completed interactions. Quarterly is the common choice and is a judgment, not a sourced rule. Raising the cadence lowers responses per period, which widens the critical difference and makes real changes harder to detect.

The bottom line

If your question is what customers need from something you have not built yet, run the 1993 method: 10 to 30 one-on-one interviews, coded into a needs hierarchy, sized by saturation.

If your question is how a defined population experiences what you already ship, run a frozen closed battery with semi-open probes, on a stated frame, with the response rate recorded.

Either way, analyze who did not answer before you report what the people who did answer said.

The federal standard asks for it below an 80% response rate, very few programs clear that threshold, and the analysis typically takes an afternoon once the composition data exists.

Then build the loop, in the operational system that already tracks ownership and resolution, using the operating spec above.

The action layer is where the industry definition of this category actually lives, and it is the part a survey platform generally does not supply.

The next study worth running is a properly sampled relational read with the response rate recorded and the non-response worksheet filled in. The customer satisfaction template is the concrete artifact to start from.

Back to top
Solutions
Experience measurementStrategic & foundational discoveryJourney & behavioral researchMarket & consumer insightsConcept & prototype testing
Agents
DesignFieldSynthesize
Deploy
EmailPanelsWeb apps and websitesMobile app
Pricing
Community
EventsBlogGuides
CustomersIntegrationsCompare
Company
About usCareersService agreementPrivacy policyData addendumSystem status
Socials
LinkedInX