Introduction
Running a customer satisfaction survey takes five decisions in order. Decide whether you are measuring a transaction or a relationship, because they are different instruments fielded to different populations.
Pick a scale and freeze it, since changing the number of points shifts the level of the score on its own.
Report the mean and the top-two-box share together, and publish the distribution behind them. Size the sample from the precision you need. Then benchmark against your own prior periods, because no comparable cross-company benchmark exists.
Most programs report top-two-box alone. That convention merges two groups a 1995 Harvard Business Review study, practitioner work rather than peer-reviewed, found were six times apart on repurchase at Xerox.
At 200 responses a top-two-box near 80 percent carries a margin of error of roughly 5.5 points, and detecting a four-point move between periods takes closer to 1,450.
This guide covers:
- Separate the transactional read from the relationship read before writing a question
- Freeze the scale, because scale length moves the score without anything changing
- Report the mean, the top-two-box and the distribution together, never the box score alone
- Derive the sample from the precision you need instead of copying a threshold
- Read the American Customer Satisfaction Index as an index, because it is not a customer satisfaction score
Platform capabilities and documentation cited here were verified against sprig.com and docs.sprig.com on September 17, 2026. No prices appear anywhere in this guide.
What customer satisfaction measures
Customer satisfaction, commonly abbreviated CSAT, is a stated evaluative judgment: how satisfied a person says they are, at one moment, with a transaction, a product, or a relationship.
The judgment is evaluative rather than behavioral. It records an opinion about an experience instead of counting what the person did next, and most of what follows in this guide comes from that distinction.
There is no canonical source, and that matters
Unlike Net Promoter Score, customer satisfaction has no founding article, no owned question wording, and no licensing regime. There is a theoretical lineage and a set of convergent conventions, and that is all.
The lineage runs through three papers. Cardozo, R. N. (1965), "An Experimental Study of Customer Effort, Expectation, and Satisfaction," Journal of Marketing Research 2(3), 244 to 249, is the earliest experimental satisfaction study, and it is about effort and satisfaction together.
Oliver, R. L. (1980), "A Cognitive Model of the Antecedents and Consequences of Satisfaction Decisions," Journal of Marketing Research 17(4), 460 to 469, established expectancy-disconfirmation theory, which every subsequent satisfaction measure rests on.
Fornell, C., Johnson, M. D., Anderson, E. W., Cha, J. and Bryant, B. E. (1996), "The American Customer Satisfaction Index: Nature, Purpose, and Findings," Journal of Marketing 60(4), 7 to 18, produced the first national standardized satisfaction measurement.
The practical consequence is that every convention in circulation is a convention rather than a standard. Nobody can tell you that your wording is wrong by pointing at a specification, and nobody can tell you that your number is comparable to theirs either.
The question wordings in common use
Two forms dominate, and neither is owned by anyone.
The transactional form asks "How satisfied were you with [the interaction]?" on a five-point scale anchored Very dissatisfied to Very satisfied. The relationship form asks "Overall, how satisfied are you with [the product or company]?" on a five-point or seven-point scale.
The score is typically reported either as the mean of all responses or as the top-two-box percentage, which is the count of 4s and 5s divided by total responses, multiplied by 100.
What the score does not capture
A satisfaction score reports what a person says about an experience rather than what they subsequently do about it, and the gap between those two things is the subject of most of the research in the next several sections.
It also carries no attribution. The number tells you the level and never the cause, which is why a satisfaction program without a driver study attached usually reports the same finding for years.
When to use customer satisfaction
Use a satisfaction measure when the customer has made an evaluative judgment you want recorded, when you can field the same instrument repeatedly to a population you can name, and when your response volume supports precision narrower than the movement you plan to act on.
Four conditions that make it the right instrument
The first is evaluative content. Satisfaction generally fits where the customer formed an opinion about quality, value, or fit, which is different from measuring how hard something was to do.
The second is comparability over time. A frozen satisfaction item on a stable population is one of the few measures that typically means roughly the same thing to a researcher and to an executive.
The third is breadth of application. The same instrument works across a purchase, a support interaction, an onboarding flow and an annual relationship read, provided you keep those four populations separate.
The fourth is volume. Satisfaction items commonly attract higher response rates than longer instruments, which is what makes the precision arithmetic in this guide achievable for most programs.
Transactional or relationship, and why they are not the same study
The transactional read fires after a specific event and measures that event. The relationship read fires on a schedule and measures the accumulated impression.
They produce different numbers from different populations, and combining them generally produces a figure that describes neither. Run both if you need both, report them separately, and never average them.
Where the metric typically earns its place
Satisfaction is often the right instrument when the decision in front of you is about priorities rather than about diagnosis. A driver model built on a satisfaction item will rank what to fix, which is a question most other instruments in this family cannot answer at all.
What kind of evidence this produces
A satisfaction survey produces directional evidence about a population's stated opinion. It does not produce causal evidence about what drives that opinion, and it does not produce behavioral evidence about what the population will do.
Treating it as the first of those three is the source of most of the disappointment programs report with this metric.
When not to use customer satisfaction
Five things a satisfaction score cannot tell you. Each names the instrument that answers the question instead.
It cannot tell you whether the customer will stay or buy again
This is the load-bearing limitation, and the evidence is consistent across two large studies.
Kumar, V., Dalla Pozza, I. and Ganesh, J. (2013), "Revisiting the satisfaction-loyalty relationship," Journal of Retailing 89(3), 246 to 262, report that satisfaction explains less than 25 percent of the variance in repeat purchase, and that satisfaction alone explains approximately 8 percent.
Models adding moderators, mediators and antecedents raise explained variance to an average of 34 percent.
Their conclusion is that satisfaction is a necessary but not a sufficient condition for predicting loyalty.
Keiningham, T. L., Cooil, B., Aksoy, L., Andreassen, T. W. and Weiner, J. (2007), Managing Service Quality 17(4), 361 to 384, followed a two-year longitudinal panel of over 8,000 United States customers across retail banking, mass-merchant retail and internet service providers.
Their multi-predictor retention models reached adjusted R-squared of 7.3 percent, 10.3 percent and 7.5 percent respectively.
Those were multi-predictor models rather than single-item ones.
The instrument that answers this is observed retention, renewal and repeat-purchase data taken from the customer record and modelled directly. Where an attitudinal leading indicator is genuinely wanted, Keiningham's own recommendation is a multiple-indicator model instead of a single question.
It cannot tell you which of your satisfied customers are about to leave
This limitation is mandatory reading for anyone reporting top-two-box, and it is the sharpest argument in the guide.
Jones, T. O. and Sasser, W. E. Jr. (1995), "Why Satisfied Customers Defect," Harvard Business Review, November to December 1995, is a practitioner article rather than a peer-reviewed study, and it should be cited as one.
Verbatim: "Its totally satisfied customers were six times more likely to repurchase Xerox products over the next 18 months than its satisfied customers." In retail banking, "completely satisfied customers were nearly 42% more likely to be loyal than merely satisfied customers."
The scale in question was a five-point one, where 4 meant satisfied and 5 meant completely satisfied.
Top-two-box scoring merges those two groups by construction. A program reporting only the box score is deliberately discarding the distinction that study found was worth a factor of six.
The instrument that answers this is the full response distribution, with the 5-share tracked as its own series. For diagnosis, a follow-up open-ended item or a small set of depth interviews with 4-raters.
It cannot tell you why the score is what it is
A satisfaction item is a single summary judgment with no attribution to any driver. It reports a level and never a cause.
The instrument that answers this is a key-driver study, which models attribute-level ratings against overall satisfaction, or task-based usability research where the object is a specific flow.
There is structural support for putting the effort upstream. Otto and colleagues, in the meta-analysis cited later in this guide, conclude that satisfaction is more appropriately depicted as mediating the effects of marketing-strategy variables on firm performance, which means the actionable variance sits above the satisfaction item rather than inside it.
It cannot tell you how you compare to anyone else
No comparable cross-company satisfaction benchmark exists, for reasons the benchmark section sets out in full.
Scale length alone shifts the aggregate mean, so two organizations running different instruments are not producing comparable numbers even before sampling differences are considered.
The instrument that answers this is your own trend on a frozen instrument. Where an external comparison is genuinely required, the American Customer Satisfaction Index covers the industries it covers, and it must be read as an index rather than as a satisfaction score.
It cannot tell you how the people who did not respond feel
Peterson, R. A. and Wilson, W. R. (1992), "Measuring customer satisfaction: Fact and artifact," Journal of the Academy of Marketing Science 20(1), 61 to 71, state the problem directly.
Verbatim from the abstract: "Self-reports of customer satisfaction invariably possess distributions that are negatively skewed and exhibit a positivity bias. Examination of the customer satisfaction literature and empirical investigations reveal that measurements of customer satisfaction exhibit tendencies of confounding and methodological contamination and appear to reflect numerous artifacts."
Dawes, in the scale study cited later, independently corroborates the skew across three different scale formats.
The instrument that answers this is a probability-sampled relationship survey of the full customer base instead of an intercept on responders, paired with behavioral instrumentation for the population that never answers.
No verified study quantifying satisfaction non-response bias specifically was located for this guide. Make the argument from the artifact finding and from sampling logic, and do not attach a number to it.
Study design, and the fork that decides everything else
Three decisions carry the design: which read you are running, which scale, and what you report.
The recommendation
Run the transactional read as a frozen five-point item fired immediately after the event, run the relationship read as a separate scheduled study, and report the mean, the top-two-box and the distribution together on both.
Each element of that answers a specific limitation named above, and the research behind each is set out below.
The transactional and relationship fork
Decide this before writing a question, because it determines the population, the cadence, the channel and the interpretation.
A transactional study fires on an event. Its population is everyone who had that event, its cadence is continuous, and its number describes the quality of a process.
A relationship study fires on a schedule. Its population is your customer base or a defined segment of it, its cadence is quarterly or annual, and its number describes the accumulated standing of the brand with those customers.
The common failure is running a transactional instrument, reporting it as a relationship metric, and then being unable to explain why it disagrees with renewal behavior. A post-ticket satisfaction score describes your support organization and generalizes to nothing else.
Why the scale length matters more than the scale choice
Preston, C. C. and Colman, A. M. (2000), "Optimal number of response categories in rating scales," Acta Psychologica 104(1), 1 to 15, tested scales from two to eleven points plus a 101-point scale across 149 respondents.
Their finding: two-point, three-point and four-point scales performed relatively poorly on reliability, validity and discriminating power. Indices rose with the number of categories up to about seven. Test-retest reliability tended to decrease beyond ten.
Two consequences follow for a satisfaction program.
The near-universal five-point satisfaction scale is defensible rather than optimal. Five sits above the poorly performing two-to-four band, but nothing in the evidence establishes it as the best choice, and a guide that tells you five is correct is overstating what is known.
The three-point good-neutral-bad widget common in helpdesk and ticketing tools sits squarely in the band their study found performs worst. Teams typically adopt it because it is the default in the tool rather than because anyone chose it.
Why changing the scale breaks the trend
Dawes, J. (2008), International Journal of Market Research 50(1), 61 to 77, randomly assigned respondents to five-point, seven-point and ten-point versions of the same items, with group sizes of 250, 185 and 300.
After rescaling to a common metric, the ten-point format produced an aggregate mean 0.3 points lower, at p equals 0.04. Skewness and kurtosis did not differ significantly, and all formats produced negatively skewed data.
The shape of your data is generally robust to scale length. The level is not.
Freeze the instrument at wave one and treat any change to it as a series break. This matters more than it appears to, because scale changes are usually made for good local reasons: a new tool defaults to five points, a redesign shortens the survey, a regional team standardizes on a different form.
Each change is defensible on its own and each one breaks the comparison.
Whether to label every point
The answer depends on what you will do with the number, and there is a clean rule almost no vendor guide states.
Weijters, B., Cabooter, E. and Schillewaert, N. (2010), International Journal of Research in Marketing 27(3), 236 to 247, manipulated labeling and category count across 1,207 respondents.
Fully labeling all categories increases acquiescence and decreases both extreme responding and misresponse. Endpoint-only labeling shows better criterion validity in regression models. Their stated recommendation for modeling work is a five-point or seven-point scale with endpoint labels only.
So label every point when you will report the distribution to stakeholders, because a fully labeled scale is what makes a distribution chart interpretable to a non-specialist. Label endpoints only when the score will feed a model.
The scale decision, as a table
Four choices, and each one has an answer in the evidence rather than in convention.
| Decision | Recommendation | Why |
|:-------------------------:|:-------------------------------------------------------------------:|:-------------------------------------------------------------------------------------------------------------------------:|
| Number of points | Five, or seven for a relationship read | Indices improve up to about seven and two-to-four point scales perform poorly. Three-point widgets sit in the worst band |
| Labeling | All points if you report distributions, endpoints only if you model | Full labeling raises acquiescence and lowers misresponse. Endpoint labeling shows better criterion validity in regression |
| Midpoint | Include it, and track its share separately | A substantial minority are genuinely neutral, but most midpoint selections are hidden non-answers |
| Changing any of the above | Treat as a series break | Scale length alone shifted the aggregate mean by 0.3 points at p equals 0.04 |
The fourth row is the one that costs programs the most, because the first three are usually decided once and the fourth is decided repeatedly without anyone noticing.
Whether to include a neutral midpoint
Sturgis, P., Roberts, C. and Smith, P. (2014), "Middle Alternatives Revisited," Sociological Methods and Research 43(1), 15 to 38, probed respondents who selected the middle option across a sample of 3,113.
Between 69 and 84 percent of middle-alternative responses turned out to be hidden "don't knows" instead of genuine neutrality. The authors nonetheless recommend keeping the midpoint, because a substantial minority do hold genuinely neutral positions.
One caveat belongs with that finding, and the guide states it rather than burying it. Those were low-salience attitude items, not post-transaction ratings where the respondent has just had the experience.
The generalization to satisfaction measurement is an inference and not a finding.
The practical habit is to track the midpoint share alongside the score instead of folding it into the distribution. A midpoint that moves while the tails do not is a signal about who is answering.
Writing the instrument
A satisfaction instrument is short by design. Three or four questions is generally the working ceiling for a transactional read, and each one has to earn its place.
The core item
For a transactional study: "How satisfied were you with [the interaction]?"
For a relationship study: "Overall, how satisfied are you with [the product or company]?"
Name the object explicitly instead of relying on context. A question that says "How satisfied were you?" after an email subject line about a support ticket is typically ambiguous to anyone who opens the email an hour later, and ambiguity in the object is the most common wording defect in this instrument.
Scale: five points, Very dissatisfied to Very satisfied, with all five labeled for a program that will report distributions.
The instrument spec, ready to lift
The four-question transactional form below is complete. Substitute the bracketed terms and field it.
| Order | Question | Type | Notes |
|:-----:|:----------------------------------------------:|:----------------------------------------------------------------:|:-----------------------------------------------------------------:|
| 1 | How satisfied were you with [the interaction]? | 5-point, Very dissatisfied to Very satisfied, all points labeled | The scored item. Freeze the wording and the scale |
| 2 | Did you get what you needed? | Yes / No / Not yet | Separates satisfaction from outcome |
| 3 | What would have made this better? | Open text, shown only on a 1, 2 or 3 | The cause. Conditional display is what keeps the instrument short |
| 4 | What worked well? | Open text, shown only on a 5, retired after the baseline period | Establishes what to protect, then stops earning its place |
For the relationship form, replace question 2 with a renewal or continuation intention, hold questions 1, 3 and 4 constant, and add the attribute battery that feeds the driver model.
Why the probe is worded that way
"What would have made this better?" asks for a specific change and generally returns codeable causes. "Why did you give that score?" asks the respondent to justify themselves, and typically returns defensiveness or a shrug.
Avoid the generic "Any additional feedback?" entirely. It commonly returns a mix of praise, unrelated product requests and blanks, and the yield of usable material is frequently a fraction of what a targeted probe produces.
The outcome item, and why it is not optional
Satisfaction and outcome are different constructs. A program measuring only the first cannot distinguish a customer who is satisfied because the process was pleasant from one who is satisfied because the problem is solved.
Three options rather than two matters. A customer whose issue is open but progressing is neither a success nor a failure, and forcing them into one distorts both groups.
The "Not yet" share is also a useful operational signal on its own. A rising proportion typically indicates cases are being closed in the system before they are closed for the customer.
Fields to attach instead of asking
Attach these from your own systems: interaction type, channel, tenure, plan or segment, product area, and an identifier that lets you join the response to the operational record.
Every one is a cut you will want at analysis time and every one is already in your systems. A survey that spends a question asking what your customer record already knows is trading the only attention you get for data you already had.
The join identifier is the field most often forgotten and the hardest to add later. Without it the driver analysis in this guide is not possible at all.
What to leave out
Leave out demographics on a transactional survey. Leave out a second open-text box, which typically produces one answered field and one blank.
Leave out any question whose answer will not change a decision, which on a four-question instrument is a stricter filter than it sounds.
One wording trap worth naming
Do not use a satisfaction scale to ask about agreement, and do not mix the two within one program.
"How satisfied were you?" on a Very dissatisfied to Very satisfied scale and "I was satisfied with this interaction" on a Strongly disagree to Strongly agree scale are different items with different response distributions.
Mixing them across a program generally produces a series that is not internally comparable, and the change is easy to make during a redesign without anyone registering it as a measurement change.
One more trap belongs here. Do not let a translated instrument drift. Translated satisfaction scales are known to produce different response distributions across languages, and a multi-market program that compares markets on an unvalidated translation is measuring the translation as much as the experience.
Scale and scoring choices
Report two numbers and publish a third view. This section explains why one number is never enough.
Top-two-box
The convention is the percentage selecting 4 or 5 on a five-point scale. It is simple, it communicates to an audience that will not read a distribution, and it is what most stakeholders expect.
Its weakness is structural. Collapsing five categories into two discards the distribution, and the discarded distinction is the one Jones and Sasser found mattered most.
Consider three programs, all reporting a top-two-box of 80 percent:
| Rated 5, completely satisfied | Rated 4, satisfied | Top-two-box |
|:-----------------------------:|:------------------:|:-----------:|
| 55 percent | 25 percent | 80 percent |
| 40 percent | 40 percent | 80 percent |
| 25 percent | 55 percent | 80 percent |
Those three programs report an identical score. On the Jones and Sasser evidence they have materially different repurchase prospects, and the box score cannot distinguish them.
The mean
Report the arithmetic mean alongside the box score, to one decimal place.
The mean uses the full distribution and moves when the box score does not. In the three programs above, the means differ even though the box scores do not, which is the point.
A rising top-two-box with a falling mean generally indicates the distribution is polarizing instead of improving. That pattern is common when a change helps most customers and badly breaks the experience for a minority.
The distribution
Publish the full five-point distribution at least quarterly, and put the 5-share on its own line in any executive report.
Tracking the 5-share separately is the single cheapest improvement available to most satisfaction programs. It costs nothing, it uses data you already have, and it restores the distinction that the headline convention throws away.
A distribution chart also changes the conversation in a review. A single number invites the question of whether it went up, while a distribution invites the question of which group moved, and the second question is the one with an action behind it.
Which is more predictive, and why the literature disagrees
The honest answer is that it depends on the unit of analysis, and the disagreement is worth carrying rather than resolving.
Morgan, N. A. and Rego, L. L. (2006), Marketing Science 25(5), 426 to 439, analyzed American Customer Satisfaction Index data from 1994 to 2000 against future business performance.
Verbatim: "Average satisfaction scores have the greatest value in predicting future business performance and that Top 2 Box satisfaction scores also have good predictive value."
de Haan, E., Verhoef, P. C. and Wiesel, T. (2015), International Journal of Research in Marketing 32(2), 195 to 206, found the opposite ordering at the customer level, with top-two-box correlating with two-year retention at .184 and a multivariate coefficient of .482, against mean satisfaction at .151 and .117.
The correlations look close. The coefficients do not.
Morgan and Rego use firm-level index data to predict firm performance. de Haan uses customer-level data to predict individual retention. The disagreement tracks the unit of analysis, which is why reporting both numbers is the defensible practice rather than a hedge.
Sample size and precision
No methodology source publishes a satisfaction-specific sample-size rule with a stated derivation. Not the American Customer Satisfaction Index, not any peer-reviewed source located for this guide, and not any vendor page read for it.
That absence is worth stating plainly, because the thresholds circulating in vendor content have nothing behind them. A number presented without a derivation is a convention someone repeated, and a convention cannot tell you whether your own result is precise enough to act on.
Derive the sample from the precision you need instead.
The arithmetic for a box score
Top-two-box is a proportion, so the interval around it is a binomial proportion interval. The defensible published method is the adjusted Wald interval described by Agresti, A.
and Coull, B. A. (1998), "Approximate is Better than 'Exact' for Interval Estimation of Binomial Proportions," The American Statistician 52(2), 119 to 126.
It outperforms both the standard Wald interval and the exact Clopper-Pearson interval at small sample sizes.
The adjustment is simple enough to do by hand. Add two successes and two failures to the observed counts, then compute the ordinary interval on the adjusted proportion.
Margins of error at 95 percent confidence, derived using that method:
| Responses per period | At a top-two-box of 80 percent | At a top-two-box of 65 percent |
|:--------------------:|:------------------------------:|:------------------------------:|
| 50 | plus or minus 11.1 points | plus or minus 12.8 points |
| 100 | plus or minus 7.8 points | plus or minus 9.2 points |
| 200 | plus or minus 5.5 points | plus or minus 6.6 points |
| 400 | plus or minus 3.9 points | plus or minus 4.7 points |
| 800 | plus or minus 2.8 points | plus or minus 3.3 points |
| 1,200 | plus or minus 2.3 points | plus or minus 2.7 points |
One row worked in full, so you can reproduce the rest. At 200 responses with 160 in the top two boxes, add two successes and two failures to get 162 of 204, an adjusted proportion of 0.794.
The margin is 1.96 times the square root of 0.794 times 0.206 divided by 204, which is 0.055, or 5.5 points.
The table is derived rather than cited, and the derivation is the asset. No source publishes a rule, so the honest alternative to an invented threshold is showing the arithmetic and letting the reader set their own floor.
Note the second column. Precision is worse near the middle of the range and better at the extremes, so a program scoring 65 percent needs more sample than one scoring 80 percent to make the same claim.
That asymmetry has a practical consequence worth planning for. A program that improves from 65 to 80 percent gains precision as it goes, so the sample that was barely adequate at the start will generally be comfortable later, and the sample that is adequate for a struggling segment is more than adequate for a healthy one.
Comparing two periods needs more than this
A movement between periods is a difference between two proportions, and its interval is wider than the interval around either figure. This is where most satisfaction reporting goes wrong.
Take a move from 80 to 84 percent. At 100 responses per period the margin on the difference is about 10.6 points. At 200 it is about 7.5 points.
At 400 it is about 5.3 points, so a four-point move is still inside it. At 800 it is about 3.8 points, and only there does the move clear.
Those are intervals around a difference you have already observed. Planning is a different question, and a study sized at 800 per period will only see a genuine four-point move about half the time.
Detecting a four-point movement at 80 percent power takes roughly 1,450 responses per period. Detecting a two-point movement takes roughly 6,050.
That second figure is the one worth sitting with.
Most satisfaction programs report to one decimal place and discuss movements of a point or two, and almost none has the sample to support that conversation. Publishing the required sample alongside the target is generally the fastest way to end it.
For the mean
Use a t-interval driven by the observed standard deviation, computed from your own data rather than assumed. Dispersion on a satisfaction item varies with the skew of the distribution, and satisfaction distributions are reliably skewed.
The mean is generally the more sensitive of the two statistics at a given sample size, because it uses the full distribution instead of a binary split. Where sample is tight, the mean will often detect a movement the box score cannot.
What to do when the volume is not there
Below a few hundred responses per period, a satisfaction score stops supporting period-over-period comparison.
The honest response is to change what you report rather than to report it with more confidence. Move to quarterly reporting, report the level with its interval and stop drawing a trend line, or switch the effort to qualitative research on the segment you care about.
Audience, targeting and screening
The population for a satisfaction study is defined by the fork in the design section. Get that right and targeting is mostly mechanical.
Sampling a transactional study
Everyone who had the qualifying event is eligible. Sample a fixed random fraction of them rather than surveying all of them, because a program that surveys every event trains customers to ignore it.
Set the fraction from the precision arithmetic. If 400 responses per month gives you the precision you need and your response rate is 20 percent, you need roughly 2,000 invitations.
Random selection matters more than it appears to. "Every fifth ticket" is not random when ticket order correlates with anything, and in most operational queues it does.
Sampling a relationship study
The population is your customer base or a named segment of it, and the sampling problem is representativeness rather than volume.
Draw a probability sample from the full base rather than surveying whoever is most reachable. The most reachable customers are typically the most engaged ones, and a satisfaction score built on them is measuring your best customers and reporting it as your customer base.
Stratify by the dimensions you will cut the results on. A study that reports satisfaction by plan tier needs enough sample in each tier to support the claim, and that is a design decision rather than an analysis one.
Small segments are the usual casualty. An enterprise tier with forty customers will never support a quarterly satisfaction figure at useful precision, and the honest treatment is to report it annually or to research it qualitatively rather than to publish a number with an interval nobody reads.
Who to exclude
Exclude customers inside your recontact window. Exclude interactions the customer did not initiate, unless you are deliberately measuring proactive outreach as its own category.
Do not exclude unresolved or negative interactions. Those are frequently the most informative cases in the dataset, and removing them is how a program produces a healthy score alongside an unhealthy operation.
Screening
Most satisfaction programs need no screener, because eligibility is established by the event or by the sample frame.
The exception is a relationship study spanning multiple products. There, one screening question establishing which product the respondent is rating belongs at the top, and it should be a single-select.
The non-response problem
Satisfaction surveys are answered by a self-selected minority, and Peterson and Wilson's positivity bias operates on top of that selection.
Track response rate by segment as a finding rather than as an operational detail. A segment whose response rate is falling while its score is rising deserves scrutiny, because the most frequent explanation is that dissatisfied customers in that segment have stopped answering.
Fielding and delivery
An in-product intercept and an emailed invitation reach different halves of your base, and the difference shows up in the score before it shows up anywhere else.
In the product
For a transactional read on a digital interaction, firing inside the interface immediately after the event is generally the strongest option. The experience is unmediated by recall, and the response rate is typically higher than email.
This is also the channel where behavioral context attaches most reliably, because the survey and the behavior happen in the same system. That join is what later makes the driver analysis possible.
Email
For agent-resolved interactions and for relationship studies, email commonly remains the practical channel.
Send transactional invitations within the same business day, for the recall reason set out in the timing section. For a relationship study, timing matters less and consistency matters more, so field in the same window every wave.
Keep the first question in the email body where the channel supports it, because every click between the customer and the first answer is a place to lose one.
Avoid sending from a no-reply address, because customers often reply to satisfaction surveys with the detail that explains their score and a program that discards those replies is throwing away its own verbatims.
Panels, for a market-level read
Where the question is about a market rather than about your own customers, a research panel supplies a sample frame you do not own.
That is a different study with a different interpretation, and the resulting number is generally not comparable to a score from your own customer base. Report it as a separate series.
What to avoid
Avoid mixing channels inside a reported segment without checking the mix, because channel differences are frequently larger than the operational differences a program is trying to measure.
Avoid measuring a specific transaction inside a periodic relationship survey. A customer asked in March about a February interaction is reconstructing an event, and the accuracy of that reconstruction is not something the survey can check.
Running this in Sprig
Sprig fires in-product surveys from behavioral events, which is the trigger a transactional satisfaction read needs, and the recontact waiting period gives a documented fatigue control for a high-volume program. The Rating Scale question type supports the five-point and seven-point forms this guide discusses.
Two limitations to know before you build. The recontact window is enforced only in production environments, so a development environment will over-survey during testing. And Sprig publishes no range specification for the Rating Scale in its documentation, so confirm the configuration in the product instead of relying on a documented maximum.
Timing and cadence
Transactional and relationship studies have opposite timing requirements, which is one more reason to keep them separate.
Timing a transactional study
Fire as close to the event as the channel allows. Satisfaction with a specific interaction is a judgment about a recent experience, and the judgment degrades as the experience recedes.
One deliberate exception. Where the outcome is not yet known to the customer, surveying at the moment of internal closure measures your process rather than their experience.
Wait until the customer-facing resolution is confirmed, and record the delay so the series stays comparable.
Timing a relationship study
Quarterly or annual, fielded in the same window each time.
Seasonality is real in most categories, and a relationship study that moves its fielding window has introduced a confound it cannot separate from the result.
How often to report
Set reporting cadence from the precision table rather than from the meeting calendar. This is frequently the easiest improvement available to an existing program.
A monthly number on 150 responses cannot support a trend claim. A quarterly number on 450 can, and the quarterly version will be believed for longer.
How often to survey the same customer
Set a recontact window and enforce it in the tooling rather than in a policy document. Derive it from your own contact-frequency distribution: the interval inside which a given customer would be surveyed twice is the interval to suppress.
For customers who interact frequently, sample one interaction rather than all of them, and choose which one deliberately.
Quality control and data hygiene
Four checks, run before any number leaves the analysis.
Straight-lining and speed
Flag responses completed faster than a plausible reading time for the instrument. Time it yourself by reading it aloud at speed and set the threshold from that rather than from a convention.
Flag rather than delete automatically, and report how many you flagged and what the score looks like with and without them. A deletion rule applied silently is generally indistinguishable from a thumb on the scale.
On a four-question instrument, straight-lining is harder to detect than on a long battery, because there is little to straight-line. Speed is the more useful signal here.
Non-response handling
Decide the rule before looking at the data: responses missing the satisfaction item are excluded from the satisfaction calculation, and responses missing other items stay in the dataset for their own questions.
State the rule in the report. An analysis that quietly drops partial responses produces a different number than one that does not, and nobody reading the chart can tell which they have.
Mix stability
Check the composition of each period against the previous one, on interaction type, channel and segment. A score that moved because the mix moved is not a change in customer satisfaction.
This is commonly the largest source of unexplained movement in a satisfaction program, and a simple table of volume share by segment usually reveals it in a few minutes.
Duplicate and test responses
Check for duplicate responses against the same interaction identifier, and remove internal test submissions. Both are more frequent in the weeks after a survey is rebuilt than at any other time.
Read the impact off the precision table. A handful of test responses barely moves a figure built on 800 and visibly moves one built on 50.
Run all four checks in the same pass and record the outcome, even when nothing is flagged. A hygiene log is what lets you rule out a data explanation the next time a number moves unexpectedly.
Analysis: the core calculation
The arithmetic is simple. The discipline is in the denominator, the segmentation and the comparison.
Computing the score
Top-two-box equals the count of 4 and 5 responses divided by the count of valid responses to the satisfaction item, multiplied by 100. Exclude blanks and non-responses from the denominator instead of treating them as zeros.
The mean is the arithmetic mean of the same valid responses, to one decimal place. The 5-share is the count of 5s over the same denominator.
Report all three with the response count beside them. A score without its denominator is not interpretable, and a score published without one is frequently a score nobody checked.
The reporting template
The three numbers belong in one place, every period, in this shape.
| | This period | Prior period | Change | Margin on the difference | Real? |
|----------------|:-----------:|:------------:|:--------------:|:------------------------:|:----------------:|
| Mean | 4.1 | 4.2 | minus 0.1 | 0.13 | No |
| Top-two-box | 80 percent | 78 percent | plus 2 points | 7.5 points | No |
| Share rating 5 | 40 percent | 46 percent | minus 6 points | 9.1 points | No, but watch it |
| Share rating 4 | 40 percent | 32 percent | plus 8 points | 9.0 points | No, but watch it |
| Share rating 3 | 12 percent | 13 percent | minus 1 point | 6.3 points | No |
| Share rating 2 | 5 percent | 6 percent | minus 1 point | 4.4 points | No |
| Share rating 1 | 3 percent | 3 percent | no change | 3.4 points | No |
| Response count | 200 | 200 | | not applicable | |
Read the 4-row and the 5-row together. Jones and Sasser found those two groups were six times apart on repurchase, and top-two-box merges them. In the worked example above the box score rose two points while the 5-share fell six, which is the pattern the headline number is built to hide.
Neither movement clears its interval at 200 responses per period, which is why the final column says so rather than leaving the reader to guess.
The final column is the one that changes behaviour. A change smaller than the margin on the difference is reported as no change, in writing, rather than discussed as a movement.
Publishing the margin beside the number is the single cheapest credibility improvement available to a satisfaction program, and almost no program does it.
Segmenting
Cut by interaction type, channel and segment at minimum, and add tenure for a relationship study.
Report the 5-share by segment as well as the box score. That cut frequently reveals that two segments with the same headline number have quite different distributions behind them, which is generally the most useful thing a segmentation pass produces.
Add agent or team as a cut only with deliberate care. It is operationally useful and it changes how the survey is perceived internally, and a program that becomes an individual performance measure will typically start producing managed results.
Comparing periods and segments
Compute the interval on the difference before comparing, using the figures in the sample-size section. Two proportions whose intervals overlap substantially are not different.
Sprig documents no statistical significance testing anywhere in the product, so this comparison generally happens in a spreadsheet, a statistics tool or an AI client.
Where you scan many segments at once, expect some to look different by chance. At 95 percent confidence one false positive is expected per twenty comparisons, and the chance of at least one across twenty is about 64 percent, so treat multi-segment scans as hypothesis generation instead of as findings.
The driver analysis that makes the score useful
Join the response to the operational record and to any attribute battery you field, then model overall satisfaction against the attribute ratings.
That model is what turns a level into a priority list. A satisfaction program that never runs one will commonly report a number for years without identifying anything that would move it.
Use derived importance from the model rather than stated importance from a direct question. Respondents typically rate almost everything important, which leaves stated-importance data with too little discrimination to rank anything.
The output to take to a review is a short list of attributes with high derived importance and low performance. That quadrant is where the available improvement is, and it is the only output of a satisfaction program that reliably survives contact with an operations team.
Rebuild the model rather than reusing last quarter's coefficients. Driver importance moves as the experience changes, and an attribute that mattered when it was broken typically stops mattering once it is fixed.
Benchmarks and what a good result looks like
There is no sourced cross-industry customer satisfaction benchmark that is comparable to a score from your own survey. The one genuine national index is real, documented, and measuring something else.
The American Customer Satisfaction Index is real and it is not a customer satisfaction score
This is the most valuable distinction in the section and it is almost universally got wrong.
The index is methodologically documented, and described in Morgeson, F. V. III, Hult, G. T. M., Sharma, U. and Fornell, C. (2023), "The American Customer Satisfaction Index (ACSI): A sample dataset and description," Data in Brief 48, 109123.
Its satisfaction construct is not one question. It is three, each on a ten-point scale, quoted verbatim from that paper:
"Considering all of your experiences to date with the company/brand, how satisfied are you?"
"Considering all of your expectations, to what extent has the company/brand fallen short of or exceeded your expectations?"
"Forget the company/brand you bought for a moment. Imagine an ideal product. How well do you think the company/brand you bought compares with that ideal?"
Those three sit inside a six-construct structural model covering expectations, perceived quality, perceived value, satisfaction, complaints and loyalty.
The score is then produced, verbatim, by "a proprietary and patented Partial Least Squares structural equation modeling approach," with proprietary weighting, and reported after transformation on a 0 to 100 scale.
So an index score of 78 and a top-two-box of 78 percent are not the same kind of number. They come from different instruments, different estimation, different sampling and different scaling.
Any guide that puts them in the same table is wrong. Citing the index accurately and saying what it is not is more useful than borrowing its authority.
On sample, the index's own published figures differ. The 2023 peer-reviewed description states roughly 400,000 consumers annually across more than 400 companies, while current press materials describe roughly 200,000 interviews.
Both are the index's own numbers, three years apart, and this guide reports the discrepancy rather than picking one.
The vendor benchmark tables
Several vendors publish satisfaction benchmarks by industry. The sourcing is consistent across them and it is thin.
Retently, SI Labs, Fullview, ProProfs, Simplesat, Giva, Bandwidth and Dovetail all publish tables in this shape. The most widely circulated is sourced to the publisher's own customer database, with no stated sample size, no methodology, no respondent count and no collection period.
Its own page offers no more methodological statement than a note that the publisher turned to customer data for insights.
A second class of table cites regional industry associations without sample sizes or methods. A larger class, spanning most of the pages ranking for this query, reproduces figures without sourcing them at all.
One plausible candidate for a real ticket-level database exists in the helpdesk category, computed from actual survey responses across customer accounts. Its methodology, question wording, benchmark population and sample size could not be retrieved for this guide, so its status is unverified and it is not cited here as a benchmark.
What to do instead
Benchmark against your own prior periods, on a frozen instrument, within the same interaction type and channel.
Set your first three months as the baseline and report movement against it with the interval attached. That comparison is valid, available immediately, and defensible in a review.
Where an external comparison is genuinely demanded, explaining why one does not exist typically lands better than a sourced-looking number that collapses under a single question.
Interpreting and acting on the result
Read the distribution, not the headline
Start every review with the five-point distribution and the 5-share, then look at the box score.
Reading in that order changes what gets noticed. A program whose 5-share is falling while its box score holds is losing its most loyal customers into the merely satisfied band, which is the pattern Jones and Sasser found mattered most and the one the headline number is built to hide.
The reverse is also informative. A flat box score with a rising 5-share is genuine improvement that box scoring is structurally blind to, and a program reporting only the collapsed figure will typically miss its own success.
Read movement against the interval
Compare any period-over-period movement to the margin on the difference before treating it as a change.
Most reported monthly movement in a satisfaction program sits inside that interval. A program that reacts to it is chasing noise, and it will generally lose credibility when the number reverts and the intervention takes the blame or the credit either way.
Act on the drivers, not the score
The score tells you which segment to open. The driver model and the verbatims tell you what to fix.
Work in that order every time. Opening the verbatims before the segmentation produces a list of vivid complaints with no sense of how many people hold them, which is how a program ends up fixing the loudest problem rather than the largest one.
Theme the low-score verbatims, then cross-tabulate the themes against interaction type and segment. Rank the themes by mean satisfaction rather than by count, because the largest theme is frequently a minor annoyance affecting everyone while a smaller one is often a serious failure affecting a specific group.
What a genuine improvement looks like
Satisfaction improves when a specific driver improves for a specific population. It does not improve because the number was put on a dashboard.
Tie every intervention to a driver the model identified, then look for movement in that segment instead of in the blended figure. Name the segment and the expected direction in advance, so the check afterwards is a test of a stated prediction rather than an open-ended interpretation of whatever moved.
A blended satisfaction score is usually too insensitive to detect a fix to one journey, and expecting it to move is how good work commonly gets judged ineffective.
Expect the aggregate to move slowly even when the work is going well. A fix to one interaction type in one channel is typically invisible in the blended number, which is why segment-level reporting is what keeps an improvement program credible over more than a quarter.
Set the expectation before the work starts rather than after the first flat quarter. A program that has told its stakeholders in advance where the movement will appear is generally given time to produce it.
Running the analysis with Claude or ChatGPT
The score itself is a division. What an AI client adds is the theming of open-text responses, the cross-tabulation of themes against segments, and the driver model that most programs never build.
This chapter covers what is specific to satisfaction data. The general mechanics of coding open-ended responses into themes and cross-tabbing them are covered in depth in the cross-tab analysis guides and are not rebuilt here.
What the analysis produces
Four outputs: the mean, top-two-box and 5-share by period and segment, the interval on any comparison you intend to report, a themed list of causes from the low-score probe with satisfaction attached to each theme, and a ranked driver list where you field an attribute battery.
The theme list with satisfaction attached is the output worth the effort. A theme list alone tells you what customers said. A theme list with the mean score attached tells you which complaint is costing you the most.
Getting your data in
Export responses as CSV. Include the satisfaction item, the outcome item, the open-text probe, and every attached field: interaction type, channel, segment, tenure and the interaction timestamp.
Confirm before you start that the satisfaction column holds the numeric response rather than the label text. A column of "Very satisfied" strings needs an explicit mapping step, and a model asked to infer that mapping will occasionally invert it.
If your survey data lives in Sprig, the Sprig MCP connector passes studies and responses directly into a supported AI client, which skips the export. Responses are retrieved in pages, so a larger dataset takes more than one call.
Prompt one: the score, with its interval
What this prompt does: computes customer satisfaction scores with confidence
intervals, overall and by segment, from a survey export.
What it returns: a markdown table of mean, top-two-box, 5-share, response count
and 95 percent margin of error, by period and segment.
Use code to calculate this, not estimation. Write and run Python. Do not compute
any percentage, mean or interval by reading values.
Data: [ATTACHED CSV]
Satisfaction column: [COLUMN NAME]
Scale: 5-point, 1 = Very dissatisfied, 5 = Very satisfied. Higher is better.
Segment columns: [INTERACTION TYPE, CHANNEL, SEGMENT]
Period column: [DATE COLUMN], group by [MONTH]
Rules:
1. Exclude blanks, nulls and non-numeric values from both numerator and
denominator. Do not treat a blank as a zero. Report how many you excluded.
2. Top-two-box = count of 4 or 5 divided by count of valid responses, times 100,
to one decimal place. 5-share = count of 5 divided by count of valid
responses, times 100. Mean = arithmetic mean, to one decimal place.
3. Margin of error: use the adjusted Wald interval at 95 percent confidence.
Add two successes and two failures, then compute the ordinary interval on the
adjusted proportion. Do not use the plain Wald interval.
4. Report the response count beside every figure.
5. Any segment with fewer than [30] valid responses is reported as "insufficient
responses" instead of as a number. A segment with zero valid responses is also
below this threshold and must be reported the same way, not omitted from the
table.
6. After computing, recompute the overall top-two-box by a different method: sum
the per-segment numerators and denominators and divide. Compare to the direct
calculation. If the two disagree by more than 0.1 points, report both figures
and state that they disagree. Do not silently reconcile them.
7. Save the output as csat-scores.md.
Rule 6 is the one most often missing from a first draft, and it is the one that catches the errors that matter.
A figure computed twice by genuinely different routes surfaces a mistake that a model asked to "double-check its work" will not.
Rule 5 exists because of a documented failure pattern. A model asked for a threshold will sometimes count thin cells and empty cells separately, then report only the first as the total.
Neither count is wrong. Combining them is, and naming zero explicitly is what prevents it.
Rule 3 matters more here than it looks. Asked for a confidence interval without a named method, a model will usually produce the plain Wald interval, which is the one the published comparison found performs worst at small samples.
Prompt two: themes with satisfaction attached
What this prompt does: codes open-text responses into themes and attaches the
satisfaction score for each theme.
What it returns: a theme table with response counts, mean satisfaction per theme,
5-share per theme, and three verbatim examples each.
Use code to calculate this, not estimation. Write and run Python for every count
and every mean. Assign themes by reading, then compute the arithmetic with code.
Data: [ATTACHED CSV]
Open-text column: [COLUMN NAME]
Satisfaction column: [COLUMN NAME], 5-point, higher is better
Rules:
1. Exclude blanks, single characters and responses under three words from
theming. Report how many you excluded.
2. Build the theme list from the data rather than from a list I supply. Include
an "Other or unclear" bucket and report its size.
3. Assign each response to exactly one theme.
4. For each theme report: response count, mean satisfaction, 5-share, and three
verbatim quotes. Do not paraphrase the quotes.
5. Any theme with fewer than [15] responses is reported as "low count,
directional only" instead of as a finding. A theme with zero responses is
below this threshold and must be reported the same way.
6. Recompute each theme's mean satisfaction by a second route: group the raw rows
by your theme assignment and recalculate. If any theme's two means differ,
report the discrepancy rather than correcting it silently.
7. If the "Other or unclear" bucket exceeds [15] percent of responses, say so and
recommend rebuilding the theme list instead of proceeding.
8. Save the output as csat-themes.md.
Reading the output
Read the theme table by mean satisfaction rather than by count. A small theme with a very low mean is frequently a broken experience affecting one segment, and it will not surface in a count-ranked list.
Check the "Other or unclear" bucket first. A large unclear bucket means the theme list does not fit the data, and every downstream number inherits the problem.
What to verify before reporting
Five checks, every time.
- Confirm the excluded-response counts are plausible when checked against the raw export file
- Confirm the theme counts sum to the total response count minus the stated exclusions
- Confirm the scale direction by reading three verbatims from the highest-scoring theme
- Confirm every segment you are about to report clears the stated response threshold
- Confirm the interval was computed on the difference before calling any period comparison a change
The third check catches the most embarrassing error available. If the verbatims attached to your highest satisfaction scores describe frustration, the scale has been inverted somewhere in the pipeline.
Pitfalls
Six documented failure modes.
Discovering themes is harder than applying them. Work published in PLOS Digital Health in 2026 by Hill and colleagues found models matched human analysts on deductive coding against an existing codebook, at 93.5 percent against 92.7, and were materially worse at inductive discovery.
Strict fabrication ran at 1.2 percent and comprehensive error at 12.4 percent once partial matches and misattributions were counted. The practical rule is to theme inductively once, rebuild the list yourself, then have the model apply your list every period after.
Rare themes get over-predicted, and rare themes are what get escalated. Ashwin, Chhabra and Rao, publishing in Sociological Methods and Research in 2025, found non-random bias in 10 of 19 codes, correlated with respondent characteristics, with sparse codes systematically over-predicted.
The same prompt does not produce the same answer twice. Work published by Thinking Machines Lab in September 2025 found 1,000 completions at temperature zero produced 80 unique outputs.
Neither major chat client exposes a seed, so any number you intend to report should be produced twice and compared.
Instructions placed after a large data block get lost. Liu and colleagues documented this in Transactions of the Association for Computational Linguistics in 2024. Put instructions above the data and split large files.
Models agree with you. Sharma and colleagues documented sycophancy at ICLR in 2024, and OpenAI publicly withdrew a model update in April 2025 for being overly agreeable. Never ask a model to confirm a driver you already suspect.
Verbatims contain personal information nobody asked for. Satisfaction verbatims routinely include names, account numbers and complaint details about identifiable staff. Strip them before the data leaves your environment, and check which data-handling tier your account sits in rather than which vendor it is.
The critique you should know
The literature on satisfaction measurement is larger than on any other metric in this family, and more critical than most practitioners realise.
Satisfaction explains less of behavior than programs assume
The Kumar and Keiningham findings set out earlier are the core of it: satisfaction alone explains approximately 8 percent of variance in loyalty, and multi-indicator retention models reached adjusted R-squared between 7.3 and 10.3 percent.
Those are not fringe results. They come from a Journal of Retailing review and a two-year longitudinal panel of over 8,000 customers, and they are the reason the guide recommends against presenting satisfaction as a retention forecast.
The relationship is not linear
Anderson, E. W. and Mittal, V. (2000), "Strengthening the Satisfaction-Profit Chain," Journal of Service Research 3(2), 107 to 120, established non-linearity and asymmetry in the satisfaction to retention to profit chain.
The specific coefficients could not be retrieved for this guide, so the paper is cited for the finding of non-linearity and asymmetry and no statistic is attached to it.
The practical implication is that a one-point improvement at the bottom of the scale and a one-point improvement at the top are not worth the same thing, which most satisfaction targets implicitly assume they are.
The measurement itself is partly an artifact
Peterson and Wilson's finding, quoted earlier, is that satisfaction self-reports invariably show negative skew and positivity bias and reflect numerous artifacts.
That is a strong claim from a Journal of the Academy of Marketing Science paper, and it should temper any interpretation of an absolute level. It is a better reason to track your own movement than any argument about benchmarks.
The counter-critique
Three arguments run the other way, and the first is the best single citation in this guide.
Otto, A. S., Szymanski, D. M. and Varadarajan, R. (2020), "Customer satisfaction and firm performance: insights from over a quarter century of empirical research," Journal of the Academy of Marketing Science 48(3), 543 to 564, meta-analyzed 251 correlations from 96 studies published between 1991 and 2017.
Verbatim: "the satisfaction-performance relationship is positive and statistically significant on average (r = .101)," rising to "r = .349" under the authors' most favorable contingencies.
That single result concedes and defends in one line. The average effect is small and it becomes substantial under specified conditions, which is a more defensible position than either enthusiasm or dismissal.
Satisfaction also outperforms the metric most often proposed to replace it. Morgan and Rego (2006), using index data across 1994 to 2000, found that metrics based on recommendation intentions and behavior "have little or no predictive value," and concluded verbatim that "recent prescriptions to focus customer feedback systems and metrics solely on customers' recommendation intentions and behaviors are misguided."
Whatever satisfaction fails to predict, it is not inert.
One claim to handle carefully
The argument that satisfaction moves stock prices is contested, and a guide that cites only one side of it is overstating.
Fornell, C., Mithas, S., Morgeson, F. V. III and Krishnan, M. S. (2006), Journal of Marketing 70(1), 3 to 14, reported high returns and low risk for a high-satisfaction portfolio. Jacobson, R. and Mizik, N. (2009), Marketing Science 28(5), 810 to 819, reanalyzed it and concluded that significant evidence of mispricing "is limited to firms in the computer and Internet sector," reporting monthly abnormal returns of 0.027 in that sector at t equals 2.25, against 0.0045 elsewhere at t equals 1.42, which does not reach significance.
Cite both or use the Otto meta-analysis instead, which is cleaner. The published rebuttals in the same issue were not read for this guide.
The position this guide takes
Customer satisfaction is a durable, well-validated measure of a stated evaluative judgment, and a weak standalone predictor of behavior.
Run it to track the level and movement of that judgment on a frozen instrument, pair it with a driver model to make it actionable, report the distribution rather than a single collapsed figure, and do not present it to a board as a forecast of retention.
Common mistakes
Eight of the ten failures below are covered in full earlier in the guide. This section exists so a reader who skipped to it still meets them, and the table says where the detail lives.
| Mistake | What goes wrong | Covered in |
|:-------------------------------------------:|:----------------------------------------------------------------------:|:---------------------------------:|
| Reporting top-two-box alone | The 4s and 5s are merged, and those two groups behave differently | Scale and scoring choices |
| Mixing transactional and relationship reads | One number describing two populations, explaining neither | Study design |
| Changing the scale mid-program | The level shifts without the experience changing | Study design |
| Reading movement inside the interval | Noise reported as trend, credibility spent when it reverts | Sample size and precision |
| Comparing to a vendor benchmark | A target borrowed from a table with no stated sample or method | Benchmarks |
| Comparing to the national index | An index score and a survey percentage are not the same kind of number | Benchmarks |
| Excluding negative or unresolved cases | A healthy score alongside an unhealthy operation | Audience, targeting and screening |
| Never building a driver model | A level reported for years with no identified cause | Analysis |
Two further failures are not covered elsewhere, and both are common in otherwise mature programs.
Setting a target without a distribution
A satisfaction target set as a top-two-box number can be met by moving 3s to 4s while the 5-share falls. The target is hit and the most loyal group has shrunk.
Where a target is required, set it on the 5-share or on the mean, both of which respond to the movement that matters. Publishing the target with the sample required to detect it is what keeps the target honest.
A box-score target is generally the easiest of the three to hit by accident and the least informative when it is.
Targets set before a baseline exists are worse again, because the figure is usually borrowed from a vendor table and tends to be either trivially met or unreachable.
Treating the open text as colour
Verbatims are routinely quoted in decks to illustrate a number rather than coded to explain it. That treatment wastes the only part of the instrument that carries a cause.
Code them, attach the satisfaction score to each theme, and rank by score rather than by count. The output of that exercise is a priority list, which is what the program was supposed to produce in the first place.
What these ten have in common
All ten share a root cause, and naming it is more useful than the list.
Each comes from treating the number as the finding. A program that treats the score as an index pointing at a segment will check the interval, keep the negative cases in, and build the driver model.
A program that treats the score as the deliverable will do none of those, because each one makes the number harder to report.
Synthetic respondents
Teams increasingly ask whether AI-generated respondents can stand in for real ones on a satisfaction measure. The evidence generally says aggregate means track reasonably and subgroup structure does not.
Bisbee, J. and colleagues, publishing in Political Analysis in 2024, compared 7,530 real respondents against 3.6 million synthetic responses. Averages corresponded closely to the human data.
But 48 percent of regression coefficients differed significantly and 32 percent of those flipped sign, and variance was artificially deflated, which breaks power calculations.
For a satisfaction program the constraint is sharper than that general finding suggests. A transactional satisfaction score measures a specific person's judgment of a specific event that occurred, and there is no synthetic equivalent of an event that happened.
A synthetic panel can tell you how a described persona says it would feel about a described experience. That is a different construct, and reporting it as a satisfaction score means publishing a number that resembles the real one and measures something else.
There is a further problem specific to this metric. The deflated variance the study documents would suppress exactly the distributional spread that the 5-share reporting in this guide depends on, so the one view most worth having is the one a synthetic panel is least able to produce.
Use synthetic respondents to pretest wording and to check that an instrument is comprehensible. Do not use them to produce a score, and do not use them to fill a segment where real response volume is thin, which is precisely the case where the subgroup findings say they are least reliable.
How Sprig supports a customer satisfaction survey
Sprig is an enterprise survey platform powered by AI agents, and several documented capabilities map onto the requirements this guide sets out.
Behavioral event triggering fires a transactional instrument at the moment the interaction completes, which is the timing a transactional read depends on. The Rating Scale question type supports the five-point and seven-point forms, and it allows all points to be labeled, which is what a program reporting distributions needs.
The recontact waiting period provides a documented control on the survey fatigue that typically degrades a high-volume satisfaction program, and it operates across studies rather than within one.
Open-Text AI Analysis themes the low-score probe and exports the themes alongside the responses, which is the input the driver analysis needs. AI Follow Ups can probe a low score while the respondent is still present, which is generally the highest-value recovery available on a four-question instrument.
What Sprig does not do for this method
Sprig documents no statistical significance testing anywhere in the product. Every interval and every period comparison in this guide happens outside it, in a spreadsheet or an AI client.
Sprig also documents no in-product trend chart and no cross-study numeric comparison, which matters for a metric whose whole value is its movement. There is no documented respondent identity key persisting across separate studies, so following the same customer's satisfaction across waves is not an in-product operation. No minimum response count and no accuracy figure are published for theming, and no accuracy or reproducibility claim is published for AI-generated follow-ups, so treat both as a first pass to be checked.
Response export is CSV only, capped at 950,000 rows, with a download link that expires after one week. For a program of this size that is rarely a constraint, and it is worth knowing before anyone builds a recurring pipeline against it.
The Sprig MCP connector is the alternative to an export for the analysis in this guide, passing studies and responses directly into a supported AI client. Responses are retrieved in pages, so a larger dataset takes more than one call.
Alternatives and adjacent methods
Four measures sit near this one, and each answers something satisfaction does not.
Customer Effort Score measures perceived ease on a transactional interaction instead of an evaluative judgment. The evidence generally supports it as a supplementary diagnostic beside satisfaction and not as a standalone loyalty measure, and it is the natural companion where the interaction is a problem the customer came to you to solve.
Net Promoter Score measures stated advocacy at the relationship level. Morgan and Rego found satisfaction outperformed it at predicting firm performance, it typically reports quarterly rather than per interaction, and it carries a licensing regime that satisfaction does not.
A key-driver study is the natural companion rather than an alternative. It uses the satisfaction item as the dependent variable and identifies which attributes move it, which is the step most programs skip and the one that converts a level into a plan.
Observed retention and repeat purchase, taken from the customer record, are the outcome measures satisfaction is a weak proxy for. Where you can measure the behavior directly, measure the behavior and use satisfaction to explain it rather than to predict it.
The general rule across all four is that satisfaction answers what people think and the adjacent measures answer what they did or how hard it was. A program that runs satisfaction alone is choosing to know only the first of those.
Task-based usability research is the better spend when the object of study is one screen rather than a whole service, because it answers why a specific flow is hard rather than how people feel about the service.
Where the question is about a whole market rather than your own customers, a brand or category tracking study is the instrument, and its satisfaction reading is not comparable to a score drawn from your own customer base.
Frequently asked questions
What is a good CSAT score?
There is no sourced cross-industry customer satisfaction benchmark comparable to a score from your own survey, so no external number qualifies as good.
Scale length, question wording, sampling and channel all shift the result, which means two organizations' figures are rarely on the same instrument.
Set your own first three months as a baseline and report movement against it with the margin of error attached.
How is CSAT calculated?
Customer satisfaction is most commonly calculated as the top-two-box percentage, which is the count of respondents selecting 4 or 5 on a five-point scale divided by the total count of valid responses, multiplied by 100.
Exclude blanks from the denominator rather than counting them as zeros. Report the arithmetic mean and the share selecting 5 alongside the percentage, because the box score hides movement that both of those capture.
Is a CSAT score the same as an ACSI score?
No. The American Customer Satisfaction Index is a three-item latent construct estimated with a proprietary and patented partial least squares model and transformed onto a 0 to 100 scale, drawn from a sample the index draws itself.
A customer satisfaction score is typically a single self-administered question reported as a mean or a top-two-box percentage.
An index score of 78 and a satisfaction score of 78 percent are not comparable numbers and should never appear in the same table.
How many responses do I need for CSAT?
No methodology source publishes a satisfaction-specific sample-size rule, so derive the figure from the precision you need. At 200 responses a top-two-box near 80 percent carries a margin of error of roughly 5.5 points at 95 percent confidence, at 400 roughly 3.9 points, and at 800 roughly 2.8 points.
Comparing two periods needs substantially more, because the interval on a difference is wider than the interval on either figure.
Should CSAT use a 5-point or 7-point scale?
Either is defensible and the choice matters far less than freezing it. Scale research finds reliability and validity improve up to about seven categories and that two-point to four-point scales perform poorly, so five and seven both sit in acceptable territory while the three-point widget common in helpdesk tools does not.
Changing scale length mid-program shifts the level of the score by itself, so pick one and treat any change as a series break.
What is the difference between transactional and relationship CSAT?
A transactional study fires after a specific event, measures that event, and describes the quality of a process. A relationship study fires on a schedule, measures the accumulated impression of the product or company, and describes brand standing with existing customers.
They draw different populations and produce different numbers, so run them as separate studies and never average them together.
Does CSAT predict whether customers will stay?
Weakly. Satisfaction alone explains approximately 8 percent of the variance in loyalty, and multi-predictor retention models in a two-year panel of over 8,000 customers reached adjusted R-squared of 7.3, 10.3 and 7.5 percent across three industries.
A meta-analysis of 251 correlations puts the average satisfaction to performance relationship at r equals .101, rising to .349 under favorable conditions.
Use observed retention data where you have it, and treat satisfaction as a leading indicator rather than a forecast.
Do I need a license to use CSAT?
No. No one owns the term or the standard question wordings. The acronym is registered, serial 78151839 and registration 2827719, by the International Institute for Trauma and Addiction Professionals, for training courses in the field of sex addiction.
Unrelated field, no constraint on using it to mean customer satisfaction.
The one caution is the American Customer Satisfaction Index, whose name and patented estimation method belong to its operator, so describe and cite it freely but never call your own three-item battery an index score.
Bottom line
Customer satisfaction is the most extensively studied measure in this family and a weak predictor of behavior, and most programs get into trouble by treating it as the second thing.
Separate the transactional read from the relationship read. Freeze the scale. Report the mean, the top-two-box and the 5-share together, because the collapsed figure hides the distinction the evidence says matters most.
Size from the precision you need, and compare movement against the interval on a difference before calling it a trend.
If your goal is to know whether the judgment customers hold about you is moving, this is the right instrument and the measurement literature behind it runs back to Cardozo in 1965.
If your goal is to predict who will renew, the evidence points at observed retention and at a multi-indicator model.
Your next study is a key-driver study on the attributes behind the score, fielded on a frozen instrument to one clearly defined population and reported with its distribution and its interval. That is the step that turns a level into a priority list, and it is the one most programs never take. The customer satisfaction template is the concrete artifact to start from. Ignore any sample-size figure printed on that template or any other, and size from the precision table above, which is derived rather than inherited.