Running an NPS survey takes six steps: decide whether you are measuring the relationship or a single interaction, field the standard zero-to-ten recommendation question with one open-text follow-up, pick a delivery channel and hold it constant, collect enough responses that your margin of error is smaller than the change you want to detect, calculate the score as the percentage of promoters minus the percentage of detractors, and read the verbatim responses to learn why it moved. The metric answers one question well: whether loyalty is trending up or down in a population you can survey repeatedly. At 100 responses it carries a margin of error of roughly 16 points.
The default this guide recommends is a relational survey fielded quarterly, delivered in-product, with event-triggered surveys kept separate and separately reported.
What Net Promoter Score measures
Net Promoter Score℠, commonly called NPS®, measures how willing a population of customers says it is, at one moment, to recommend a company, product, or service to someone else.
Where the metric came from
Fred Reichheld introduced it in "The One Number You Need to Grow", published in Harvard Business Review in December 2003, reporting that the share of customers enthusiastic enough to refer a company correlated with relative growth rates inside an industry.
The metric was developed with Bain and Company and Satmetrix Systems, which NICE acquired in 2017. That is why current trademark filings now name NICE instead of Satmetrix.
The bands and the calculation
Promoters score 9 or 10. Passives score 7 or 8. Detractors score anywhere from 0 through 6, which is why the band is often described as over-broad.
The score is the percentage of promoters minus the percentage of detractors. Passives are excluded from the arithmetic entirely, though they remain in the denominator.
NPS = (Promoters / Total x 100) - (Detractors / Total x 100)
The result runs from negative 100 to positive 100. It is an index rather than a percentage, and writing it with a percent sign is a frequently made error.
What the published standard leaves out
Bain's published pages define the question, the bands and the formula. They say nothing about cadence, minimum sample size, administration mode, or what score counts as good.
Every rule below is therefore derived, not sanctioned, and each is typically stated with its derivation so you can argue with the reasoning instead of the authority.
What the score does not capture
Stated intent and actual recommending diverge, so the score generally reports what respondents say they would do instead of counting referrals that happened. It also does not measure why anyone gave the score they gave.
When to use NPS
Use NPS when the question is directional and recurring, when you can reach substantially the same population every period through the same channel, and when your response volume supports a margin of error narrower than the movement you intend to act on.
Four conditions that make it the right instrument
The first is comparability across periods. The recommendation question has the rare property of meaning roughly the same thing to a researcher and to a board member.
The second is drift instead of diagnosis, which is what a tracking instrument typically detects well.
The third is the open-text prize. A zero-to-ten tap costs a respondent almost nothing and typically earns you the sentence that follows.
The fourth is volume. Your responses per period have to deliver precision narrower than the change you plan to act on.
What kind of evidence this produces
NPS produces directional evidence about a population rather than causal evidence about a mechanism. It cannot typically tell you that a change you shipped caused the movement.
Do not run it because a board deck has a slot for it. That is often the reason a program exists, and it produces measurement nobody in the room actually trusts.
When not to use NPS
Six situations call for a different instrument. In each case the better instrument is named, because a limitation with no alternative attached is a complaint instead of advice.
Three of these are the ones that most often get ignored in practice: the diagnostic question, the account-level question, and the low-volume case.
A program that runs NPS anyway in any of the three produces a number, and the number will be defended in a meeting where nobody can say what it rules out.
When you need to know why something is failing
NPS will not tell you why users abandon a specific step.
The survey asks about the relationship in general, arrives at a moment unconnected to the failure, and reaches whoever happened to be present, so the responses commonly describe an average rather than an event.
The better instrument is a targeted in-product study triggered at that step, asked of the users who abandoned it. A study of 40 people who just failed at something typically beats 400 recalling how they feel about you overall.
When you need to know which account is at risk
A single contact's rating is generally not account health. In business-to-business software the respondent is frequently not the buyer, the administrator, or the person who signs the renewal, so an administrator's 9 and a finance leader's 4 are counted identically.
The better instrument is account-level aggregation with revenue weighting and a segment cut by role, read alongside usage and support signals. None of those three inputs comes from the survey, which is the point.
When your response volume is low
Below a certain volume the index often returns a number that cannot be read. The floor is not a fixed figure, and no methodology source publishes one, so the sampling chapter derives it rather than asserting it.
For a team collecting 60 responses a quarter, the better instrument is the mean likelihood-to-recommend reported with a confidence interval, with the index withheld.
Sauro made that argument at MeasuringU in July 2014, and Dawes reached the same conclusion from different evidence in 2024.
When you need to compare across channels or regions
The metric is frequently not stable across administration modes or across languages, for reasons the fielding chapter quantifies. Two numbers collected differently are not comparable, and no amount of methodological care after the fact repairs that.
The better instrument is a single third-party administration under one mode and one language, which removes the mode effect by design rather than by adjustment.
Where you cannot pay for that, compare each region only against its own prior period and never against another region.
When you need to measure a single interaction
A post-support score typically answers a narrower question than a relational score, and the two are not interchangeable. Blending them is how a support backlog ends up looking like a loyalty problem.
The better instruments are Customer Satisfaction (CSAT) and Customer Effort Score, asked at the interaction. A narrower question collects less noise and needs a smaller sample to answer.
When you need to know what people will pay
NPS says nothing about willingness to pay, and teams frequently reach for it during pricing work because it is the tracking metric already in place.
The better instruments are Gabor-Granger and Van Westendorp, and for tradeoffs across bundled features, conjoint analysis. All three have launched analysis guides.
Study design and the relational fork
Decide first whether you are measuring the relationship or a transaction, because that single fork determines cadence, sampling frame, channel, and what your number can be compared against.
The fork, and the recommendation
A relational survey asks the whole population at fixed intervals, deliberately away from moments of active purchasing or support. A transactional survey fires after a specific event, commonly a support interaction or an onboarding milestone.
The two produce structurally non-comparable numbers, because a transactional score is conditioned on the respondent having just had an interaction. Blending them into one trendline generally produces a number that means nothing.
Recommendation: run the relational survey as your tracked instrument on a quarterly cadence, delivered in-product, and keep transactional NPS as a separate metric with a separate name, owner and report.
Every figure in this guide, including the precision table and every benchmark comparison, assumes the quarterly relational read.
Four decisions that cannot be corrected later
- Question wording
- Scale construction
- Delivery channel
- Survey language
Each is an instrument change and not a settings change, so the periods either side of it are two different measurements sharing a chart.
The design fork, rendered
Q: Do you need a number you can compare across quarters?
NO -> You do not need NPS. Run a targeted study at the moment in question.
YES -> continue
Q: Can you reach substantially the same population every period,
through the same channel, in the same language?
NO -> Fix the frame first, or report per-region against itself only.
YES -> continue
Q: Does your volume per period give a margin of error
narrower than the movement you would act on?
NO -> Report mean likelihood-to-recommend with an interval.
Withhold the index.
YES -> continue
Q: Is the question about the relationship, or about one interaction?
INTERACTION -> Run CSAT or Customer Effort Score at that moment.
Keep it out of the relational trendline.
RELATIONSHIP -> Run relational NPS, quarterly, one fixed channel.
The third branch is the one most programs typically discover in their third quarter rather than before their first.
Writing the instrument
Use one of Bain's two published wordings, ask exactly one open-text follow-up, and then leave both alone permanently.
The canonical question
The longer of Bain's two published wordings, from The Ultimate Question 2.0, reads:
On a zero-to-ten scale, how likely is it that you would recommend us (or this product/service/brand) to a friend or colleague?
The shorter version, from Bain's own measurement page, reads:
How likely are you to recommend us to a friend or colleague?
What you may change, and what you may not
Substituting a product name for the company name is the one safe variation, provided it is consistent. Asking about "Acme" in one quarter and "the Acme dashboard" in the next typically produces two different measurements.
Rather than treating a rewording as an improvement, treat it as a decision to restart the trendline.
The follow-up question
Reichheld's canonical follow-up, from The Ultimate Question 2.0, is a single open-ended question:
What is the primary reason for your score?
His rationale is that an open question lets a company hear reasons in the customer's own words, avoiding the distortions of preconceived response categories.
Ask exactly one follow-up and make it optional. Forcing a comment commonly produces filler text that pollutes the thematic analysis later.
Band-conditional follow-ups often break comparability of the verbatim corpus and leak the scoring model to the respondent. Where a team insists on branching, keep the neutral question first and add the conditional probe second.
Run this in Sprig AI Follow-Ups are documented as available on the NPS question type, generating one contextual follow-up per response based on what the respondent just wrote. The limitation: the feature skips automatically when it cannot generate a safe and relevant question quickly, so coverage is not guaranteed. Follow-up text is a non-random and incomplete subset of your sample, and counting it as though it covered everyone overstates whatever it contains.
The instrument spec, ready to lift
Q1. Type: 0-10 rating scale, 11 points
Required: yes
Text: How likely are you to recommend [COMPANY] to a friend or colleague?
Left anchor (0): Not at all likely
Right anchor (10): Extremely likely
Point labels: endpoints only
Randomization: none
Position: first question, nothing in front of it
Q2. Type: Open text, single field
Required: no
Text: What is the primary reason for your score?
Placeholder: none
Character limit: none
Branching: none
Attach silently as attributes, do not ask:
plan or account tier | account revenue band | tenure in days
role or persona | region | survey language | delivery channel
Do not add:
a third question | a satisfaction battery in front of Q1
band-conditional follow-ups as the only follow-up
a "how can we improve" variant in place of Q2
Scale and scoring choices
Use eleven points, zero through ten, labeled at both ends only, with the canonical anchors "Not at all likely" at zero and "Extremely likely" at ten.
Why eleven points and not ten
Never present the scale as one through ten. A one-to-ten variant silently removes an entire detractor position while leaving the bands nominally intact, and it is a frequently shipped error.
Label the endpoints only, and label them the same way in every period.
Whether to move the band cutoffs
Recommendation: keep Bain's bands. Switching cutoffs forfeits every external comparison and restarts your trendline exactly as a rewording would.
Rather than changing the split, report the full distribution and the mean alongside the score, and treat the alternative cutoffs described in the critique chapter as a sensitivity check.
What to report beside the index
Report four numbers every period rather than one: the index, the three band shares, the mean likelihood-to-recommend, and the response count.
The mean uses all eleven scale points and is what significance tests should run against.
Sample size and precision
Your sample size should be set by the size of the change you intend to detect, not by a round number and not by whatever the response rate happens to deliver.
The precision table
The variance of an NPS estimate depends on how responses distribute across the three bands.
For a sample split 40 percent promoters, 30 percent passives and 30 percent detractors, producing a score of 10, the 95 percent margin of error runs as follows.
| Responses | Promoters | Passives | Detractors | Reported score | Margin of error, 95 percent |
|:---------:|:---------:|:--------:|:----------:|:--------------:|:---------------------------:|
| 10 | 4 | 3 | 3 | 10 | plus or minus 51.5 points |
| 100 | 40 | 30 | 30 | 10 | plus or minus 16.3 points |
| 400 | 160 | 120 | 120 | 10 | plus or minus 8.1 points |
| 1,000 | 400 | 300 | 300 | 10 | plus or minus 5.1 points |
This table is derived, not sourced. It applies the published variance formula to one stated distribution, and the arithmetic below lets you rerun it against your own split.
At 100 responses a score of 10 carries an interval running from roughly negative 6 to positive 26. Most teams typically report quarterly movements of 3 to 8 points off samples in the low hundreds and treat them as real.
They are noise.
The arithmetic, shown
Promoters count as 1, passives as 0 and detractors as negative 1. The formulas, as published by Genroe:
Var(NPS) = (1 - NPS)^2 x (P/T) + (0 - NPS)^2 x (Pa/T) + (-1 - NPS)^2 x (D/T)
Standard Error = sqrt(Var(NPS)) / sqrt(T)
Margin of Error (95%) = sqrt(Var(NPS)) / sqrt(T) x 1.96 x 100
Worked at T equals 100: variance 0.690, square root 0.831, divided by the square root of 100 gives 0.083, multiplied by 1.96 and then by 100 returns 16.3 NPS points.
Setting your own floor
The floor is the response count at which your margin of error falls below the movement you intend to act on. No methodology source publishes a universal minimum, so the derivation is the rule rather than a number.
Work it backward. A team acting on a 10-point shift needs a margin of error under 10 points, which the table places nearer 400 responses than 100. Detecting a 5-point shift takes roughly 1,000.
A volume of 150 a quarter detects movements of about 13 points or larger, and that expectation generally belongs with stakeholders before the first report rather than after the third.
Why the table understates your uncertainty
This interval covers sampling variation alone. It does not cover the mode effect quantified in the fielding chapter, the language effect described there, or non-response bias.
Read it as the least uncertainty you have rather than all of it, which is also the argument for holding channel and language constant.
How large a population you need
Refiner's in-app dataset, updated November 13 2025 and covering 1,382 surveys and 50.3 million views, puts NPS at a 21.71 percent response rate. On that rate, 400 responses requires roughly 1,800 eligible users reached inside the window.
Email performs worse. Writing for Bain in January 2021, Ilker Carikcioglu noted that response rates have been declining and that any given survey now might have response rates in the low single digits.
Where that arithmetic says your reachable population cannot support an index at a usable interval, that is the finding and not an obstacle to route around.
Run this in Sprig Audience sampling lets you field to a fixed number of responses or to a continuous daily rate, with confidence presets from 80 to 98 percent and custom sampling from 1 to 100 percent. The limitation: response-based quotas appear in the changelog as coming soon rather than as shipped, so balancing segment sizes while responses arrive is typically a manual watch rather than an automatic control.
Audience, targeting and screening
Define the sampling frame before the first send, because the frame is the population you can actually reach and it is rarely the population you want.
What each frame excludes
An in-product survey reaches only users who signed in during the window, which oversamples engaged users and excludes the dormant accounts closest to churning.
An email survey reaches the full contact list including dormant users, and is commonly answered disproportionately by people with strong feelings in either direction.
Neither frame is neutral, so name the bias you accepted rather than claiming representativeness.
Attributes to attach without asking
Attach plan tier, revenue band, tenure, role, region, language and channel as targeting attributes rather than as questions. Every attribute you ask generally costs response rate and introduces self-report error.
Screening and weighting
Screen only when a wrong respondent would corrupt the number instead of diluting it. For a relational read among your own customers the frame is the screen.
Weight when your respondent mix differs materially from your population mix on a dimension correlated with the score, which is most often account size.
Publish the scheme alongside the number, because a weighted score with an undocumented scheme is not reproducible by whoever inherits the program.
Fielding and delivery
Choose one delivery channel, document it, and change it only at a deliberate re-baseline. Channel choice often affects the score more than a product change does.
The mode effect, quantified
Fred Van Bennekom of Great Brook Consulting reported a case in which an identical survey was administered two ways at a business-to-business technology company handling 750,000 service transactions a year.
Telephone administration produced 58.1. The web form produced 18.0. The gap was 40.1 points on the same instrument, at a p-value of 5.09 x 10 to the negative 16, which rules out sampling error.
The mechanism he describes is scale truncation. Phone respondents hear only the endpoints and typically gravitate toward them, while web forms display every point visually.
Treat this as what it is: a single-company consulting case study from 2014, unnamed client, raw data unpublished, not peer-reviewed. It remains the clearest available evidence on mode effects and does not need overstating.
What each channel biases toward
In-product surveys reach signed-in active users during the window, which is the most engaged sample available to you and the least representative of the accounts closest to churning.
Email reaches the whole contact list including dormant accounts, and commonly skews toward whoever is the named contact rather than the person using the product.
Research panels generally reach people outside your customer base entirely, which makes a panel the wrong tool for a relational score and the right tool for a competitive benchmark you administer yourself.
The channel decision rule
Use in-product delivery when your product has regular authenticated usage and the population you care about signs in. Accept that you are typically measuring active users and say so in the report.
Use email when a material share of your customer relationships sit with people who rarely or never sign in, which is common in business-to-business software sold to a buying committee.
Never compare scores across channels, and never change channel mid-program without restarting the baseline. Any external benchmark that mixes channels is unusable for the same reason.
Run this in Sprig Delivery covers in-product surveys on web and native mobile, native email, shareable links, QR codes and research panels. Sprig publishes up to a 2x improvement in completion rates when moving from a static form to a conversational format, which is its own figure with no published baseline. The limitation: there is no native short message service delivery, and a link can be sent through your own messaging tool instead. The conversational format is also an instrument change, so adopting it mid-program moves your score for reasons unrelated to loyalty.
Localization moves the score
Anne-Wil Harzing's "Response Styles in Cross-national Survey Research", in the International Journal of Cross Cultural Management in 2006, examined response styles across 26 countries. English-language questionnaires elicited higher middle responses, while native-language surveys produced more extreme patterns.
Rolling out multi-language support mid-year will typically move your reported score, and the movement will not be a change in loyalty. Localize at a period boundary and report the two eras separately.
Qualtrics XM Institute's work across 18 countries and 17,509 consumers points the same direction. The per-country adjustment factors are not published, so treat this as direction only and never apply an invented correction.
Timing and cadence
Field a relational survey quarterly at most, enforce a recontact window across every study your company runs, and treat respondent attention as a company-wide budget rather than as your team's alone.
Why quarterly
Quarterly is a ceiling rather than a target. It generally gives most programs enough volume per period to clear a usable margin of error while leaving room for the other studies competing for the same contacts.
Where volume is too low to support quarterly reads, move to twice a year and report a wider interval rather than a more frequent guess.
The contradiction in the standard advice
One widely repeated recommendation is to survey every contact in every account every quarter. Another puts the comfort threshold for an individual respondent at three to four surveys a year in total.
Both circulate freely and they are incompatible. The second comes from vendor-published research with no stated sample or method, so treat it as a convention rather than a measurement.
Cadence and the recontact window are one decision
Your relational survey competes with product-market fit surveys, satisfaction surveys, research recruitment and marketing research, all aimed at the same people.
Choosing quarterly without setting a recontact window means a respondent can receive the relational survey and three other studies in the same month. Decide the allocation once at the company level.
Run this in Sprig A recontact waiting period sets the minimum number of days before a user receives another study, applies account-wide as a default, and can be overridden per study. Published cadence guidance starts at 30 days, with 7 days as the shortest recommended window. The limitation: the shipped default and the configurable range are not stated in the documentation, so confirm what your account is set to rather than assuming a window is already in force.
Quality control and data hygiene
Decide your exclusion rules, your bot controls and your de-identification step before the first response arrives, because retrofitting any of them means your first period is not comparable to the rest.
What to exclude, and what to keep
Exclude blank and non-substantive open-text answers from the theme analysis and report how many you excluded. Keep those respondents in the score denominator, since they answered the rating question.
Exclude nothing from the score on the basis of the score itself. A rule that drops extreme responses simply manufactures the stability it then reports.
Turn on bot and fraud filtering before fielding to any open link and record that you did.
De-identify verbatim responses before they leave your platform, because respondents type names, email addresses and account details into a free-text box without being asked to.
Pre-launch checklist
Run through this before the first send, in this order:
- Confirm the question wording matches the canonical text exactly and is recorded permanently
- Confirm the scale runs zero through ten with eleven points and endpoint labels only
- Confirm the recommendation question sits first with nothing in front of it
- Confirm the open-text follow-up is present and optional
- Confirm the instrument contains no third question
- Attach plan tier, revenue band, tenure, role, region, language and channel as silent attributes
- Set the target response count from the movement you intend to detect
- Set a recontact window and confirm it applies across every other study in flight
- Confirm the delivery channel matches the prior period, or document the re-baseline
- Confirm the survey language matches the prior period for each region
- Turn on bot filtering and record that you did
- Decide who receives detractor responses and within how many hours
- Decide the theme reporting threshold below which a verbatim cluster is not reported
- Confirm the trademark attribution footnote appears on any published output
- Field to a small holdout first and read every verbatim response it returns
The last item is the one most often skipped and the one that most often catches a broken trigger, a truncated scale on mobile, or a follow-up question nobody can answer.
Analysis: the core calculation
Calculate the score as the percentage of promoters minus the percentage of detractors, then attach a confidence interval before anyone reads it.
The calculation, worked
Take a quarter that collected 240 responses: 108 promoters, 66 passives and 66 detractors.
Promoters are 108 divided by 240, or 45 percent. Detractors are 66 divided by 240, or 27.5 percent. Passives are excluded from the subtraction but stay in the denominator.
The score is 45 minus 27.5, which rounds to 18. Report it as 18 and not as 18 percent, since a difference between two percentages is a number of points.
Attaching a confidence interval
For NPS specifically the recommended method is the adjusted Wald interval.
Turk, Cinderich and McNeill published "Coverage and Precision of Net Promoter Score Confidence Intervals Across Sampling Distributions" in Stats in 2026, volume 9, issue 2, article 45, comparing Wald, bootstrap t and adjusted Wald across four population shapes.
They found adjusted Wald with triangular and uniform weights particularly robust, and recommended avoiding Wald and bootstrap t at small samples. The CRAN package NPS implements nps.test for intervals and significance tests.
Testing whether two numbers actually differ
Jim Lewis and Jeff Sauro published a worked significance test for NPS in February 2021, using an adjusted Wald interval converted to a Z-test, comparing two products on samples of 36 and 31.
The gap between them was 33 points, which typically looks decisive in any dashboard.
The standard error of the difference was 0.180, giving a Z of 1.815 and a p-value of 0.07, with a 90 percent interval running from 3 to 62 points.
A 33-point gap did not reach significance at the conventional threshold. If that can fail the test, a 6-point quarterly movement on a few hundred responses is not a finding.
Sauro's earlier analysis, published in July 2014, gives the practical rule: run comparisons on the raw mean and standard deviation, and keep the index as the reporting layer rather than the analytical one.
Benchmarks and what a good result looks like
There is no usable cross-industry NPS benchmark: published figures for software and software-as-a-service diverge by roughly 20 points across sources, and by far more once self-reported company figures are included, depending entirely on who was sampled and who did the sampling.
The published figures, side by side
| Source | Software or software-as-a-service figure | Sample and method |
|:-------------------------------------------------------------------------------:|:-----------------------------------------------------------:|:------------------------------------------------------------------------------------------------------------------------------:|
| Qualtrics XM Institute, XMI Customer Ratings NPS 2024, published September 2025 | Software Firm 21.1 | 10,000 United States consumers, 354 organizations, 22 industries, disclosed method, independent panel |
| Retently, published April 2026 | Business-to-business software and software-as-a-service, 41 | Client-base aggregate. Industries qualify on as few as 11 clients. No per-industry sample sizes, no size or geography controls |
| CustomerGauge, page titled 2025, survey dates undisclosed | Software-as-a-service average, 36 | 38 companies. No per-company sample sizes, dates, channel, or relational-versus-transactional type |
| CustomerGauge, same page | One named company at 92 | Whether the named company scores are self-reported or independently measured is not disclosed |
The independently panelled figure sits roughly 20 points below both vendor client-base aggregates, and that gap is the whole argument.
The 92 comes from the same page as the 36, so it demonstrates inconsistency inside one vendor's own benchmark rather than adding a fourth independent source.
The strongest of these figures is also the oldest. The Qualtrics XM Institute set covers 2024 fieldwork published in September 2025, so it lags the period you are reporting.
What to benchmark against instead
Recommendation: benchmark against your own prior period with a stated confidence interval, and treat any external figure as context rather than as a target.
Bain's own commercial answer to "what is a good score" is a paid, double-blind panel averaging over 20,000 respondents per industry. A paid panel of that size is a reasonable thing to sell only if self-reported benchmarks are not usable.
Interpreting and acting on the result
Decide who receives a detractor response and within how many hours before the first one arrives, because retrofitting routing means the first cohort of detractors went unanswered.
Reading the number honestly
Report the interval next to the score every time. A movement inside the margin of error is generally not a result, and the reliable defense against reading it as one is having the interval in the same sentence.
Read the distribution alongside the index. A 30 built from 40 percent promoters and 10 percent detractors is a different business from a 30 built from 65 percent promoters and 35 percent detractors.
Segmentation, and the tradeoff nobody states
Segment by plan tier, tenure, role, journey stage, region, survey language and delivery channel. The last two are mandatory and not optional, because both typically move the score independent of sentiment.
Segmentation is required for the metric to be actionable and it simultaneously destroys the sample sizes that make it readable.
A quarter with 400 responses carries a margin of error near 8 points, and cutting it four ways into segments of 100 gives each roughly 16.
Recommendation: pick the two cuts that would change a decision, size the sample for those, and report every other dimension as a qualitative pattern in the verbatim responses. Where a segment falls below your floor, publish the response count and the themes and withhold the index.
For an account-level view, take the mean of the raw scores inside each account rather than computing an index per account, and report the revenue-weighted and unweighted aggregates side by side.
Closing the loop, justified without vendor statistics
Most published figures on the value of closing the loop come from vendor research with no stated sample sizes, controls or methodology. Two arguments stand without them.
The first is obligation. You asked someone for their time and their candid opinion, which creates a debt a silent dashboard does not discharge.
The second is measurement. A respondent who never hears back is often less likely to answer next time, which shrinks the denominator every subsequent period depends on.
What closing the loop requires, and where Sprig does not help
Closing the loop requires a named owner, a response inside a stated window, a record of what was said, and a way to see whether the underlying issue was fixed.
Sprig publishes no case management, no response-level assignment and no resolution tracking, and the Slack integration notifies on study events rather than on response content. Alerting an owner the moment a detractor responds is therefore not a published capability.
This is the real structural gap against the four Leaders in the 2026 Gartner® Magic Quadrant™ for Voice of the Customer Platforms, where case management is table stakes.
For a program whose value depends on closing the loop, that workflow belongs in a platform built for it, or in your existing support and customer success tooling with responses routed there.
The bias closing the loop does not fix
Respondents self-select toward strong feeling in one direction or the other, and increasingly so as survey volume rises.
A score can be arithmetically precise and still misleading when the people who answered are systematically happier or angrier than the people who did not.
Closing the loop improves your future response rate and does not make the current sample representative. Where a neutral read matters enough to pay for, the better instrument is a third-party administration the respondent does not associate with you.
Running the analysis with Claude or ChatGPT
The score tells you the direction. The verbatim responses tell you why, and reading 240 by hand is the step most programs quietly skip.
What the analysis produces
You end up with two artifacts and one ranking.
The first is a theme table: five to eight recurring reasons drawn from the verbatim responses, each with a respondent count, a share of the sample, the mean score of the people in that theme, and quotes you can check yourself.
The second is the score arithmetic: the index, the three band shares, the mean likelihood-to-recommend and a confidence interval.
The ranking is typically what you act on. Themes sorted by how far their mean sits below the overall mean are your candidates, because a theme mentioned by many people who still scored you a 9 is not a problem.
What you do not get is a cause.
Prompt one: themes with a score attached
Attaching the score to each theme is the instruction teams forget, and it is what turns a list of topics into a ranked list of problems.
What this prompt does: groups NPS verbatim responses into themes created
from the data, attaches the mean score of the respondents in each theme,
and tells you which themes are large enough to act on.
What it returns: a theme table with counts, shares, mean scores and
verbatim quotes; a list of themes held below a reporting threshold; a
ranking by score gap; and a saved file of the table.
Data: [EXPORT FILE, default nps-export.csv] with columns response_id,
score, verbatim, plan_tier, tenure_days, role, region, language, channel.
1. Read every non-blank verbatim response and group them into [THEME COUNT,
default 5 to 8] themes created from the data, not from a predefined
list. Give each theme a label of five words or fewer and a one-sentence
description. Put anything that does not fit into an "Other or unclear"
bucket rather than forcing it into a theme.
2. Use code to calculate this, not estimation.
3. Exclude blank, whitespace-only and non-substantive verbatim responses
from every theme calculation rather than counting them as zero. Report
how many you excluded. Keep those respondents in the score denominator.
4. Flag any theme with fewer than [REPORTING THRESHOLD, default 5 percent
of non-blank responses] respondents as NOT REPORTABLE and place those
rows below a divider. A theme with ZERO respondents also counts as
under the threshold. Do not delete either kind.
5. Recompute every count, share and mean using a genuinely different
method than the one you used first, for example a manual row-by-row
tally against a formula. Re-running the same calculation does not count.
6. If a recomputed number does not match, report both numbers and say
which you trust and why. Do not silently correct it.
7. Output a table with theme, respondent count, share of non-blank
responses, mean score of respondents in that theme, and 2 to 3 verbatim
quotes copied exactly. Save it as [OUTPUT FILE, default nps-themes.csv]
and give me the download link.
8. State explicitly what these responses cannot tell me, and name two
follow-up studies that would answer it.
Step 8 is the one people delete, and it is the one that keeps the analysis honest.
Prompt two: the index, the interval and the period comparison
The interval and the significance test are the parts teams skip when they compute the score in a spreadsheet, which is the argument for running the arithmetic here.
What this prompt does: computes the NPS index, the band shares, the mean
likelihood-to-recommend, a confidence interval, and a test of whether this
period differs from the last one.
What it returns: a single statistics block, a stated smallest detectable
change, and a saved file of every number so the next period can be
compared against it.
Data: [EXPORT FILE, default nps-export.csv], same columns as above.
Prior period: [PRIOR PERIOD FILE, default none, in which case skip step 5].
1. Use code to calculate this, not estimation.
2. Count promoters (9 to 10), passives (7 to 8) and detractors (0 to 6),
with counts and percentages. Exclude blank and non-numeric score values
from the denominator rather than treating them as zero, and report how
many you excluded.
3. Report the index as promoter percentage minus detractor percentage, on
a scale of -100 to 100, not as a percentage. Report the mean
likelihood-to-recommend and its standard deviation across all eleven
points.
4. Compute a 95 percent confidence interval on the index using the
adjusted Wald method and state it in NPS points.
5. If a prior period file is attached, test whether the two periods differ.
Run the comparison on the raw means, report the p-value, and say in one
sentence whether the difference is distinguishable from noise at this
sample size.
6. If any segment in [SEGMENT COLUMN, default plan_tier] has fewer than
[SEGMENT THRESHOLD, default 100] respondents, mark it NOT REPORTABLE and
report its response count instead of its index. A segment with ZERO
respondents also counts as under the threshold.
7. Recompute steps 2 through 4 by a genuinely different method than you
used first, and report any mismatch with both numbers instead of
correcting it silently.
8. State the smallest change in the index this sample size could detect at
95 percent confidence. Save every number above to [OUTPUT FILE, default
nps-stats.csv] and give me the download link.
Step 8 answers the question every stakeholder asks after the first report, and having it in writing beforehand stops a noise movement becoming a roadmap.
Setup and getting your data in
Turn on code execution before running either prompt, or the instruction to calculate rather than estimate has nothing enforcing it.
Export a file carrying the score column, the verbatim column and every segmentation attribute you want to cut by, because adding a column later means rerunning the analysis. Exporting raw responses works anywhere.
Or connect the data live, which removes the export step and the staleness window entirely.
Run this in Sprig The Theming Agent groups open-text responses into named themes written from the underlying responses, shows how many responses support each theme, and lets you open the responses behind any theme. AI Study Reports require at least 10 responses, and Sprig MCP connects studies, responses and generated themes into Claude or ChatGPT, capped at 1,000 responses per call. The limitation: no accuracy or reproducibility figure is published for the theme counts, and no significance testing or cross-study trend view exists, so the interval, the period comparison and the theme verification all happen in your own analysis layer. Cross-tabs and significance reasoning run in the connected AI client rather than inside the platform.
Reading the output
Read the table for density before you read it for meaning. On a 240-response quarter split across eight themes, it is common for only three or four to clear a five-percent floor.
A run where everything clears usually means the themes are too broad to act on.
The five-percent default in the prompt above is this guide's convention and not a sourced standard. Set your own and record it, because a threshold chosen after seeing the themes is not a threshold.
Then read the ranking rather than the counts. A theme carried by 40 respondents whose mean is 8.1 describes something people tolerate, while a theme carried by 18 respondents whose mean is 3.4 describes the thing costing you the score.
Two numbers belong beside every theme table: how many responses were excluded as blank, and how many themes were withheld as too small.
What to verify before reporting
Three checks, run by hand, every time:
- Confirm the theme counts sum to your non-blank response total, or the shares are wrong
- Filter the export to one theme and count those rows yourself against the reported count
- Search your data for one quoted verbatim per reportable theme, worded exactly as shown
A quote tidied into cleaner English commonly signals paraphrase where you needed a copy, and a theme count that does not reconcile means responses were double-counted or dropped.
Researchers remain responsible for setting the threshold, validating themes against the underlying responses, and deciding which are decision-relevant. The analysis removes the labor of reading 240 comments and not the judgment about which three matter.
Pitfalls
Six failure modes, each with what to do about it.
Discovering themes is harder than applying them. Hill and colleagues, in PLOS Digital Health in April 2026, found models matched human analysts on deductive coding against an existing codebook, 93.5 percent against 92.7, and were materially worse at inductive discovery, with a 1.2 percent strict fabrication rate and a 12.4 percent comprehensive error rate.
Theme inductively once, rebuild the list yourself, then have the model apply it every period after.
Rare themes get over-predicted, and rare themes are what get escalated. Ashwin, Chhabra and Rao, in Sociological Methods and Research in May 2025, found non-random bias in 10 of 19 codes, with sparse codes systematically over-predicted.
The same prompt does not return the same answer. Thinking Machines Lab reported in September 2025 that 1,000 completions at temperature zero produced 80 unique outputs.
Neither chat client exposes a seed, so a reportable number should generally be produced twice and reconciled.
A large context window is not attention across it. Liu and colleagues, in Transactions of the Association for Computational Linguistics in 2024, found performance degraded when relevant material sat in the middle of the input.
Put the data above the instructions.
Leading questions get agreeable answers. Sharma and colleagues documented sycophancy at the 2024 International Conference on Learning Representations, and OpenAI withdrew a model update in April 2025 for being overly agreeable.
Never ask a model to confirm the theme you already suspect.
Verbatim responses contain personal information nobody asked for. Whether that content can be used for training turns on which tier you paste it into and not on which vendor, so de-identify before upload.
Where this chapter stops
Creating and recoding themes is covered in the guide to coding open-ended responses into themes, and cross-tabbing a score band across segments in the guide to cross-tab analysis.
The critique you should know
The academic and practitioner criticism of NPS is substantial, and a team running the metric should be able to state it before a skeptical researcher does.
The original growth claim did not replicate
Keiningham, Cooil, Andreassen and Aksoy published "A Longitudinal Examination of Net Promoter and Firm Revenue Growth" in the Journal of Marketing in 2007, volume 71, issue 3, pages 39 to 51.
Using longitudinal Norwegian data covering 21 firms and more than 15,500 interviews, and deliberately testing the industries Reichheld cited as success stories, they failed to replicate the claimed superiority of Net Promoter over other measures.
The paper won the 2007 Marketing Science Institute H. Paul Root Award, so this is not a blog rebuttal.
Editorial note for review, remove before publish. Assumption A7. The Morgan and Rego citation and its verbatim quotation were added at draft stage on the strength of the September 11 instruction. Confirm the quotation against the published article before this page goes live.
The satisfaction mean outpredicts it
Morgan and Rego published "The Value of Different Customer Satisfaction and Loyalty Metrics in Predicting Business Performance" in Marketing Science in 2006, volume 25, issue 5, pages 426 to 439.
Their finding, stated verbatim: "average satisfaction scores have the greatest value in predicting future business performance and that Top 2 Box satisfaction scores also have good predictive value. ... metrics based on recommendation intentions (net promoters) and behavior (average number of recommendations) have little or no predictive value."
A top-tier journal saying the satisfaction mean outpredicts the recommendation index is the most inconvenient finding in this literature, and it predates the Keiningham replication failure by a year.
The metric's creator disowned its self-reported form
Reichheld, Darci Darnell and Maureen Burns published "Net Promoter 3.0" in the November to December 2021 Harvard Business Review. The admission is direct: unaudited, self-reported Net Promoter Scores undermined the usefulness of NPS.
Their successor is earned growth rate, splitting revenue growth into what came from returning customers and their referrals against what was bought through advertising and sales.
It requires per-customer revenue attribution plus an acquisition-source question on every new customer, and most product and customer experience teams are not instrumented for either.
The bucketing adds noise rather than only losing information
John Dawes published "The net promoter score: What should managers know?" in the International Journal of Market Research in 2024, volume 66, issues 2 to 3, pages 182 to 198.
The counting method introduces additional variation compared with mean likelihood-to-recommend scores, both across brands and across survey periods. Collapsing eleven scale points into three buckets adds noise to exactly the period-over-period tracking most programs typically exist to produce.
The band cutoffs lack statistical support
Cazzaro and Chiodini published "Statistical validation of critical aspects of the Net Promoter Score" in The TQM Journal in 2023, volume 35, issue 9, pages 191 to 209.
Their central objection is that identical index values can correspond to different levels of loyalty, and that the original cutoff distribution lacks statistical support.
They cite Kristensen and Eskildsen's empirically derived alternative of 0 through 4, 5 through 7, and 8 through 10.
The example that lands hardest is a customer who moves from 4 to 8 and generally produces no movement in the index at all.
The counter-critique
Jaramillo, Deitz, Hansen and Babakus published "Taking the measure of net promoter score" in the International Journal of Market Research in 2024, volume 66, issues 2 to 3, pages 278 to 302, using structural equation modeling with multi-group analysis.
They found NPS corresponds to reported word-of-mouth exposure for most product categories, and that responses are invariant across demographic groupings. Both satisfaction and NPS significantly predicted differences in financial performance, with satisfaction explaining slightly more variance.
The fair reading is that NPS works. It is simply not uniquely good, and a program that treats it as uniquely good will defend a number the literature does not defend for it.
Where this leaves a practitioner
Recommendation: run NPS as a tracking and learning instrument, report the mean likelihood-to-recommend alongside it, and pair both with targeted studies that explain movement. Do not present it as a growth predictor.
Common mistakes
Ten failure modes, each with what it costs and how to prevent it.
Reading movement inside the margin of error as a result
A team reports a 5-point gain on 180 responses and builds a quarterly plan against it. The margin of error at that volume is roughly 12 points, which makes the movement indistinguishable from nothing.
Publish the interval next to the score and the mistake becomes impossible to make quietly.
Blending relational and transactional scores in one trendline
The two instruments reach different populations, so combining them produces a line that moves when support volume changes. Give them separate names, separate owners, and separate reports.
Changing the delivery channel mid-program
An email program moved in-product will move the score for reasons unrelated to loyalty. Channel migrations frequently arrive as an efficiency project, and the measurement consequence is rarely raised in that conversation. Change channels only at a declared re-baseline.
Rewording the question to make it clearer
Any rewording restarts the trendline. Teams typically do this in good faith, commonly to add a product name, then compare the new number to the old one.
Lock the wording permanently and record it where the next person will find it.
Localizing without re-baselining
Native-language surveys produce more extreme responses than English-language ones, so a localization rollout moves the number independent of sentiment. Ship localization at a period boundary and report the eras separately.
Reporting a segment below its own floor
A 30-response segment produces an interval wider than the range anyone would act within. Dashboards render any segment, and a rendered number often reads as a real one. Set a floor before the first report.
Adding questions in front of the recommendation question
A preceding question changes what the respondent has just been considering, so the instrument no longer matches the one that produced every prior period. Put the recommendation question first, always.
Treating the score as a diagnosis
The index reports that the direction changed and not why. Teams frequently spend a quarter debating causes the instrument never collected. Where the verbatim responses are thin, the answer is a targeted study instead of a longer argument.
Surveying the same contacts more often than the company agreed
A team ships a quarterly program with no recontact window, and a contact receives four studies in five weeks. Set the window across all studies and treat it as a company-level control.
Tying the score to compensation
Reporting in Customer Experience Dive on April 30 2025 cites Forrester research finding that 50 percent of customer experience leaders say their firms tie employee pay to NPS targets.
In the same reporting Reichheld calls the practice fundamentally flawed and says it will destroy the business in 95 percent of cases.
Report the metric, do not compensate on it, and set process goals such as closing the loop on every detractor within 48 hours.
A target of five points inside a noise band of twelve can be hit or missed by chance, which makes it unfair as an accountability mechanism and useless as a signal.
Synthetic respondents
Do not field an NPS survey to synthetic respondents. The instrument's entire value is that real customers answered it on a repeating schedule, and a simulated score has nothing to be compared against.
The deeper problem is validation.
A synthetic concept test can be checked against the real study that follows it, while a synthetic loyalty score has no ground truth to check against until the real survey runs, at which point the simulation was unnecessary.
Where simulation does earn its place is upstream of fielding.
Pressure-testing question wording or checking that a follow-up prompt reads clearly are drafting tasks rather than measurement tasks, and a wrong answer there costs a revision rather than a quarter of reporting.
How Sprig supports this method
Sprig ships NPS as a first-party question type that categorizes responses into promoters, passives and detractors automatically, alongside open-text theme analysis with traceability from a theme back to the responses behind it.
What a recurring program needs from a platform
Nine capabilities determine whether a recurring program runs itself or becomes a quarterly scramble:
- Applies the three bands natively to an eleven-point rating question
- Themes open-text answers with respondent counts and traceability back to responses
- Attaches plan tier, tenure, role and region to every response without asking
- Triggers an interaction study at the journey stage the score points toward
- Caps a fielding window at a defined number of responses
- Enforces one recontact window across every study in the account
- Delivers through more than one channel, since non-users cannot be reached in-product
- Exports raw responses with their attributes for your own analysis layer
- Routes a single response to whoever owns following up on it
The platform covers the first eight. It does not publish the ninth, which is typically where a closed-loop program breaks.
Where this platform is the weaker fit
There is no published trend view carrying one score across studies over time, so period-over-period reporting means exporting responses into your own reporting layer and rebuilding the view there.
For a tracking program, that is the reporting layer, and it is absent.
Native email delivery runs through Enterprise link surveys, carries a 24-hour per-recipient cooldown, and cannot operate alongside panel distribution on the same survey.
No benchmark database is published, and no native short message service, offline, contact-centre or telephony delivery exists. That last one matters here because the largest documented score difference in this guide is the one between telephone and web administration.
Sprig also does not appear in the 2026 Gartner® Magic Quadrant™ for Voice of the Customer Platforms, published 9 March 2026, whose Leaders are Qualtrics, Medallia, Sprinklr and Press Ganey Forsta.
A voice-of-customer buying committee will generally notice, and conceding it costs less than omitting it.
Editorial note for review, remove before publish. Assumptions A4 and A6. The 2026 Magic Quadrant is verified: Gartner document 7534085, published 9 March 2026, twelve vendors, Sprig absent. The four named Leaders are the correct comparison set. InMoment does not appear in the 2026 report under its own name and has deliberately been kept out of any sentence citing the Magic Quadrant. Counsel sign-off on the Gartner attribution and disclaimer below is still pending and is a pre-publish blocker.
Capability claims about this platform were verified against its published documentation on September 11 2026. The platform's G2 product-page rating was 4.3 out of 5 across 199 reviews as of August 17 2026.
Worth reading before committing: the same documentation set carries an argument against tracking this metric at all.
Next step: where the tracked score has pointed at a specific journey stage, the study to run is a targeted in-product survey at that step, asked of the users who dropped out of it.
Alternatives and adjacent methods
Pair NPS with three other instruments rather than replacing it, because each answers a question the index cannot.
What to run alongside it
Read the mean likelihood-to-recommend beside the index, since it uses every scale point and is what significance tests should run against.
Add Customer Satisfaction or Customer Effort Score at the two or three interactions carrying the most weight.
Add event-triggered studies at whichever journey stage the index and the verbatim themes both point toward. That is the study this guide has been building toward, and it is the one that changes a roadmap.
When depth interviews are the better instrument
Where the question is why a specific experience fails rather than whether sentiment moved, moderated or asynchronous depth interviews answer it from a handful of participants.
Depth interviews are a different category answering a different question. They produce rich causal detail from few people and cannot produce a comparable quarterly index, so they typically sit beside a tracking program rather than inside it.
The publisher's own position on this metric
Run this in Sprig This guide's publisher argues in its own documentation that the metric is too generic and that teams should track journey-stage questions and simple averages instead, and it makes the case for moving past it and argues it is not the right metric for monitoring customer experience. The limitation: that position is a reason to read this guide sceptically as well. A survey vendor publishing against a metric it also sells a question type for has an interest in routing you toward the studies it is strongest at, and the honest reading is that the two positions coexist rather than resolve.
This guide takes a narrower line than the documentation does: run the metric well, then instrument beyond it.
Frequently asked questions
What is Net Promoter Score?
Net Promoter Score is an index measuring how willing customers say they are to recommend a company or product. Respondents answer one zero-to-ten question. Promoters score 9 or 10, passives 7 or 8, detractors 0 through 6. The score is the promoter percentage minus the detractor percentage, running from negative 100 to positive 100.
How do you calculate NPS?
NPS is the promoter percentage minus the detractor percentage. Divide promoters by total responses, do the same for detractors, then subtract. Passives scoring 7 or 8 stay in the denominator but not the subtraction. For 240 responses with 108 promoters and 66 detractors: 45 percent minus 27.5 percent gives 18.
What is a good NPS score?
There is no usable cross-industry NPS benchmark. Published software figures run from 21.1 in an independently panelled study of 10,000 consumers to 41 in a vendor client-base aggregate, and one vendor page lists a company at 92 with no disclosed method. Compare your score against your own prior period with a stated confidence interval.
How many responses do you need for an NPS survey?
The number of responses an NPS survey needs is set by the movement you intend to detect. At 100 responses the 95 percent margin of error is roughly 16 points, and at 1,000 it falls to about 5. A team acting on a 10-point shift needs nearer 400 responses per period. A team collecting 60 should report a mean with an interval instead.
How often should you run an NPS survey?
Quarterly is a reasonable ceiling for a relational NPS survey, and twice a year suits low-volume populations better. Pair the cadence with a recontact window applied across every study your company runs, since the same contacts also receive recruitment, satisfaction surveys and product feedback requests.
What is the best NPS question wording?
Use the canonical wording and never change it. Bain publishes "How likely are you to recommend us to a friend or colleague?" alongside a longer variant naming the product or brand. Substituting your company or product name is safe if done consistently. Any other rewording restarts your trendline and makes prior periods incomparable.
Should you send NPS surveys by email or in your product?
In-product delivery reaches signed-in active users and generally achieves higher response rates. Email reaches dormant accounts and contacts who never sign in, which is the churn-risk population an in-product survey misses. Pick the one matching the population you need, then hold it constant.
What is the difference between relational and transactional NPS?
Relational NPS asks the whole customer population at fixed intervals about the overall relationship. Transactional NPS fires after a specific event such as a support ticket or an onboarding milestone. A transactional score is conditioned on the respondent having just had an interaction, so the two should never share a trendline.
How do you know whether a change in your NPS is real?
Calculate a confidence interval using the adjusted Wald method and run significance tests on the raw mean instead of on the index. In one published worked example, a 33-point gap between two products at samples of 36 and 31 produced a p-value of 0.07 and missed significance. A 6-point quarterly movement on a few hundred responses is generally noise.
How should you segment NPS results?
Segment by plan tier, tenure, role, journey stage, region, survey language and delivery channel. The last two are mandatory because both change scores independent of sentiment. Every cut widens the margin of error on the resulting segments, so pick the two cuts that would change a decision and report the rest as verbatim themes.
What should you do with detractors?
Assign an owner and a response window before the first survey goes out. Contact them, record what was said, and track whether the issue was fixed. Most survey platforms do not provide case management, so name the support or customer success system that owns the workflow.
What are the main criticisms of NPS?
Five criticisms of NPS hold up. The original growth claim failed to replicate in an award-winning 2007 Journal of Marketing study. A 2006 Marketing Science paper found the satisfaction mean outpredicts recommendation intention. The band cutoffs lack statistical support. Bucketing eleven points into three adds noise to period tracking. And the metric's creator wrote in 2021 that unaudited self-reported scores had undermined its usefulness.
Does NPS predict revenue growth?
NPS does not reliably predict revenue growth. The claim that it does comes from Reichheld's 2003 Harvard Business Review article, and a 2007 Journal of Marketing study covering 21 firms and more than 15,500 interviews failed to replicate it. A 2006 Marketing Science paper found average satisfaction scores predict future business performance better than recommendation intention. Treat NPS as a tracking instrument and not as a growth forecast.
Should NPS be tied to bonuses or compensation?
No, and the metric's creator is the strongest voice against it. Reichheld calls the practice fundamentally flawed and says it will destroy the business in 95 percent of cases, allowing one narrow exception for senior executives measured on relative competitive performance. Forrester research nonetheless finds around 50 percent of customer experience leaders tie pay to it.
Do you need permission to use NPS?
Net Promoter, NPS and related marks are registered trademarks and service marks owned jointly by Bain and Company, NICE Systems and Fred Reichheld. Bain's published terms encourage non-commercial use of the methods and state the marks may not be used commercially without a license. Confirm your own position with counsel before publishing commercially.
Is NPS still worth measuring?
Net Promoter Score is still worth measuring for most teams, provided it runs as one instrument rather than as the answer. It gives you a comparable direction across periods, a low-friction prompt that collects open-text explanation, and a number stakeholders already understand. The failure mode is treating a directional index with a wide margin of error as a precise diagnosis.
Bottom line
Net Promoter Score is worth running as a tracking instrument and is not worth treating as an answer.
Four decisions determine whether your program produces a usable signal. Lock the question wording permanently. Fix the delivery channel and change it only at a declared re-baseline.
Size the sample from the movement you intend to detect. Publish a confidence interval next to the score in every report.
The failure mode most teams hit is not a bad question or a low response rate.
It is typically acting on movement inside the margin of error, quarter after quarter, because the interval was never calculated and the noise looked like information.
Run the metric if you have a population you can reach repeatedly and a decision a directional signal would inform.
Do not run it if your volume cannot support an interval narrower than the change you care about, or if the question you actually have is why something specific is failing.
Write down what you will do differently if the score falls, and what you will do differently if it rises, before you launch instead of after.
The next study to run is not another tracking survey. It is the one aimed at whichever journey stage your score and your verbatim themes are both pointing at, asked of the people who experienced it, while they still remember.
Net Promoter®, NPS®, NPS Prism®, and the NPS-related emoticons are registered trademarks of Bain & Company, Inc., NICE Systems, Inc., and Fred Reichheld. Net Promoter Score℠ and Net Promoter System℠ are service marks of Bain & Company, Inc., NICE Systems, Inc., and Fred Reichheld.
GARTNER is a registered trademark and service mark, and MAGIC QUADRANT is a registered trademark, of Gartner, Inc. and its affiliates in the United States and internationally. Gartner does not endorse any vendor, product or service depicted in its research publications, and does not advise technology users to select only those vendors with the highest ratings or other designation. Gartner research publications consist of the opinions of Gartner's research organization and should not be construed as statements of fact. Gartner disclaims all warranties, expressed or implied, with respect to this research, including any warranties of merchantability or fitness for a particular purpose.