Introduction
Running a Customer Effort Score survey takes five decisions in order. Pick the 2.0 statement instead of the 1.0 question, because the two run in opposite directions. Field it on a seven-point scale immediately after the interaction. Ask a resolution question and a satisfaction question alongside it. Score it as the percentage selecting 5, 6, or 7. Then size the sample from the precision you need, because no sourced benchmark exists.
Most programs report effort alone. That is why the score is so often presented as a loyalty forecast, when the strongest test of it against observed retention found it performed worst of the five measures compared.
At 100 responses a top-three-box near 70 percent carries a margin of error of roughly 8.9 points.
This guide covers:
- Choose between the 1.0 question and the 2.0 statement first, because the polarity reverses
- Pair every effort item with a resolution item and a satisfaction item
- Derive the sample from the precision you need instead of copying a threshold
- Compare periods against the interval on a difference, not on a single figure
- Read the verbatims, because the score locates friction and never explains it
Platform capabilities and documentation cited here were verified against sprig.com and docs.sprig.com on September 17, 2026. No prices appear anywhere in this guide.
What Customer Effort Score measures
Customer Effort Score measures perceived ease. It captures how hard a customer felt they had to work, not whether the work succeeded.
The metric originates with Matthew Dixon, Karen Freeman, and Nicholas Toman in "Stop Trying to Delight Your Customers," published in Harvard Business Review in July 2010.
The research behind it came from the Corporate Executive Board and covered more than 75,000 people, plus hundreds of structured interviews with customer service leaders.
The argument was contrarian at the time. Rather than exceeding expectations, the authors found that service organizations create loyalty primarily by removing obstacles.
Their supporting figures, as printed in that article: 94 percent of customers reporting low effort expressed an intention to repurchase, and 88 percent said they would increase their spending.
Among customers who had a hard time solving their problems, 81 percent reported an intention to spread negative word of mouth.
The two versions, and why the difference matters
Two distinct questions circulate under the same name, and they run in opposite directions.
Customer Effort Score 1.0, as printed in the 2010 article, asks: "How much effort did you personally have to put forth to handle your request?" It is scored from 1, very low effort, to 5, very high effort. A low score is the good outcome.
Customer Effort Score 2.0 presents a statement instead of a question: "The company made it easy for me to handle my issue." Respondents rate agreement on a seven-point scale, and the score is the percentage selecting the top three boxes. A high score is now the good outcome.
The revision is attributed to Gartner, which acquired the Corporate Executive Board in 2017. Several practitioner sources date it to 2013, the year Dixon, Toman, and Rick DeLisi published The Effortless Experience.
The Gartner research note that would settle the exact publication date sits behind a paywall, so treat the date as approximate and the wording as well attested instead of officially sourced.
A question form, "How easy was it to resolve your issue today?", also circulates widely. Treat the declarative statement as the canonical item and the question form as a common variant.
The two versions side by side
| | Customer Effort Score 1.0 | Customer Effort Score 2.0 |
|:--------------------------:|:------------------------------------------------------------------------------:|:-------------------------------------------------------------------------:|
| Form | Question | Agreement statement |
| Wording | "How much effort did you personally have to put forth to handle your request?" | "The company made it easy for me to handle my issue." |
| Scale | 5 points, 1 very low effort to 5 very high effort | 7 points, Strongly disagree to Strongly agree |
| Direction | Low is good | High is good |
| Standard scoring | Mean, or share selecting 1 or 2 | Share selecting 5, 6, or 7 |
| Source | Dixon, Freeman and Toman, Harvard Business Review, July 2010 | Attributed to Gartner, commonly dated to 2013, primary document paywalled |
| Contains the word "effort" | Yes | No |
The last row is the one teams miss. The 2.0 item never mentions effort, which means a customer answering it is rating ease of handling rather than estimating how hard they worked.
Those are related judgments and they are not the same judgment, and the change was deliberate.
The difference between Customer Effort Score 1.0 and 2.0 is a reversal, not a refinement. Version 1.0 asks how much effort the customer expended, on a five-point scale where a low score is good. Version 2.0 asks whether the company made things easy, on a seven-point scale where a high score is good. A team that migrates between them without noticing has inverted its own trend line, and a falling number now means an improving experience.
A widely repeated statistic that is not where people say it is
One figure appears on most vendor pages covering this metric: 96 percent of customers with a high-effort service interaction become more disloyal, against 9 percent who have a low-effort experience.
It is commonly attributed to the 2010 Harvard Business Review article. A full-text search of that article for both figures returns nothing. The numbers come from the later Corporate Executive Board work published in The Effortless Experience in 2013.
Attribute it to the book instead of the article. A separate claim in circulation, that Customer Effort Score is 1.8 times more predictive than satisfaction and twice as predictive as Net Promoter Score, could not be traced to any retrievable primary source and should not be used.
When to use Customer Effort Score
Customer Effort Score earns its place in transactional, problem-resolution contexts where the customer arrived with a task and either completed it or did not.
Strong fits:
- Support ticket resolution
- Self-service help center tasks
- Account changes and cancellations
- Onboarding and setup steps
- Returns, claims, and billing disputes
- Password and access recovery
The common thread is that the customer had a goal, the goal was utilitarian, and friction is unambiguously unwanted. Those three conditions are what make a low score actionable, because each one points at work an operations team can actually do.
Timing is what makes the measurement work. Fire the question immediately after the interaction, while the experience is still available to the respondent, instead of burying it in a monthly relationship survey where they are reconstructing an event from memory.
Where the metric pays for itself
Customer Effort Score is often the right instrument when your improvement work is operational instead of strategic.
Channel switching, repeat contacts, and knowledge gaps are all effort-generating, and a metric that moves when you fix them is more useful to a support organization than one that does not.
That is a reason to run both rather than a reason to run effort alone.
The metric also works well as a comparison across channels serving the same task. Chat, phone, and self-service handling the same password reset should produce comparable effort scores, and a large gap between them identifies a channel that needs work instead of a customer segment that is harder to please.
One more context deserves mention. The verbatims attached to low effort scores are specific and quotable in a way aggregate satisfaction data rarely is, which makes them useful to a support organization that has historically argued for investment on anecdote.
When a satisfaction measure serves better
Run satisfaction instead of effort when the experience is evaluative instead of transactional. A customer rating a product, a relationship, or a purchase decision is making a judgment that effort does not capture.
Run both when the interaction is transactional and the outcome matters commercially. That combination is what the evidence supports, and the reason appears in the next section.
When not to use Customer Effort Score
Five things Customer Effort Score cannot tell you. Each names the instrument that answers the question instead.
It cannot tell you whether the customer actually accomplished anything
Customer Effort Score measures perceived ease, not outcome. A customer can rate an interaction easy and still leave with the problem unresolved, and a support organization optimizing on effort alone can improve the score by making failure faster.
This is the failure mode most likely to damage a program, because the metric moves in the right direction while the experience gets worse.
The instrument that answers this is a binary resolution measure taken from system data instead of from the customer, paired with a stated "was your issue resolved?" item in the same survey.
The XM Institute treats task success and effort as separate components of its experience rating for this reason.
It cannot tell you whether the customer will stay
This is the load-bearing limitation and the evidence is unambiguous.
Evert de Haan, Peter Verhoef, and Thorsten Wiesel tested five customer feedback metrics against observed retention in "The predictive ability of different customer feedback metrics for retention," published in the International Journal of Research in Marketing in 2015.
The design drew 6,649 respondents producing 8,924 firm evaluations across 93 firms in 18 industries, then followed up two years later. Retention was measured as observed behavior rather than stated intent.
Their correlations with two-year retention: top-two-box satisfaction .184, Net Promoter Score .170, mean satisfaction .151, and Customer Effort Score negative .073. In the multivariate model, Customer Effort Score was the only metric in the set that did not reach statistical significance.
Their conclusion, stated directly: "the CES in itself has little to no predictive power and performs the worst of all CFMs studied."
One scope limit belongs with that finding. The study operationalized the 1.0 effort question, not the 2.0 ease statement, so the result speaks to the original instrument.
The instrument that answers this is observed retention or renewal from the customer record. Where an attitudinal leading indicator is wanted, top-two-box satisfaction had the strongest association with two-year retention in that same study.
It cannot tell you where in the journey the effort occurred
A single post-interaction rating collapses hold time, channel switching, repeat contacts, and agent knowledge into one number. The score tells you the episode was hard. It does not tell you which part was hard.
The instrument that answers this is channel-switch and repeat-contact counts taken from instrumentation, plus a task-based usability study on the specific flow.
For a digital self-service step, the System Usability Scale or a per-task single ease question measures the step rather than the whole episode.
It cannot tell you whether effort should be minimized here at all
The premise that less effort always produces more satisfaction does not survive contact with every category.
Caroline Ardelet and Christophe Benavent analyzed 314,194 customer-brand interactions across 96 brands, collected weekly over 102 weeks, in "Does making less effort entail satisfaction?", published in the International Journal of Market Research in 2022.
Effort strongly affected satisfaction overall, with a Cohen's f of .401 across all sectors pooled.
But in hedonic sectors including e-commerce, automotive, retail food, and do-it-yourself, they found satisfaction is highest when effort is either low or intense, and lowest when effort is moderate. At point of sale, satisfaction increased with intense effort.
In utilitarian sectors including banking, insurance, internet service, and mobile, effort reduced satisfaction consistently across channels. Their conclusion is that the relationship depends on the brand, the channel, and the level of effort required.
The instrument that answers this is a segmented satisfaction study modeling effort as a curvilinear rather than linear predictor, or qualitative in-context research on what the effort meant to the customer.
It cannot tell you anything about customers who never contacted you
Customer Effort Score is a post-contact instrument by construction. Customers who tried, failed silently, and left are invisible to it, and they are the population the program most needs to understand.
The instrument that answers this is funnel instrumentation on the self-service path, plus exit or churn interviews with customers who never opened a ticket.
Study design, and the three decisions that carry it
Three decisions determine whether the resulting number means anything: which version of the question, which scale, and what you pair it with.
The recommendation
Run Customer Effort Score 2.0, the seven-point agreement statement, scored top-three-box, alongside a satisfaction item and a resolution item. Field it immediately after the interaction.
That configuration follows from the evidence instead of from convention, and each element answers a specific limitation named above.
Why the seven-point scale
Scale length is not arbitrary, and the measurement literature is unusually clear about the range that works.
Carolyn Preston and Andrew Colman tested scales from two to eleven points plus a 101-point scale across 149 respondents, in "Optimal number of response categories in rating scales," published in Acta Psychologica in 2000.
Their finding: two-point, three-point, and four-point scales performed relatively poorly on reliability, validity, and discriminating power, indices improved up to about seven categories, and test-retest reliability tended to decrease beyond ten.
That result indicts the three-point good-neutral-bad widget commonly built into helpdesk tooling. A three-point scale sits squarely in the band their study found performs worst, and teams adopt it because it is the default in the tool rather than because anyone chose it.
The same result supports the seven-point form the 2.0 item already uses. Their indices improved up to about seven categories and did not improve further, with test-retest reliability declining past ten. That plateau, not a peak, is what makes seven defensible.
Why changing scale length breaks your trend
John Dawes randomly assigned respondents to five-point, seven-point, and ten-point versions of the same items in a study published in the International Journal of Market Research in 2008.
After rescaling to a common metric, the ten-point format produced an aggregate mean 0.3 points lower, at p equals 0.04, while distributional shape did not differ significantly.
The shape of your data is robust to scale length. The level is not.
Freeze the instrument. A team that switches scale length mid-program has introduced a step change that has nothing to do with customer experience, and no amount of analysis afterward can separate the two.
This matters more than it appears to, because scale changes are often made for good local reasons. A new tool defaults to five points, a redesign shortens the survey, a regional team standardizes on a different form.
Each change is defensible on its own and each one breaks the series.
Whether to label every point
The answer depends on what you will do with the number, and there is a clean rule.
Bert Weijters, Elke Cabooter, and Niels Schillewaert manipulated labeling and category count across 1,207 respondents in a study published in the International Journal of Research in Marketing in 2010.
Fully labeling all categories increases acquiescence and decreases both extreme responding and misresponse. Endpoint-only labeling shows better criterion validity in regression models.
Their recommendation for modeling work is a five-point or seven-point scale with endpoint labels only.
So label every point when you will report the response distribution to stakeholders, because a fully labeled scale is what makes a distribution chart interpretable to a non-specialist audience. Label endpoints only when the score will feed a model.
Pick one and hold it. The choice matters less than the consistency, and teams change labeling during a redesign without registering it as a measurement change.
Whether to include a neutral midpoint
A seven-point agreement scale has a midpoint by construction, and that midpoint is doing less work than it appears to.
Patrick Sturgis, Caroline Roberts, and Patten Smith probed respondents who selected the middle alternative across a sample of 3,113, in work published in Sociological Methods and Research in 2014.
Between 69 and 84 percent of middle-alternative responses turned out to be hidden "don't knows" instead of genuine neutrality.
They still recommend keeping the midpoint, because a substantial minority do hold genuinely neutral positions.
One caveat belongs with that finding. Those were low-salience attitude items, not post-transaction ratings where the respondent has just had the experience. The generalization to effort measurement is an inference and the guide treats it as one.
The practical consequence is a reporting habit. Track the midpoint share alongside the score rather than folding it into the distribution, because a midpoint that moves while the tails do not is a signal about who is answering, and the cause is worth investigating rather than assuming.
What to pair it with
Pair every effort item with a resolution item and a satisfaction item. This is the single highest-value design decision in the guide.
The resolution item separates ease from outcome, which closes the first limitation above. The satisfaction item is what the de Haan evidence says makes effort useful at all.
Three questions is also close to the practical ceiling for a post-interaction survey. Each additional question costs responses, and on a post-interaction survey the marginal question is usually worth less than the responses it costs.
Writing the instrument
Three questions is the whole instrument. Each one is there because a specific limitation named above requires it.
The core item
Present the statement exactly as written: "The company made it easy for me to handle my issue."
Substitute your own organization name for "the company" and adjust "issue" to match the interaction type. A billing dispute is an issue. A password reset is more naturally a request, and an onboarding step is a task.
Do not rewrite the statement into a question, and do not add qualifiers. Every modification moves you further from the only version of this item with published research behind it, and teams make several small changes that collectively produce a new instrument nobody has validated.
Scale: seven points, Strongly disagree to Strongly agree. Score as the percentage selecting 5, 6, or 7.
The resolution item
Ask it as a separate closed question instead of folding it into the effort item: "Was your issue resolved?" with Yes, No, and Not yet as the options.
Three options rather than two matters here. A customer whose ticket is open but progressing is neither a success nor a failure, and forcing them into one distorts both groups.
The "Not yet" share is also a useful operational signal in its own right. A rising proportion indicates cases are being closed in the system before they are closed for the customer.
The satisfaction item
Use a five-point or seven-point satisfaction item and hold it constant across the program. Which of the two you choose matters less than never changing it.
Place satisfaction after effort. Asking about ease first and satisfaction second is preferable to the reverse, because an overall judgment stated first tends to anchor the more specific rating that follows.
The open-text probe
One probe, conditional on a low effort score: "What made this harder than it needed to be?"
That wording is deliberate. It presupposes the difficulty the respondent has already reported instead of asking them to justify it, and it asks for a cause rather than a feeling.
Compare it to the common alternative, "Do you have any additional feedback?", which returns a mix of praise, unrelated product requests, and blank responses. A targeted probe returns codeable causes, and the difference in usable yield is often substantial.
Add a second probe on high scores in the first month of a new program, "What made this easy?", then retire it. Positive verbatims are useful for establishing what to protect and they stop earning their place quickly.
Fields to attach instead of asking
Attach these from your own systems instead of spending respondent attention on them: channel, interaction type, handle time, number of contacts on the same issue, whether the customer switched channels, and agent or team identifier.
Every one of these is a diagnostic dimension you will want at analysis time, and every one is already in your systems. A survey that asks customers what your ticketing system knows is wasting the only three questions you get.
Attach a stable interaction identifier as well. That single field is what later allows the survey response to be joined to the operational record, and adding it afterward is impossible.
What to leave out
Leave out demographics, relationship tenure, and product holdings. All of them are attachable, none of them belongs in a three-question post-interaction survey, and each one measurably reduces completion.
Leave out a second open-text field. Two free-text boxes produce one answered box and one blank, and the blank is usually the more specific question.
Scale and scoring choices
Two scoring conventions circulate, and the guide recommends reporting both.
Top-three-box
The standard convention for the seven-point item is the percentage selecting 5, 6, or 7, expressed as a whole number. It is simple, it is what most published discussion assumes, and it communicates easily to an audience that will not read a distribution.
Its weakness is the weakness of every box score. Collapsing seven categories into two discards the distribution, and a program where 6s are converting into 7s looks identical to one where nothing is happening.
Top-two-box is also occasionally used. It is defensible and it is not the convention, so if you adopt it, label every chart explicitly, because a reader who assumes three boxes will misread your number as a decline.
The mean
Report the arithmetic mean alongside the box score, to one decimal place.
The two numbers answer different questions. The box score tells you what share of customers had a low-effort experience. The mean tells you where the distribution sits, and it frequently moves when the box score does not.
A rising top-three-box with a falling mean typically means the distribution is polarizing instead of improving. That pattern is common when a fix helps most customers and breaks the experience badly for a minority, and a program reporting only the box score will not see it.
The reverse pattern is also informative. A flat box score with a rising mean indicates customers already inside the top three boxes are moving upward, which is real improvement that box scoring is structurally blind to.
Reporting the distribution
Publish the full seven-point distribution at least quarterly, even when monthly reporting uses the two summary figures.
The distribution is where a bimodal pattern becomes visible, and bimodality in effort data usually means two different experiences are being averaged into one number.
That is a segmentation failure rather than a measurement failure, and it points directly at the cut you should be making.
What to do about the polarity trap
If your organization previously ran the 1.0 question, do not convert historical scores into 2.0 equivalents. The scales differ in length, direction, and wording, and no published conversion exists.
Start a new series, keep the old one visible with a clear label, and state the changeover date on the chart. A vertical rule on the chart at the changeover date is enough to stop the comparison being made accidentally.
Sample size and precision
No methodology source publishes a Customer Effort Score sample-size rule with a stated derivation. Not the Corporate Executive Board, not Gartner's public materials, and not any peer-reviewed paper located for this guide.
That absence is worth saying plainly, because the figures circulating in vendor content have nothing behind them. A number presented without a derivation is a convention someone repeated, and conventions do not tell you whether your own result is precise enough to act on.
What you can do instead is derive the sample from the precision you need.
The arithmetic for a box score
Top-three-box is a proportion, so the interval around it is a binomial proportion interval. The defensible published method is the adjusted Wald interval described by Alan Agresti and Brent Coull in "Approximate is Better than 'Exact' for Interval Estimation of Binomial Proportions," published in The American Statistician in 1998.
It outperforms both the standard Wald interval and the exact Clopper-Pearson interval at small sample sizes.
The adjustment is simple enough to do by hand. Add two successes and two failures to the observed counts, then compute the ordinary interval on the adjusted proportion.
Margins of error at 95 percent confidence, derived using that method, for a top-three-box result near 70 percent:
| Responses per period | Margin of error |
|:--------------------:|:-------------------------:|
| 50 | plus or minus 12.4 points |
| 100 | plus or minus 8.9 points |
| 200 | plus or minus 6.3 points |
| 400 | plus or minus 4.5 points |
| 800 | plus or minus 3.2 points |
One row worked in full, so you can reproduce the rest. At 50 responses with 35 in the top three boxes, add two successes and two failures to get 37 of 54, an adjusted proportion of 0.685.
The margin is 1.96 times the square root of 0.685 times 0.315 divided by 54, which is 0.124, or 12.4 points.
The table is derived rather than cited, and the derivation is the asset. No source publishes a rule, so the honest alternative to an invented threshold is showing the reader the arithmetic and letting them set their own floor.
Note that the interval widens as the score approaches 50 percent and narrows at the extremes. A program scoring near 90 percent has more precision at the same sample size than the table above suggests, and one scoring near 50 percent has less.
What the numbers mean in practice
Read the table as a constraint on what you can claim, not as a target.
A movement between two periods is a difference between two proportions, and its interval is wider than the interval around either figure. That distinction is where most reporting goes wrong.
Take a move from 70 to 74 percent. At 100 responses per period the margin on the difference is about 12.4 points.
At 400 it is about 6.2 points. At 800 it is about 4.4 points, so a four-point move is still inside it.
Detecting a four-point movement at 80 percent power takes closer to 2,000 responses per period. For most programs the honest conclusion is that a monthly number cannot support the comparison and a quarterly one can.
A program reporting monthly on a few hundred responses per segment is reporting mostly noise, and the table above says how much.
The practical consequence is usually a reporting-cadence decision rather than a sampling decision. A team that cannot reach the volume in a month can often reach it in a quarter, and a quarterly number you can defend is worth more than a monthly one you cannot.
Below a few hundred responses per period the score stops supporting comparison at all. At that volume the honest move is to stop reporting a trend and run task-based qualitative research on the flow instead.
For the mean
Use a t-interval driven by the observed standard deviation. A seven-point agreement item produces a standard deviation between 1.3 and 1.8 in practice, and you should compute your own rather than assuming a figure.
The mean is the more sensitive of the two statistics at a given sample size, because it uses the full distribution instead of a binary split. Where sample is tight, the mean will often detect a movement the box score cannot.
Sizing for a comparison
When the question is whether two channels differ instead of what the score is, size for the difference rather than the level. A comparison needs substantially more sample than a point estimate, because two intervals have to separate rather than one interval having to be narrow.
Decide the smallest difference worth detecting before fielding. A team that decides afterward will choose the threshold that makes the observed result significant, which is not a decision at all.
Audience, targeting, and screening
The population for a Customer Effort Score survey is defined by an event instead of by a customer attribute. Everyone who had the interaction is eligible, and nobody else is.
There is no screener to write. The whole design problem is which qualifying interactions you invite and which you suppress.
Sample every interaction or a fraction
Sample a fixed random fraction of qualifying interactions instead of surveying everyone. A program that surveys every ticket trains customers to ignore the survey, and the response rate decay is difficult to reverse.
Set the fraction from the sample-size arithmetic above rather than from a round number. If 400 responses per month gives you the precision you need and your response rate is 15 percent, you need roughly 2,700 invitations.
Support queues are ordered by escalation, shift, and skill routing, so most non-random selection rules you can think of are correlated with effort.
A fraction taken as "every fifth ticket" is not random when ticket order correlates with anything, and in support queues it frequently does.
Who to exclude
Exclude customers who have already received a survey within your recontact window. Exclude interactions the customer did not initiate, such as proactive outreach, unless you are deliberately measuring those separately.
Do not exclude unresolved tickets. Those are frequently the highest-effort experiences in the dataset, and removing them is how a program produces a healthy score alongside an unhealthy operation.
Do not exclude short interactions either. A two-minute contact that should never have been necessary is often a higher-effort experience than a twenty-minute one that resolved something genuinely complex.
Screening
Most Customer Effort Score programs need no screening question, because eligibility is established by the event that triggered the invitation.
The exception is a program running across multiple interaction types from a single invitation. There, one screening question establishing which interaction the respondent is rating is necessary, and it should be the first question.
The non-response problem
Post-interaction surveys are typically answered by a self-selected minority, and that minority commonly skews toward the extremes of the experience.
There is no clean fix. What you can do is track response rate by segment as a diagnostic in its own right, and treat a falling response rate as a signal instead of an inconvenience.
A segment whose response rate is falling while its score is rising deserves particular attention. The most frequent explanation is that dissatisfied customers in that segment have stopped answering, which makes the improvement an artifact.
Fielding and delivery
In-product measurement reaches the customer who completed the task. Email reaches the one willing to come back for it, which is a different population.
In the product or the interface
For a self-service task, the strongest measurement fires inside the interface immediately after the task completes. The respondent has not left, the experience is unmediated by recall, and the response rate is higher than email.
This is also the only channel where you can reliably attach behavioral context to the response, because the survey and the behavior happen in the same system.
Email
For support interactions resolved by an agent, email commonly remains the practical channel. Send within hours instead of days.
Keep the first question in the email body where the channel supports it. Every click between the customer and the first answer costs response rate, and the effect is often larger than teams expect.
Avoid sending from a no-reply address. Customers frequently reply to these surveys with the detail that explains their score, and a program that discards those replies is throwing away its best verbatims.
Other channels
Short message service works for interactions that happened by phone, and it produces fast responses and very short verbatims. Treat the open-text probe as optional there.
In-channel widgets inside a chat or messaging session reach the customer while they are still in the conversation, which is why response rates there tend to be high.
Expect a positive skew for the same reason, because the agent is often still present.
What to avoid
Avoid measuring effort in a periodic relationship survey. A customer asked in March about a February ticket is reconstructing an event, and the accuracy of that reconstruction is not something the survey can check.
Avoid mixing channels within a reported segment without checking the mix. Channel differences are frequently larger than the operational differences a team is trying to measure, which is why an unchecked channel mix can swamp the finding.
Running this in Sprig
Sprig fires in-product surveys from behavioral events, which is the correct trigger for a self-service effort measurement, and the recontact waiting period provides a documented fatigue control across a multi-touchpoint program. The Rating Scale question type supports the seven-point form the 2.0 item requires.
The limitation worth knowing before you build: the recontact window is enforced only in production environments, so testing in a development environment will over-survey. Sprig also publishes no range specification for the Rating Scale in its documentation, so confirm the configuration in the product instead of relying on a documented maximum.
Timing and cadence
Fire the survey as close to the interaction as the channel allows. For in-product measurement that means immediately. For email, within the same business day.
The reason is measurement quality rather than convenience. Effort is a judgment about an experience, and the judgment typically degrades as the experience recedes.
One deliberate exception
Delay the invitation where resolution is not immediate. A ticket marked resolved but awaiting customer confirmation is often not resolved, and surveying at the moment of internal closure measures the agent's view of the outcome instead of the customer's.
Waiting 24 hours in those cases generally produces a more honest score, and the tradeoff against recall decay is usually worth it.
How often to report
Report monthly for most programs, and weekly only where volume supports the precision. A weekly number on a few dozen responses is typically a random walk presented as a trend.
Set the reporting cadence from the sample-size table instead of from the meeting calendar. This is frequently the single easiest improvement available to an existing program.
How often to survey the same customer
Set a recontact window and hold it. Ninety days is a common starting point, and the right number depends on your interaction volume per customer.
A customer who contacts support four times in a month should generally be surveyed once. Which of the four is a decision worth making deliberately, and sampling the first rather than the most recent typically avoids selecting for unresolved cases.
Enforce the window in the tooling rather than in a policy document. Recontact rules that live in a runbook are commonly broken by the next person who builds a survey.
Quality control and data hygiene
Four checks, run before any number leaves the analysis.
Straight-lining and speed
Flag responses completed faster than a plausible reading time for the instrument, which is commonly a sign of a straight-lined submission. For a three-question survey, anything under about eight seconds generally deserves scrutiny.
Flag but do not automatically delete. Report how many you flagged and what happened to the score with and without them, because a deletion rule applied silently is indistinguishable from a thumb on the scale.
Non-response handling
Decide the rule before you look at the data: responses missing the effort item are excluded from the effort calculation, and responses missing other items remain in the dataset for their own questions.
State the rule in the report. An analysis that silently drops partial responses produces a different number than one that does not, and nobody reading the chart can tell which they are looking at.
Channel and interaction-type balance
Check the mix of channels and interaction types in each period against the previous one. A score that moves because the ticket mix moved is generally not a change in customer experience.
This is commonly the largest source of unexplained movement in a Customer Effort Score program, and it is trivially detectable if anyone looks. A simple table of period-over-period volume share by channel is typically enough.
Duplicate and test responses
Remove internal test submissions before analysis, and check for duplicate responses against the same interaction identifier.
Both are more frequent than teams expect, particularly in the weeks after a survey is rebuilt. Read the impact off the precision table: a handful of test responses barely moves a figure built on 800 and visibly moves one built on 50.
Analysis: the core calculation
Top-three-box is a division. Almost everything that goes wrong goes wrong in the denominator or the polarity.
Computing the score
Top-three-box equals the count of responses selecting 5, 6, or 7, divided by the count of responses to the effort item, multiplied by 100. Exclude blanks and non-responses from the denominator instead of treating them as zeros.
The mean is the arithmetic mean of the same responses, to one decimal place.
Report both, always, with the response count beside them. A score without its denominator is not interpretable, and a score reported without one is frequently a score nobody checked.
Segmenting
Cut the score by channel, interaction type, and resolution status at minimum. Those three cuts typically answer most of the questions a support organization will ask.
Resolution status is the most informative and the least commonly reported. The effort score among resolved interactions and the effort score among unresolved ones are different numbers measuring different things, and the blended figure hides both.
Add agent or team as a fourth cut with care. It is operationally useful and it changes how the survey is perceived internally, and a program that becomes an individual performance measure will generally start producing managed results.
Comparing periods and segments
Compute the interval around each figure before comparing them. Two proportions whose intervals overlap substantially are not different, and the sample-size table above is what turns that into a rule you can apply.
Sprig does not perform significance testing anywhere in the product, so this comparison typically happens in a spreadsheet, a statistics tool, or an AI client. The next chapter covers the last of those in detail.
Where you compare many segments at once, expect some to appear different by chance. Twenty segments compared at 95 percent confidence will generally produce one apparent difference that is not real, and the fix is to treat multi-segment scans as hypothesis generation rather than as findings.
Linking to operational data
Join the survey response to the interaction record on a ticket or session identifier. Joining on the ticket identifier is what lets you model effort against handle time, contact count, and channel switching, which is where a score becomes a diagnosis.
A Customer Effort Score program that never performs this join will typically report a score for years without ever identifying a cause.
Contact count is worth modeling first. The relationship between effort and a second contact on the same issue is the one most programs find earliest, and it is the one that converts most readily into an operational target.
Benchmarks and what a good result looks like
There is no sourced cross-industry Customer Effort Score benchmark, and the reason is structural instead of accidental.
Customer Effort Score 1.0 is a five-point scale where low is good. Customer Effort Score 2.0 is a seven-point scale where high is good. Scoring conventions in circulation include the mean, top-two-box, top-three-box, and a Net Promoter Score style net calculation.
Two organizations' Customer Effort Score numbers are therefore only rarely on the same instrument. A cross-industry benchmark table for this metric is a category error rather than merely an uncited one.
What the published tables actually are
Several vendors publish industry ranges. The sourcing is consistent across them, and it is thin.
One widely cited table describes its figures as compiled from customer experience research by Gartner, Qualtrics, and industry reports, with no sample size, no collection period, and no named report.
The same page concedes that the most useful benchmark is your own score over time.
Gartner, the organization that owns the research lineage, publishes no public norm table at all. That absence is informative, because a benchmark would be straightforward for them to produce if the instrument supported one.
The closest real thing, and why it is not this
The XM Institute publishes a customer ratings benchmark with a stated sample of 10,000 consumers across 354 companies and 22 industries. It scores three elements on a seven-point scale: success, effort, and emotion, then averages them.
Effort is one of three components of a composite. That benchmark is not a Customer Effort Score benchmark and should not be presented as one, and the exact question wording behind it is not published.
A real distribution that is also not a benchmark
The de Haan study reports effort means across 19 industry categories in the same Dutch sample ranging from 1.0 to 3.3 on the five-point scale.
That is a real, sourced distribution. It is also 2010 Dutch data on the 1.0 instrument drawn from a research sample, which makes it a useful illustration of variance across industries and not a target for anyone.
What to do instead
Benchmark against your own prior periods, on a frozen instrument, within the same channel and interaction type.
Set your first three months as the baseline and report movement against it. That comparison is valid, it is available immediately, and it is generally the only one anyone can defend in a review.
Where an external comparison is genuinely demanded, the honest answer is to explain why one does not exist. That explanation typically lands better than a sourced-looking number that falls apart under a single question.
Interpreting and acting on the result
The score tells you which segment to open. The verbatims attached to it tell you what to fix.
Read the score against resolution
Start every review by putting effort and resolution side by side. Four combinations, and each implies different work.
High ease and high resolution is the target state, and the work is protecting it. High ease and low resolution means you have made failure efficient, which is typically the most dangerous pattern in the dataset.
Low ease and high resolution means the outcome is right and the path is expensive. Low ease and low resolution is an operational failure and needs no further analysis to act on.
The second pattern is the one to look for deliberately, because the headline metric looks healthy while the customer leaves without what they came for. It is also the pattern most likely to be produced by pressure on handling time.
Read movement against the interval
Compare any period-over-period movement to the margin of error from the sample-size table before treating it as a change.
Most reported monthly movement in this metric is inside the interval. A program that reacts to it is chasing noise and will generally burn credibility doing so, because the number moves back and the intervention gets the blame or the credit either way.
Act on the verbatims, not the number
The score typically identifies which segments to investigate. The open-text probe is what tells you why.
Theme the "what made this harder than it needed to be" responses, then cross-tabulate the themes against channel and interaction type.
That cross-tab is the output worth taking to an operations review, and the score itself is the index that points at it.
Rank themes by mean effort score rather than by count. The largest theme is frequently a minor annoyance affecting everyone, and a smaller theme with a much lower mean is often a broken path affecting one segment badly.
What a genuine improvement looks like
Effort generally improves when you remove a step, prevent a contact, or resolve on first touch. It does not improve because the team was told the number matters.
Tie every intervention to a specific friction the verbatims named, then look for movement in that segment instead of in the blended number. A blended score is typically too insensitive to detect a fix to one journey.
Expect the blended number to move slowly even when the work is going well. A fix to one interaction type in one channel commonly moves the overall score by less than a point, which is why segment-level reporting is what keeps an improvement program credible.
Running the analysis with Claude or ChatGPT
The calculation is simple enough to do in a spreadsheet. What an AI client adds is the theming of open-text responses and the cross-tabulation of those themes against segments, which is where the diagnostic value sits.
This chapter covers what is specific to Customer Effort Score. The general mechanics of theming open-ended responses and cross-tabbing them are covered in depth in the cross-tab analysis guides and are not repeated here.
What the analysis produces
Four outputs: the top-three-box score and mean by period, the same figures split by channel and resolution status, a themed list of friction causes from the open-text probe, and a count of responses per theme with the effort score attached to each.
That last output is the one worth the effort. A theme list without the effort score attached tells you what customers said. A theme list with the score attached tells you which friction is costing you the most.
Getting your data in
Export responses as a CSV. Include the effort item, the resolution item, the satisfaction item, the open-text probe, and every attached operational field: channel, interaction type, contact count, and the interaction timestamp.
Confirm before you start that the effort column contains the numeric response rather than the label text. A column of "Strongly agree" strings requires a mapping step, and a model asked to infer that mapping will sometimes get the direction wrong.
If your survey data lives in Sprig, the Sprig MCP connector passes studies and responses directly into a supported AI client, which skips the export. Responses are retrieved in pages, so a larger dataset requires more than one call.
Prompt one: the score, computed
What this prompt does: computes Customer Effort Score and mean, overall and by
segment, from a survey export.
What it returns: a markdown table of top-three-box percentage, mean, and response
count, by period and by segment, plus a flagged list of any segment below the
reporting threshold.
Use code to calculate this, not estimation. Write and run Python. Do not compute
any percentage or mean by reading values.
Data: [ATTACHED CSV]
Effort column: [COLUMN NAME]
Scale: 7-point, 1 = Strongly disagree, 7 = Strongly agree. Higher is better.
Segment columns: [CHANNEL, INTERACTION TYPE, RESOLUTION STATUS]
Period column: [DATE COLUMN], group by [MONTH]
Rules:
1. Exclude blanks, nulls, and non-numeric values from both numerator and
denominator. Do not treat a blank as a zero. Report how many you excluded.
2. Top-three-box = count of 5, 6, or 7 divided by count of valid responses,
times 100, to one decimal place.
3. Mean = arithmetic mean of valid responses, to one decimal place.
4. Report the response count beside every figure.
5. Any segment with fewer than [30] valid responses is reported as
"insufficient responses" rather than as a number. A segment with zero valid
responses is also below this threshold and must be reported the same way,
not omitted from the table.
6. After computing, recompute the overall top-three-box by a different method:
sum the per-segment numerators and denominators and divide. Compare to the
direct calculation. If the two disagree by more than 0.1 points, report both
figures and state that they disagree. Do not silently reconcile them.
7. Save the output as ces-scores.md.
The sixth rule is the one most often missing from a first draft, and it is the one that catches the errors that matter.
A model that computes a figure twice by genuinely different routes will surface a mistake that a model asked to "double-check its work" will not.
The fifth rule exists because of a documented failure pattern. A model asked for a threshold will sometimes count thin cells and empty cells separately, then report only the first as the total.
Neither count is wrong. Combining them is. Naming zero explicitly as below the threshold is what prevents it.
Prompt two: themes with effort attached
What this prompt does: codes open-text responses about friction into themes and
attaches the effort score for each theme.
What it returns: a theme table with response counts, mean effort score per theme,
and three verbatim examples per theme.
Use code to calculate this, not estimation. Write and run Python for every count
and every mean. Assign themes by reading, then compute the arithmetic with code.
Data: [ATTACHED CSV]
Open-text column: [COLUMN NAME]
Effort column: [COLUMN NAME], 7-point, higher is better
Rules:
1. Exclude blanks, single characters, and responses of fewer than three words
from theming. Report how many you excluded.
2. Build the theme list from the data rather than from a list I supply. Include
an "Other or unclear" bucket and report its size.
3. Assign each response to exactly one theme.
4. For each theme report: response count, mean effort score, and three verbatim
quotes. Do not paraphrase the quotes.
5. Any theme with fewer than [15] responses is reported as "low count, directional
only" rather than as a finding. A theme with zero responses is below this
threshold and must be reported the same way.
6. Recompute each theme's mean effort by a second route: group the raw rows by
your theme assignment and recalculate. If any theme's two means differ, report
the discrepancy rather than correcting it silently.
7. If the "Other or unclear" bucket exceeds 15 percent of responses, say so and
recommend rebuilding the theme list rather than proceeding.
8. Save the output as ces-themes.md.
Reading the output
Read the theme table by effort score rather than by count. The largest theme is not necessarily the most expensive one, and a small theme with a very low mean effort score is often a broken path affecting a specific segment.
Check the "Other or unclear" bucket first. A large unclear bucket means the theme list does not fit the data, and every downstream number inherits that problem.
What to verify before reporting
Four checks, every time.
- Confirm the excluded-response counts are plausible against the raw file
- Confirm the theme counts sum to the total minus exclusions
- Confirm the polarity is right by reading three verbatims from the highest-scoring theme
- Confirm any segment you are about to report clears the response threshold
- Confirm the resolution rate moved the same direction as effort, because an effort gain against a resolution loss means you made failure faster
The third check catches the most embarrassing possible error. If the verbatims attached to your highest effort scores describe frustration, the scale direction has been inverted somewhere in the pipeline.
Pitfalls
Six failure modes, all documented.
Discovering themes is harder than applying them. Work published in PLOS Digital Health in 2026 by Hill and colleagues found models matched human analysts on deductive coding against an existing codebook, at 93.5 percent against 92.7, and were materially worse at inductive discovery.
Strict fabrication ran at 1.2 percent and comprehensive error at 12.4 percent once partial matches and misattributions were counted. The practical rule: theme inductively once, rebuild the list yourself, then have the model apply your list every period after.
Rare themes get over-predicted, and rare themes are what get escalated. Ashwin, Chhabra, and Rao, publishing in Sociological Methods and Research in 2025, found non-random bias in 10 of 19 codes, correlated with respondent characteristics, with sparse codes systematically over-predicted.
The same prompt does not produce the same answer twice. Work published by Thinking Machines Lab in September 2025 found 1,000 completions at temperature zero produced 80 unique outputs.
Neither major chat client exposes a seed, so a number you intend to report should be produced twice and compared.
Instructions placed after a large data block get lost. Liu and colleagues documented this in Transactions of the Association for Computational Linguistics in 2024. Put instructions above the data, and split large files rather than pasting everything at once.
Models agree with you. Sharma and colleagues documented sycophancy at ICLR in 2024, and OpenAI publicly withdrew a model update in April 2025 for being overly agreeable. Never ask a model to confirm a theme you already suspect is there.
Verbatims contain personal information nobody asked for. Support verbatims routinely include account numbers, addresses, and names. Strip them before the data leaves your environment, and check which data-handling tier your AI client account sits in rather than which vendor it is.
The critique you should know
Anyone presenting this metric to a skeptical audience should know the case against it, because the case is strong and it is published.
The central finding
The de Haan, Verhoef, and Wiesel result described earlier is the one that matters: across 93 firms and 18 industries, tested against observed two-year retention, Customer Effort Score performed worst of five customer feedback metrics and was the only one to miss statistical significance in the multivariate model.
No other customer feedback metric in common use carries a refutation this direct from a design this strong. Two of the five measures compared were effort variants, so the comparison is against three rival metrics rather than four.
The design is also unusually strong for the question, because retention was observed over two years instead of being self-reported at the time of survey.
Two scope limits belong with it. The study measured the 1.0 effort question, and the sample was Dutch. Neither weakens the finding substantially, and a guide that omits them is overstating.
The premise itself is conditional
The Ardelet and Benavent finding, drawn from 314,194 interactions across 96 brands, is that the effort-satisfaction relationship is not universally negative. In hedonic categories satisfaction is highest at low or intense effort and lowest in between, and at point of sale it rises with effort.
A metric built on the premise that less effort is always better is typically measuring the wrong thing in those categories. A home improvement retailer, a car dealership, or a specialty food business may be actively damaging the experience by removing the effort customers came for.
The commentary
Forrester argued in 2017 that effort functions as a transition metric, useful while an organization fixes basic usability and insufficient for loyalty afterward.
That is commentary rather than evidence and should be labeled as such, but it matches the shape of the peer-reviewed findings.
The counter-critique
Three arguments run the other way, and the first is the most important.
The study that condemns the metric also rescues it. The same de Haan paper states that Customer Effort Score "has incremental power, given that it is used in combination with a customer satisfaction-related CFM." Effort is a poor standalone loyalty predictor and a useful supplementary diagnostic, and that distinction is the guide's stated position.
Quoting both sentences from one paper is more honest than quoting either alone, and a reader who checks the source will find the guide represented it fairly.
The effect on satisfaction is real, and it runs the expected direction in utilitarian contexts. Ardelet and Benavent report an overall main effect of Cohen's f equals .401 across all sectors pooled, and find the relationship consistently negative across channels in banking, insurance, internet service, and mobile.
They do not report a separate effect size for those sectors. If your organization runs a support or account-servicing motion, that is the peer-reviewed warrant for measuring effort at all.
No single metric wins everywhere. Gaber Agag and colleagues, analyzing 11,547 observations across 668 firms and 16 years in the Journal of Retailing and Consumer Services in 2023, concluded that the best customer feedback metric differs by industry and by unit of analysis.
That conclusion is the citable finding and the paper's effort-specific results are not. Its effort data provenance is contested: the paper reports sourcing effort from the American Customer Satisfaction Index, and published documentation of the ACSI questionnaire does not show an effort item.
The conflict is unresolved in the literature and this guide does not resolve it.
The position this guide takes
Customer Effort Score is a diagnostic, not a loyalty metric. Run it to find friction in transactional journeys, report it beside satisfaction and resolution, and do not put it on an executive dashboard as a predictor of retention.
That position is what the evidence supports. It is more defensible than the vendor claim that effort predicts loyalty better than anything else, and more useful than the dismissal that it measures nothing at all.
It also happens to be the position that produces better work. A team that treats the score as a pointer toward friction will go and read the verbatims.
A team that treats it as a loyalty forecast will put it on a slide and defend it, which is a worse use of the same data.
Common mistakes
Eight of the ten failures below are covered in full earlier in the guide. This section exists so a reader who skipped to it still meets them, and the table says where the detail lives.
| Mistake | What goes wrong | Covered in |
|:----------------------------------------------:|:-------------------------------------------------------------------------------------:|:---------------------------------------:|
| Reporting effort alone | The metric is asked to predict loyalty, which its own evidence base says it cannot do | The short answer, and the critique |
| Mixing the two versions | Two incompatible series presented as one, with the direction inverted | The two versions, and the polarity trap |
| Reading movement inside the interval | Noise reported as trend, and credibility lost when it moves back | Sample size and precision |
| Surveying after the moment has passed | Recall measured instead of effort | Fielding and delivery |
| Excluding unresolved interactions | A healthy score alongside an unhealthy operation | Audience, targeting, and screening |
| Optimizing the score instead of the experience | Failure made faster, and the metric rewards it | When not to use this method |
| Surveying the same customers repeatedly | The sample skews toward people who have stopped reading | Timing and cadence |
| Treating a vendor benchmark as a target | A target borrowed from a table that describes nothing | Benchmarks |
Two further failures are not covered elsewhere, and both are easy to make during a platform change.
Changing the scale mid-program
Distinct from mixing the two versions, and easier to do by accident. A tooling migration that moves a seven-point item to a five-point one introduces a step change in the level of the score while leaving its shape intact.
Dawes found a shift of this kind between formats, with the ten-point version running 0.3 points lower after rescaling. The direction and size of a seven-to-five shift are not established, which is the point: you will not know how much of the step change you caused.
The change is usually invisible in the survey builder and obvious in the trend line three months later, by which time nobody remembers the migration date.
Fix: record the instrument specification alongside the data, and treat any change to it as a series break.
Setting a target before establishing a baseline
A team that commits to a number before it knows its own distribution has committed to an arbitrary figure, and the figure is usually borrowed from a vendor table.
Targets set this way tend to be either trivially met or unreachable. Both outcomes discredit the program, and the second one does it faster.
Fix: run three months, look at the distribution by segment, then set a target against your own baseline.
What the pattern across all ten has in common
Nine of these ten failures share a root cause, and naming it is more useful than the list.
Each one comes from treating the number as the finding. A program that treats the score as an index pointing at a segment will check the interval before reporting movement, will keep unresolved tickets in the sample because they are the interesting ones, and will notice when the instrument changed. A program that treats the score as the deliverable will do none of those, because each of them makes the number harder to report.
The exception is the polarity trap, which is a straightforward operational accident and catches careful teams too.
Synthetic respondents
Teams increasingly ask whether an AI-generated respondent panel can substitute for real customers on a metric like this one. The evidence generally says aggregate means track reasonably and subgroup structure does not.
James Bisbee and colleagues, publishing in Political Analysis in 2024, compared 7,530 real respondents against 3.6 million synthetic responses. Averages corresponded closely to the human data.
But 48 percent of regression coefficients differed significantly and 32 percent of those flipped sign, and variance was artificially deflated, which breaks power calculations.
For Customer Effort Score specifically, synthetic respondents cannot help, and the reason is structural rather than a matter of model quality. The metric measures a specific customer's experience of a specific interaction that actually occurred, and there is no synthetic equivalent of an event that happened.
A synthetic panel can tell you how a hypothetical persona says it would feel about a described support experience. That is a different construct, and substituting it produces a number that looks like a Customer Effort Score and is not one.
Use synthetic respondents for pretesting question wording and for checking that an instrument is comprehensible. Do not use them to produce a score, and do not use them to fill a segment where real response volume is thin.
How Sprig supports a Customer Effort Score survey
Sprig is an enterprise survey platform powered by AI agents, and several of its documented capabilities map onto the requirements this guide has set out.
Behavioral event triggering fires the instrument at the moment the interaction completes, which is the timing requirement the measurement depends on. The Rating Scale question type supports the seven-point form the 2.0 item requires, and the recontact waiting period provides a documented control on the survey fatigue that typically degrades a high-volume effort program.
AI Follow Ups can probe a low score while the respondent is still present. On a three-question survey that is generally the highest-value recovery available, because the alternative is a verbatim written after the customer has decided how they feel.
Two things are undocumented and matter before you rely on it. No accuracy or reproducibility claim is published for AI-generated follow-ups, and there is no documented control over how many fire on a single response.
What Sprig does not do for this method
Sprig documents no statistical significance testing anywhere in the product, so every period-over-period and segment comparison in this guide happens outside it, in a spreadsheet or an AI client.
Sprig also documents no in-product trend chart and no cross-study numeric comparison. For a metric whose entire value is its movement over time, that matters, and the honest workflow is export or a connector into an analysis environment. There is also no documented respondent identity key that persists across separate studies, so following the same customer's effort across waves is not an in-product operation.
Open-Text AI Analysis themes the friction probe, and the themes export alongside the responses. No minimum response count is documented for theming and no accuracy figure is published for theme counts, so treat the themes as a first pass to be checked instead of as a finished analysis.
Response export is CSV only, capped at 950,000 rows, and the download link expires after one week. For a program of this size that is rarely a constraint, and it is worth knowing before anyone builds a recurring pipeline against it.
Alternatives and adjacent methods
Four methods sit near this one, and each answers a question Customer Effort Score cannot.
Customer Satisfaction measures an evaluative judgment instead of perceived ease, and it outperforms effort as a retention indicator in the evidence cited throughout this guide. Run it alongside rather than instead, because the two together are what the de Haan finding supports.
Net Promoter Score measures stated advocacy at the relationship level. It answers a different question at a different altitude, it is typically reported quarterly rather than per interaction, and it carries a licensing regime that Customer Effort Score does not.
Task-based usability testing diagnoses the specific interface failure a low effort score points at. Twenty moderated sessions will generally tell you more about why a self-service flow is hard than four hundred survey responses, and the two are complementary rather than competing: the survey finds the flow, the sessions explain it.
First contact resolution, taken from operational data instead of from customers, is the outcome measure that belongs beside effort. It is not a survey metric, and that is exactly why it is useful here, because it is not subject to the response bias that affects everything else in this guide.
One further option deserves mention. The System Usability Scale and the single ease question both measure perceived difficulty at the task level rather than the episode level, and either is frequently a better fit than Customer Effort Score when the object of study is one screen or one flow.
Choosing between the three common metrics
These are independent tests rather than a sequence. Read every row, and where more than one matches, the disqualifying rows win: a hedonic context or an outcome question overrides a transactional match.
| If the thing you are measuring is | Run | Because |
|:-------------------------------------------------------------------------------------------------:|:---------------------------------------------------------:|:-----------------------------------------------------------------------:|
| A single transaction the customer initiated to get something done | Customer Effort Score, beside resolution and satisfaction | Friction is the actionable variable and the interaction is fresh |
| A single transaction where the customer made an evaluative judgment, such as a purchase or a demo | Customer Satisfaction | The judgment is about quality, not difficulty |
| The overall relationship, reported quarterly to a leadership audience | Net Promoter Score or a relationship satisfaction measure | Neither effort nor transactional satisfaction operates at that altitude |
| One screen or one flow, with a small number of participants | System Usability Scale or the single ease question | Both measure task-level difficulty at the screen or flow level |
| Whether the customer got what they came for | First contact resolution, from operational data | No survey metric answers this reliably |
| A hedonic experience the customer chose to spend time on | Satisfaction, and model effort as curvilinear if at all | The premise that less effort is better does not hold here |
Two rows deserve a note. The last one exists because of the Ardelet and Benavent finding, and it is the row most often got wrong: a business whose customers enjoy the process is measuring the wrong thing when it tracks effort reduction.
It overrides the first row wherever both apply.
The fifth row is the one teams resist, because it points away from surveys entirely. It is still the right answer to that question.
Frequently asked questions
What is a good Customer Effort Score?
There is no sourced cross-industry benchmark for Customer Effort Score, so no external number qualifies as good. The versions of the question differ in scale length, direction, and scoring convention, which means two organizations' figures are rarely on the same instrument.
Benchmark against your own prior periods on a frozen instrument instead, using your first three months as the baseline.
How is Customer Effort Score calculated?
Customer Effort Score is calculated as the percentage of respondents selecting the top three boxes on a seven-point agreement scale, which means the count of 5, 6, and 7 responses divided by the total count of valid responses, multiplied by 100.
Exclude blanks from the denominator rather than counting them as zeros. Report the arithmetic mean alongside the percentage, because the two move independently.
Is Customer Effort Score better than Net Promoter Score?
No, and the largest study to test them against observed retention found the opposite. de Haan, Verhoef, and Wiesel tracked 93 firms across 18 industries and found Net Promoter Score correlated with two-year retention at .170 while Customer Effort Score correlated at negative .073 and missed significance entirely in their multivariate model.
Customer Effort Score is better understood as an operational diagnostic than as a loyalty predictor.
How many responses do I need?
No methodology source publishes a Customer Effort Score sample-size rule, so derive the number from the precision you need rather than copying a threshold. At 100 responses a top-three-box score near 70 percent carries a margin of error of roughly 8.9 points at 95 percent confidence, at 400 responses roughly 4.5 points, and at 800 responses roughly 3.2 points.
Comparing two periods needs more again, because the interval on a difference is wider than the interval on either figure.
Should I use the 5-point or the 7-point version?
Use the seven-point version, which is Customer Effort Score 2.0 and presents the statement "The company made it easy for me to handle my issue" rather than asking about effort directly.
Scale research finds reliability and validity improve up to about seven categories, and the seven-point form is the one most current published discussion assumes.
Switching between versions inverts the direction of the score, so migrate fully rather than running both.
When should I send a Customer Effort Score survey?
Send a Customer Effort Score survey immediately after the interaction ends, in the product for a self-service task and within the same business day by email for an agent-resolved ticket.
Effort is a judgment about a specific experience, and the judgment degrades as the experience recedes.
Measuring effort in a periodic relationship survey asks the customer to reconstruct an event from memory rather than to report one.
Do I need a license to use Customer Effort Score?
No. "Customer Effort Score" is not a registered trademark. The Corporate Executive Board applied for the mark in April 2010, serial number 85027665, and the application was abandoned in March 2011 for failure to respond to an office action. No registration exists.
This is the opposite of the Net Promoter Score position, where registered marks carry published licensing terms, so the attribution footnote that accompanies Net Promoter Score does not apply here. Gartner acquired the Corporate Executive Board in 2017 and inherited whatever common-law rights survive, which is a question for your own counsel if you intend to use the term as a product name rather than editorially.
Can Customer Effort Score tell me why the experience was hard?
No, the score is an index rather than a diagnosis, and it identifies which segments to investigate rather than what went wrong in them. Pair it with a single open-text probe asking what made the interaction harder than it needed to be, then theme those responses and cross-tabulate the themes against channel and interaction type.
That cross-tab, not the score, is the output worth taking to an operations review.
Bottom line
Customer Effort Score is a good diagnostic and a poor loyalty metric, and most programs fail because they use it as the second thing.
Run the seven-point 2.0 statement, fire it immediately after the interaction, and report it beside a resolution item and a satisfaction item. Derive your sample from the precision you need instead of from a benchmark nobody can source.
Compare movement to the margin of error before calling it a trend.
If your goal is to find and remove friction in transactional journeys, Customer Effort Score is the right instrument, and the verbatims attached to it are typically where the value sits.
If your goal is to predict who will renew, the evidence points at observed retention and at satisfaction instead.
The most useful thing this metric will ever tell you is which journey to go and look at.
So the next study is not a bigger survey. Put the three-item instrument on your next few hundred support interactions, report effort, resolution, and satisfaction together, and read the verbatims attached to the lowest-scoring segment.
The customer effort template is the concrete artifact to start from. Ignore any sample-size figure printed on that template or any other, and size from the precision table above, which is derived rather than inherited.