Solutions
Experience measurement
Track sentiment and KPIs with AI-driven gap analysis
Strategic & foundational discovery
Uncover market whitespace with AI-led foundational studies
Journey & behavioral research
Connect user actions to motivations across the lifecycle
Market & consumer Insights
Understanding markets, audiences, & opportunity
Concept & prototype testing
Test designs and prototypes with rapid feedback
Agents
Design
Structure rigorous studies
Field
Run adaptive studies at scale
Analyze
Get statistical analysis you can trust
Synthesize
Turn results into research reports
Deploy
Email
Reach external audiences with native deliverability
Panels
Recruit from 300K+ verified participants
Web apps and websites
Embed studies in web experiences
Mobile apps
Run studies in iOS and Android apps
Customers
Community
Events
Join curated gatherings shaping the future of research
Blog
Insights on integrating AI into research craft
Book icon
Guides
Ultimate playbooks for enterprise survey research
Pricing
Sign in
Book a demo
Sign in
Book a demo
Guide

How to Run: Customer Needs Assessment Survey

October 2, 2026

By The Sprig Team

Example H2
Example H3
Example H4
Example H5
Example H6

Introduction

A customer needs assessment survey measures two things about the same list of outcomes. How much customers want each one, and how well they are getting it today.

The gap between those two numbers is the prioritization signal, and everything else in the method exists to make that gap trustworthy.

Four research traditions claim this territory, and they are typically presented as competing approaches. Outcome-Driven Innovation collects solution-free outcome statements and scores each on importance and satisfaction. The Kano model sorts features by asking a functional and a dysfunctional question about each one.

Importance-performance analysis plots the same two axes on a quadrant grid and reads the upper-left corner as the action list. The switch interview works qualitatively, reconstructing what pushed someone away from an old solution and pulled them toward a new one.

They are variants of one logic rather than four separate methods. Each starts from what customers are trying to accomplish instead of from what they say they want built. Two of the four, outcome scoring and importance-performance analysis, then measure the current state and treat the distance between the two as the reason to act. Kano infers the same priority from how people react to a feature being present and absent, and the switch interview reconstructs it from a story.

Run the quantitative version when you already have a candidate list of thirty to forty outcomes and need to rank them. A longer inventory has to be split across fieldings. Run the switch interview first when you do not, because the list has to come from somewhere.

A typical fielding needs roughly three hundred responses for outcome scoring, and closer to two hundred per segment if you plan to run Kano categorization, which is generally less stable than practitioners assume.

Scores tell you where the gap is, and they do it on a scale everyone in the room can compare. What they will not tell you is what to build in response.

What a customer needs assessment actually measures

The method is often described as asking customers what they want. That description is wrong in a way that matters, because customers asked what they want will name solutions, and solutions age badly.

Anthony Ulwick's contribution was to separate the job from the solution. A customer hiring a product is trying to get something done, and the things they are trying to get done can be written as measurable statements that survive whatever technology currently satisfies them.

Strategyn, the firm Ulwick founded, describes the output this way: a single job typically generates 50 to 150 of these outcome statements, each solution-free, measurable, and stable over time. A market contains more than one job, so the full inventory is a multiple of that figure.

Four traditions, one underlying comparison:

| Tradition | What it asks | What you get | Best when | |:-------------------------------:|:-----------------------------------------------------:|:-----------------------------------------------------:|:--------------------------------------------------------------:| | Outcome-Driven Innovation | Importance and satisfaction on each outcome statement | A ranked gap score per outcome | You have a written outcome list and need prioritization | | Kano model | A functional and a dysfunctional question per feature | A category per feature, such as must-be or attractive | You have discrete features and want to know which ones delight | | Importance-performance analysis | Importance and performance on attributes | A four-quadrant map | You are reviewing an existing product against known attributes | | Switch interview | What pushed, pulled, and held someone back | A narrative of the moment of change | You do not yet know what the outcomes are |

The quantitative three share an assumption worth naming early. They assume the list you field is the list that matters, which means the survey can only rank the needs you already thought of.

A voice of customer study or a round of strategic foundational research is what produces the list in the first place.

If ranking is the whole task and the list is short, MaxDiff will usually give you a cleaner ordering than importance ratings will, because it forces trade-offs rather than allowing everything to be rated important.

When to run a customer needs assessment

The strongest case for this method is a roadmap argument you cannot settle with opinion. Two teams both believe their area is underserved, meaning customers rate its outcomes as important and rate their satisfaction with them as low. Both have anecdotes, and neither has a number that compares the two on the same scale.

A needs assessment is also the right instrument before a repositioning.

If you are about to change what the product claims to do, you generally want evidence about which unmet outcomes the new claim would speak to, rather than discovering after launch that the gap you targeted was already closed.

It fits well ahead of annual planning, where the deliverable is a ranked list rather than a decision about one feature.

It fits well when entering an adjacent segment, because the same outcome list scored by a new audience shows you which needs travel and which do not.

It also fits well when a category has commoditized and the obvious moves have all been made. When every competitor satisfies the same outcomes equally well, the underserved outcomes are the only places differentiation is still available.

Where it fits in a research program: this is foundational work rather than evaluative work. It sizes the opportunity space, usually for a year at a time. It does not test whether your answer to the opportunity is any good.

Run it at most once or twice a year for a given market. Outcome importance is generally stable over months, and refielding quarterly will typically produce noise that teams will nonetheless interpret as movement.

When not to run a customer needs assessment

Five situations where a needs assessment is the wrong instrument, each with the instrument that is actually right.

You need to choose between two or three specific concepts. A needs assessment ranks abstract outcomes, and concepts are bundles of solutions to several outcomes at once.

A ranked outcome list will not tell you whether concept A beats concept B, because both may address the same top-ranked outcome with different execution quality. Run concept testing instead, ideally monadic so that respondents evaluate one concept without the contrast effect of seeing the other.

You need to set a price or choose a feature bundle. Importance ratings have no budget constraint in them, so respondents rate everything important because nothing costs anything.

Conjoint analysis exists precisely to impose that constraint, and it returns part-worth utilities, the estimated value a respondent places on each attribute level, which you can use to simulate a bundle. A needs assessment cannot simulate a configuration, and it was never designed to.

The third case is that you need to know which words to use in the market. Outcome statements are deliberately stripped of marketing language, which is what makes them stable and also what makes them useless as copy.

If the question is how to describe the value, run message testing against the audience who will read the words.

The fourth is that you are diagnosing one specific broken flow in a live product. A needs assessment operates at the level of jobs and outcomes, which is several altitudes above a checkout error or a confusing settings page.

Instrumentation, session replay, and a targeted in-product survey at the point of failure will find the problem faster and with a fraction of the sample.

Your list of outcomes came from the team rather than from customers. This is the most common failure and the hardest to see from inside, because a team-generated list feels complete to the team that generated it.

Fielding it produces a confident ranking of the wrong things, and the fix is qualitative and unglamorous, meaning a return to interviews before you field anything.

There is a sixth case worth flagging, though it is a timing problem rather than a method problem.

If the roadmap for the next two quarters is already locked, a needs assessment will produce a ranked list that nobody can act on, and the most likely outcome is that the study gets cited selectively to defend decisions that were already made.

Designing the study, and the fork you have to take first

Every needs assessment splits at one decision, and taking it late is what causes rework. The decision is whether you already have a validated list of outcomes.

If you do not, the study is qualitative and the deliverable is the list. If you do, the study is quantitative and the deliverable is a ranking.

Trying to do both in one fielding produces an open-text question that nobody codes properly and a ranking of a list that was still being written while people answered it.

Choosing between outcome scoring, Kano, MaxDiff, and the quadrant method

Start with the list question, then work down.

Do you have thirty or more outcome statements, written by customers or derived from customer language?

No. Run switch interviews or a voice-of-customer study first, and treat the fielding as a separate project. Twelve to twenty conversations will typically surface the outcomes that matter, and the fielding comes after.

The subsection below sets out how to run those conversations.

Yes. Continue.

Are the items discrete features, or are they outcomes?

Features, and you want to know which ones customers will notice. Use Kano. The two-question structure gives you a category per feature rather than a score, which is the right shape when the decision is include or exclude.

Outcomes. Continue.

Do you need a ranking, or do you need to know where the gap is?

A ranking, and the list is under twenty-five items. Use MaxDiff. A ranking on a longer list goes to importance and satisfaction scoring, presented as tiers. Forced trade-offs beat rating scales for ordering, and the output is a single scale everyone reads the same way.

A gap, meaning you need to know both how much customers want something and how badly they are currently served. Use importance and satisfaction scoring, which is the core of this guide.

Is the product already live and the attribute list stable?

Yes. Importance-performance analysis is the lighter version of the same logic, and its quadrant output is easier to present to people who will not read a score table.

No. Stay with importance and satisfaction scoring.

Running the switch interviews first

Switch interviews are the qualitative front half of a needs assessment, and this is the shape of them.

Recruit people who changed something recently, and be strict about what recently means. Not people who use your product, and not people with opinions about the category.

People who moved from one way of doing the job to another within the last ninety days, because the memory of the decision decays quickly and reconstructed reasoning arrives in its place.

Interview around the moment of change, not the product. The structure most practitioners use reconstructs a timeline in five beats, beginning with the first thought that the current approach was not working.

Then comes the passive looking that followed, often for weeks with nothing happening, then the active looking triggered by some specific event, then the deciding, and finally the first period of actually using the new thing.

Four forces act on that timeline, and Bob Moesta and Chris Spiek developed the framing that names them. There is the push of the current situation and the pull of the new solution.

There is also the habit that holds someone in place, and the anxiety about what switching might cost. This is a synthesis of how the two describe the interview rather than a fixed instrument.

What you are listening for is not preference. It is the language people use for the outcome they were failing to achieve, because that language is what the outcome statements get written from.

Twelve to twenty interviews will typically surface the outcome set for a single job. You will know you are close when new interviews stop producing new outcomes and start producing new phrasings of outcomes you already have.

Write the statements immediately after, while the transcripts are fresh, and write them in the customer's nouns.

A statement written three weeks later in a planning document will have drifted into the company's vocabulary, and the respondents who rate it will be rating something slightly different from what they said.

Sprig's jobs-to-be-done interview template is a reasonable starting structure, and the unmoderated version will scale the recruiting when moderated sessions are the bottleneck.

Building the instrument in Sprig

The importance and satisfaction design runs natively. You build the outcome list as a matrix question scored on a rating scale, field it as two blocks within one study, one for importance and one for satisfaction, and export both.

Four constraints are worth knowing before you design rather than after.

The matrix question does not render the same way everywhere. On Desktop Web it defaults to a full-scale matrix, and on Mobile Web and Mobile Apps it is always displayed as an accordion, meaning one row at a time.

That is a materially different response task, and it is a good reason to split a fifty-item list across two or three shorter matrix blocks rather than handing a phone user a fifty-row sequence.

MaxDiff is available on Enterprise plans and in the Standard Format only. It accepts between 4 and 24 items. By default, items per set is half the total item count capped at 8, and the number of sets is the item count divided by items per set, multiplied by three. Both values can be configured within the ranges the builder allows.

Scores run from 100 to negative 100. The constraint that catches teams out is post-launch editing: once a MaxDiff study is live you can edit the text of existing items, but you cannot add items or remove them. Lock the list before you launch, because refielding is typically the only remedy.

There is no combined view that shows MaxDiff scores alongside rating-scale results in one report. If you want both, you are exporting both and joining them outside the platform.

Sprig also documents no clustering, latent class, or statistical segmentation function, so any needs-based segmentation happens in your own environment and the export is where it starts.

Sequencing the importance and satisfaction blocks

Field importance first and satisfaction second, in separate blocks within the same study, with importance always preceding satisfaction for every respondent.

The reason is contamination running in one direction, and the asymmetry is what makes the fix so cheap.

A respondent who has just rated their satisfaction with an outcome tends to adjust their importance rating to be consistent with it, which compresses exactly the gap you are trying to measure. The reverse effect is weaker, though it is not zero, so do not randomize the order of the two blocks and then treat the halves as equivalent. If you want to know how large the order effect is in your population, split the sample and measure it as its own question.

Within each block, randomize the order of the outcome statements themselves. Primacy in a fifty-row grid is substantial, and unrandomized order will hand your top rank to whatever you listed first.

Writing the instrument

The outcome statement is where the method succeeds or fails, and it is the part teams rush.

A well-formed outcome statement has four parts in a fixed order. A direction of change, a unit of measure, an object of control, and a contextual clarifier.

"Minimize the time it takes to identify which transactions need manual review, when the queue is unusually long."

Read that back and notice what is missing. There is no product in it, no interface, no feature, and no vendor. That is usually the whole point.

The statement stays true whether the job is done on paper, in a spreadsheet, or by a model, which is why it can be scored against any solution including ones that do not exist yet.

The four-part outcome statement spec

A statement opens with a direction of change, meaning minimize, maximize, increase, or reduce, kept to the same four verbs across the whole list. Avoid improve and optimize, which do not specify a direction and invite each respondent to supply their own.

Unit of measure, which means time, number, frequency, or likelihood. If you cannot attach a unit, you have written a feeling rather than an outcome, and feelings do not have satisfaction levels that customers can compare.

Object of control, which is the thing being measured, stated in the customer's vocabulary rather than yours. If your customers say invoices and your product calls them billing documents, the statement says invoices.

Contextual clarifier. The situation in which the outcome matters, introduced with when or before or after. This is the part most often dropped, and dropping it is why so many outcome lists read as interchangeable.

"Minimize the time it takes to onboard a new user" is a statement about nothing in particular. The same sentence with "when the user is joining an existing team account" is a statement someone can rate.

Rules that keep the list scoreable

One outcome per statement. A statement containing and is two statements, and a respondent who feels differently about the two halves has no way to say so.

No solutions. If the statement names a mechanism, delete the mechanism and check that the sentence still parses. "Minimize the time it takes to find a saved report using search" becomes "Minimize the time it takes to find a saved report."

No comparatives against competitors. The scale already captures relative satisfaction, and naming one in the stem makes the item about them.

Keep one grammatical shape across the list. Mixed forms make a grid harder to read, and respondents in a hurry answer by position rather than by content.

Test the list on five people first, though not for agreement and not for whether they think the outcomes matter. Ask them to say each statement back in their own words.

Any statement that comes back changed will be read differently by different respondents, which is a measurement problem and not a wording preference. The writing survey questions guidance covers the hygiene on top of this.

How long the list can be

Strategyn's figure of 50 to 150 outcome statements describes what it takes to map a single job completely, not the number you can reasonably field in one study.

In practice a single fielding holds up to around forty statements rated twice, which is eighty judgments per respondent. Past that, straight-lining rises sharply, meaning respondents give every row the same rating instead of reading it, and the later rows start to show suspiciously uniform variance.

If the inventory genuinely runs to a hundred and fifty, split it. Field the list in two or three studies with a shared anchor block of eight to ten statements appearing in all of them, then use the anchor block to put the separate fieldings on a comparable footing.

This is more work and it is the honest way to handle a list that size.

The Kano question pair

If you are running the Kano variant instead of outcome scoring, the instrument is different enough to need its own spec.

Each feature gets two questions, always in the same order. A functional question asking how the respondent would feel if the feature were present, and a dysfunctional question asking how they would feel if it were absent.

The wording is fixed by convention and worth keeping fixed, because the category assignment table depends on it.

Functional form: "If [FEATURE, described in one sentence] works this way, how do you feel?"

Dysfunctional form: "If [FEATURE, same description] does not work this way, how do you feel?"

Both questions take the same five answer options, in this order:

  1. I like it that way.
  2. It must be that way.
  3. I am neutral.
  4. I can live with it that way.
  5. I dislike it that way.

The pair of answers maps onto one of six categories. Attractive, meaning the feature delights when present and is not missed when absent. One-dimensional, meaning satisfaction rises and falls with it.

Must-be, meaning its absence causes dissatisfaction and its presence earns nothing. Indifferent. Reverse, meaning customers actively prefer it absent, and questionable, which is the code for a contradictory answer pair and is a data-quality signal rather than a finding.

Kano and his co-authors introduced the two-question structure in 1984, in the Journal of the Japanese Society for Quality Control, volume 14 issue 2, pages 147 to 156.

Berger and eleven co-authors extended it in 1993 in the Center for Quality Management Journal, volume 2 issue 4, at pages 3 to 36, and that later paper is the source most Western practitioners are actually working from.

Three cautions apply specifically to this instrument and not to the outcome-scoring version.

Keep the feature descriptions concrete and free of benefit language. A functional question that describes the feature and its payoff is asking two things at once, and respondents will answer the payoff.

Watch the questionable rate as the study fields, and not only at the end.

A questionable share above roughly ten percent on a given feature usually means the description was ambiguous, and the honest move is to report that feature as unmeasured rather than to assign it the mode of the remaining answers.

Size for the categorical output. The precision arithmetic in the sample-size chapter is why two hundred per segment is the working figure here, and why a feature whose leading category is only a few points ahead of the runner-up has not actually been categorized.

Choosing scales, and the coarseness problem

Use a 1 to 10 scale for importance and a 1 to 10 scale for satisfaction, with the same direction and the same endpoint labels on both.

Sprig's Rating Scale question supports scale configuration, though the platform documentation does not currently publish the full permitted range, so confirm the maximum in the builder before you commit a design that depends on ten points.

Ten points instead of five is not a stylistic preference, though it is frequently treated as one. The writing studies guidance covers the scale hygiene underneath it.

A five-point scale gives each respondent five places to put a judgment, and outcome lists of thirty or forty items routinely need more discrimination than that. The visible symptom is compression. Importance ratings pile up on 4 and 5, the variance you need for ranking disappears, and outcomes that differ in the population look identical in the data.

Why subtracting satisfaction from importance is contested

Subtracting satisfaction from importance is the step a quantitative reader will challenge, and it is worth naming here rather than burying it.

Importance and satisfaction are measured on separate rating scales that share a label set and nothing else. A 7 on importance and a 7 on satisfaction are not the same quantity, and rating scales of this kind are ordinal at best.

Subtracting one from the other assumes both are interval and that their units are commensurate, and neither assumption is established. The critique chapter later in this guide sets out the full objection and the answer to it.

What follows from it here is narrow. It should not stop you from using the method. It should stop you from reporting the opportunity score, the combined importance-and-gap figure defined in the calculation chapter below, to two decimal places.

Reporting the gap as an ordering

Report the gap as an ordering, not a quantity. Outcome 14 ranks above outcome 22 is a defensible claim. Outcome 14 scores 1.3 points higher than outcome 22 is a claim your scale cannot support.

Where you need a central tendency for a single outcome, report the median alongside the mean and say so.

When the two disagree noticeably the distribution is bimodal, which typically means two segments rather than one population.

A note on what the two blocks are actually asking, because the distinction gets lost once the grid is built and the study is in the field.

The importance block asks about the outcome in the abstract, independent of how anyone currently achieves it, which is why the statement must be solution-free for the question to be answerable at all. The satisfaction block asks about the respondent's current experience of that outcome with whatever they are using today, including tools that are not yours.

Respondents will generally answer the second question about their whole workflow rather than about your product, and that is the correct behavior even though it complicates attribution. If you need product-specific satisfaction, ask for it as a third block and label it unmistakably.

Sample size, and what the numbers actually buy you

Sample size for this method is usually argued from habit, though it is better argued from the precision you need on the specific comparison you intend to make.

Three different questions need three different sample sizes, and teams commonly size for the easiest one and then make claims that need the hardest.

Estimating one outcome's mean

Every figure in this chapter rests on one assumption, so state it before using them. Rating scales on a 1 to 10 range typically produce a standard deviation somewhere around 2.2 in a reasonably heterogeneous population, and that is a working assumption rather than a sourced constant.

Check it against your own pilot. At a standard deviation of 1.8 the intervals below tighten by about a fifth, and at 2.6 they widen by about as much, so the assumption is doing real work.

At that spread, the 95 percent confidence interval on a single outcome's mean is roughly plus or minus 0.43 points at n equals 100, plus or minus 0.30 at n equals 200, and plus or minus 0.25 at n equals 300.

That is precise enough for reporting a single score at almost any sample size you would realistically field. This is the easy question, and sizing for it alone is how teams end up overclaiming.

Ranking two outcomes against each other

Ranking two outcomes against each other is the question the method exists to answer, and it needs a larger sample than estimating a single mean does.

One caveat about the model before the numbers. The figures below treat the two outcome means as independently estimated, which they are not, because the same respondents rate every outcome in the same grid. The estimates are paired, and the true standard error on a difference is the square root of two times one minus the correlation, over the sample size.

Correlations across outcomes in a single grid commonly run between 0.3 and 0.6, so the independent form is a conservative bound. Plan with it and let the bootstrap in the analysis chapter recover the real figure from your data.

On that conservative bound, comparing two means carries roughly 1.4 times the error of estimating one.

At n equals 100, the 95 percent interval on the difference between two outcomes is about plus or minus 0.61 points. At 300 it is about plus or minus 0.35, so translate those intervals into practice before you commit to a sample. If your top ten outcomes are separated by less than half a point, a sample of 100 cannot order them, and a sample of 300 can order them only loosely.

The honest presentation in that situation is a tier, meaning a group of outcomes the sample cannot separate, rather than a rank. Decide in advance what gap size you will treat as a real difference.

Three hundred responses is generally a reasonable default for outcome scoring against a single population, and it follows from a stated tolerance rather than from convention. Set the tolerance at separating two outcomes half a point apart, and 300 is roughly where the conservative bound on a difference drops under it. Going to 400 buys you a modest tightening, from about plus or minus 0.35 to about plus or minus 0.30, and is generally not worth a third more field time.

Categorizing features with Kano

Kano output is categorical, which changes the arithmetic.

Each feature gets assigned to a category based on which category the largest share of respondents falls into, and the stability of that assignment depends on how far the leading category is ahead of the second.

At n equals 200, the 95 percent interval on a category share near 40 percent is about plus or minus 6.8 points. At n equals 100 it widens to plus or minus 9.6.

A feature where the leading category holds 38 percent and the runner-up holds 33 percent is, at n equals 200, a feature you have not categorized at all.

Two hundred per segment is a defensible planning figure for Kano, and it is worth being clear about where it comes from. Set the tolerance at separating a leading category from a runner-up about fifteen points behind it, and 200 is roughly where the interval on a category share drops under half that gap. It follows from that stated tolerance and from nothing else, because no authoritative minimum exists in the original literature.

Kano and his co-authors, publishing in 1984, did not specify one, and neither did Berger and colleagues when they extended the method in 1993.

Treat any source that gives you a confident Kano sample minimum without showing the precision calculation behind it as repeating a convention.

Segment-level reporting

The rule that catches teams out is that segment-level sample is not a discount on total sample.

If you intend to report the ranking separately for enterprise and self-serve customers, you need the full sample in each, because each ranking is its own estimation problem.

Four segments reported separately at 300 each is 1,200 responses, and quotas are how you make sure the smaller segments actually fill rather than arriving at 40 responses and being reported anyway.

Audience, targeting, and screening

The population you want is customers who have recently done the job, not customers who have recently used your product.

Those overlap and they are not the same set, and the difference matters most in exactly the markets where a needs assessment is most valuable.

Someone doing the job with a spreadsheet and a competitor's tool has unmet outcomes worth knowing. Someone who opened your app twice last month has usage you can measure without asking.

Screen on the job, not on the product, since the two frames increasingly diverge as a category matures.

A screening question that asks whether the respondent has performed the relevant task in the last thirty days will typically produce a cleaner frame than any behavioral trigger will.

Getting the frame right in practice

For in-product fielding, attributes are how you carry role, plan, tenure, and segment into the data so the ranking can be cut by them afterward.

Set these before you launch the study, without exception. Attributes that are not attached at response time cannot be reconstructed later from the export.

Note the boundary here, because it is easy to lose in conversation. Attributes in Sprig can be used to filter and to target, and they are what cross-tab treats as segments.

They are not a segmentation the platform derives for you, and the distinction is easy to lose in a conversation.

Custom audience sampling is the control for the opposite problem, which is over-surveying the same active minority.

A needs assessment fielded to whoever is in the product this week will systematically over-represent heavy users, and heavy users are the population whose outcomes are already best served.

If your job frame extends beyond your own customer base, Sprig Panels is the route to non-customers, with over 300 available filters for building the target.

This is the in-platform route to people who evaluated you and chose something else, a group where the largest unmet gaps frequently sit. External panel providers and a link survey distributed through a community do the same job.

One further frame problem is worth anticipating. If your product is used by more than one role, the person who experiences the unmet outcome and the person who answers your survey are frequently different people.

An admin who configures the tool and an analyst who lives in it have genuinely different outcome sets, and pooling them produces a ranking that describes neither. Capture role as an attribute and check the segment split before you report anything company-wide.

Fielding and delivery

The instrument is long by in-product standards, and that shapes the delivery choice more than anything else.

An eighty-judgment matrix study is not a contextual microsurvey. Fielding it as an in-product intercept produces high abandonment and a sample biased toward people with unusual patience, which correlates with nothing you want to generalize about.

Email delivery or a link survey is generally the right vehicle, because both set the respondent's expectation that this will take a few minutes.

Sprig supports both alongside in-product delivery, and the delivery options documentation covers what each one can and cannot target.

Use the in-product channel for recruitment instead. A short intercept that screens on the job and invites qualified respondents into the longer instrument completes better than fielding the whole thing in the product.

Tell respondents the length honestly, in minutes. Understating it raises the click rate and lowers the completion rate, and the second number decides whether you can report the study.

Timing and cadence

Field a needs assessment once or twice a year for a given market.

Outcome importance generally moves slowly, over years rather than over quarters. What moves faster is satisfaction, which responds to your releases and to your competitors', and that is the half worth refielding when you want to know whether a gap has closed.

Kano categories are the exception to the stability claim, and they migrate in a known direction. A feature that is attractive when it first appears becomes expected once competitors match it, and becomes a must-be once the category standard has moved. Nothing in the literature fixes how long each step takes, and it varies with how fast the category is moving.

Nilsson-Witell and Fundin looked at this directly in the International Journal of Service Industry Management, volume 16 issue 2, at pages 152 to 168. Comparing customers at different stages of adopting an e-service, they found it was perceived as indifferent at introduction and as attractive later by the market at large, while early adopters already treated it as one-dimensional or must-be.

Read that carefully, because it is a cross-sectional comparison and not a longitudinal one. A design of that kind cannot fully separate migration over time from differences between the people at each adoption stage, so treat it as evidence consistent with migration rather than as proof of it.

The implication is narrow and genuinely useful for planning. A Kano study is a snapshot of a moving distribution, so date it in the report and resist comparing categories across studies more than eighteen months apart.

Sprig's resurvey waiting period and its guidance on survey frequency are the controls that stop an annual study from colliding with your continuous measurement programs.

Quality control and data hygiene

Long grids attract exactly the response behaviors that ruin a gap calculation, and the calculation is unusually sensitive to them because it operates on the difference between two numbers from the same person.

Straight-lining is the first thing to catch, and it is the easiest to catch mechanically. Flag any respondent whose within-block standard deviation across the outcome list is zero, and look hard at anyone below about 0.5 on a ten-point scale.

A respondent who rated all forty outcomes a 9 for importance has contributed nothing to the ranking and a great deal to the mean.

Check completion speed against a floor you derived from the instrument itself. Time the instrument yourself, take a third of that as the floor, and review everyone below it instead of deleting automatically.

Watch for differential blank rates across the two blocks. If satisfaction has substantially more non-response than importance, that is usually a real signal and not carelessness, because respondents who have never experienced an outcome have nothing to rate.

Handle those as genuine missing values and never as zeros, because a zero says completely unsatisfied, and completely unsatisfied is the opposite of no opinion.

Place one attention check inside the longer of the two blocks, worded as an outcome statement so it does not break the rhythm of the grid.

An instruction to select a specific point on the scale works, and it works better placed around two thirds of the way down the list than at the start, because attention typically degrades rather than being absent from the beginning.

Use one. Two in a single block irritate careful respondents, which is generally a worse trade than catching a few more careless ones.

Keep the failures in the dataset and report the analysis both ways. If the ranking is stable with and without them, the question is settled and you have said so. If it moves, that is worth knowing before anyone sees a number.

Bot detection covers the automated end, which matters more for link and panel fielding than for in-product.

Report your exclusions, always. State the starting count, the number removed, and the rule, before you show anyone a ranking. A ranking with an undisclosed cleaning step behind it is a ranking nobody can check.

The core calculation, and what it is really doing

Ulwick's opportunity score is the best-known way to turn importance and satisfaction into a single number. It is also the most misunderstood piece of arithmetic in the whole method, and understanding what it actually does will change how you read it.

The formula is usually written like this:

Opportunity = Importance + max(Importance - Satisfaction, 0)

Work through the algebra and it collapses into two cases.

If Importance > Satisfaction:      Opportunity = 2 x Importance - Satisfaction
If Satisfaction >= Importance:     Opportunity = Importance

That reduction is the whole story, and it is worth sitting with.

What the two-case reduction tells you

When importance exceeds satisfaction, the formula weights importance twice and satisfaction minus one. The formula is not treating the two inputs as equals.

It is an importance score with a satisfaction penalty attached, and importance carries double the influence of satisfaction on the result.

In the second case, satisfaction drops out of the equation entirely. Any outcome where customers are as satisfied as the outcome is important returns its importance score unchanged, and it does not matter whether satisfaction was one point above importance or seven.

Worked example: the same score, different situations

Take an outcome rated 3 for importance, where a satisfaction of 4, a satisfaction of 9, and a satisfaction of 10 all return an opportunity score of 3.

Those are three genuinely different situations for a product team to be in. All three are outcomes where satisfaction already exceeds importance, which is what overserved means, and they differ substantially in how far past the line they sit.

The formula reports them identically, because satisfaction drops out of the equation above the crossover point. That is correct given what the score is designed to surface, and misleading if you read it as a description of the outcome.

Now take two different outcomes that happen to return the same score. Importance 8 with satisfaction 6 scores 10, and importance 7 with satisfaction 4 also scores 10.

The first is a high-importance outcome with a modest gap, and the second is a moderately important outcome with a large gap. A team reading only the ranked list will treat them as equivalent opportunities, and they are not.

Always report the score with its two inputs

Never present the opportunity score alone, and instead present importance, satisfaction, and the score together on the same row, every time. Right-sizing the report is the general form of the same discipline.

On the 1 to 10 scale recommended earlier, the maximum possible score is 19, which occurs at importance 10 and satisfaction 1. On a 0 to 10 scale it is 20.

The range is a property of the inputs rather than a designed scale, so do not describe the output as a 0 to 20 scale as though that were a documented feature of the method.

Be careful with ties, which are commonly more numerous than people expect. Because satisfaction drops out above the crossover point, scores cluster heavily in the middle of a typical outcome list, and ties in the middle of the ranking are common.

Break them by looking at the two input numbers in the response data, not by adding decimal places, which generally manufactures precision the scale does not have.

The quadrant alternative

Importance-performance analysis handles the same two inputs without subtracting them, and for some audiences that is the better presentation.

Martilla and James introduced it in the Journal of Marketing in 1977, across pages 77 to 79. The method plots each attribute on a grid with importance on one axis and performance on the other, then divides the grid into four quadrants.

High importance and low performance is the concentrate-here quadrant, and it is the action list. High on both is keep-up-the-good-work, which is the protect list.

Low on both axes is the low-priority quadrant, and it is usually the largest one. Low importance and high performance is possible overkill, which is the quadrant worth reading carefully because it is where you find effort you can redirect.

The appeal is that no arithmetic combines the two axes, so the unit-mixing objection does not apply. Nothing is subtracted from anything, and neither axis is weighted against the other.

The appeal is partly illusory, because the quadrant boundaries have to go somewhere and that placement is a judgment. Put the crosshair at the scale midpoint and a study where everything scored above 6 on importance lands entirely in the top half.

Put it at the sample means and the quadrants are defined relative to your own data, which makes them uncomparable across studies. Both conventions are in common use, and they frequently assign the same attribute to different quadrants.

State which crosshair you used, every time, and show the raw scatter underneath the quadrant lines so a reader can see how much the classification depends on where you drew them.

Use the quadrant version when the audience will generally not read a score table, and use the opportunity score when the list is long enough that a scatter plot becomes unreadable. Around thirty attributes is roughly where the plot stops working.

Benchmarks, and why this guide does not give you any

There is no benchmark for an opportunity score, and the absence is worth stating plainly because you will find plenty of numbers online that look like benchmarks.

The most repeated convention holds that a score above 15 marks an underserved outcome and below 10 an overserved one.

A second uses 12 and 8 for the same boundaries. A third divides the range into six labelled bands.

They cannot all be right, and they disagree because none has a published source behind it.

Neither Ulwick's own writing nor Strategyn's published material sets a numeric threshold for underserved or overserved. The thresholds circulating in secondary write-ups are conventions that acquired authority through repetition.

This matters more than a citation quibble, because a threshold is a decision rule.

Applying a 15 cutoff to a study whose top score is 13 tells a team nothing is underserved, which is a conclusion about your scale and not your customers.

Setting the cuts from your own distribution

Use relative position inside your own study. The top decile is your underserved set whatever the absolute numbers are, and the bottom decile where satisfaction exceeds importance is your overserved set.

Set the cut before you see the ranked list. Typically the hardest discipline in the method, and frequently the one skipped.

Deciding afterward how many outcomes count as underserved is a choice about how much work the team wants to sign up for, dressed as an analytical finding.

Where you do want a cross-study anchor, build your own by fielding the same core outcome block annually and comparing this year's distribution to last year's.

Your own prior wave is the only benchmark for this method that is defensible, and it becomes available on your second run.

The same discipline applies to Kano studies and to anything else you plan to track. Category shares vary enormously by product maturity and by market, and a must-be share that would be alarming in a young category is unremarkable in an established one.

There is no external table to check yourself against, and the honest version of a tracking study is one that treats its own first wave as the baseline.

One piece of housekeeping while the subject is thresholds. This guide contains working figures that are conventions rather than findings, including forty statements as a practical instrument ceiling, twelve to twenty switch interviews, ninety days as a recency window, a ten percent questionable rate for Kano, roughly thirty attributes as the point where a quadrant plot stops being readable, and the straight-lining and rank-shift defaults in the analysis prompts.

None of those has a published derivation behind it, and this guide would be inconsistent if it presented them as though it did. Treat them as starting points to adjust against your own instrument and population.

Reading the results and turning them into decisions

A ranked list is not a roadmap, and the gap between those two things is where most needs assessments quietly die.

Start by reading the three-column table instead of the score, meaning importance, satisfaction and opportunity, sorted by opportunity, with the top twenty visible on one page.

What you are looking for is not the top row but the shape of the whole top section.

Four importance and satisfaction patterns

High importance and low satisfaction is the obvious one, and it is usually already known to somebody in the building. The value of the study in this quadrant is not discovery of something new.

It is that the finding now has a number attached and can be compared against the other candidate on the list.

Moderate importance and very low satisfaction produces a high score through the gap rather than through the importance.

These are frequently real opportunities and they are frequently small ones, because the population who cares is limited. Check the segment cut before committing anything to the roadmap.

High importance and high satisfaction is the table stakes region, where there is nothing to build and a great deal to protect. A competitor who closes on one of these outcomes is attacking your retention rather than your acquisition.

High satisfaction and low importance is overinvestment, and this quadrant is the one teams skip. It is where you find the features that consume maintenance and roadmap attention out of proportion to what customers get from them.

Reading this quadrant honestly is how a needs assessment pays for itself in things you stop doing.

What the readout table looks like

A worked fragment, using the top six rows of a fictional study of forty outcomes fielded to 312 respondents.

| Outcome | Imp | Sat | Opp | Tier | n | |:---------------------------------------------------------------------------------------------:|:---:|:---:|:----:|:----:|:---:| | Minimize the time it takes to confirm a payment cleared, when the payer is a new counterparty | 8.7 | 4.1 | 13.3 | 1 | 298 | | Minimize the likelihood of discovering an error after a period has closed | 8.4 | 4.3 | 12.5 | 1 | 305 | | Reduce the number of people who must approve a routine exception | 7.9 | 4.2 | 11.6 | 2 | 288 | | Minimize the time it takes to explain a variance to someone outside finance | 7.2 | 3.4 | 11.0 | 2 | 241 | | Increase the likelihood that a forecast survives the first review unchanged | 7.6 | 4.6 | 10.6 | 3 | 302 | | Minimize the effort required to onboard a new approver, when the approver is not a daily user | 6.8 | 3.1 | 10.5 | 3 | 197 |

Four things to notice in that fragment, because they are the things a reader should be doing with every readout of this kind.

The top two sit in one tier, which means the study cannot order them and the presentation should not pretend otherwise. On this study's numbers the bootstrap interval on the difference between two opportunity means runs to roughly plus or minus 0.7 points, and the 0.8 separating rows one and two does not clear it. Presenting the first row as the top priority is a claim the data does not support.

Row four and row six have the lowest satisfaction scores on the page and neither is in the top tier, because their importance is moderate.

These are commonly the most interesting rows in a study, and they are the ones a reader scanning the opportunity column will skip.

Row six has an n of 197 against a study base of 312, which is a third of respondents leaving it blank. That blank rate is itself a finding and belongs in the readout.

A large share of this population has never onboarded an approver, and the outcome may belong to a segment rather than to the market.

Every row carries its two inputs, and if you strip the Imp and Sat columns then rows five and six both read as roughly equivalent opportunities, which they are not.

From pattern to decision

Take the top ten outcomes and ask one question of each. What would have to be true for this gap to close?

Some answers will be product work. Some will be onboarding, because the outcome is already achievable and customers do not know it.

Some will be pricing or packaging, because the capability sits on a tier the respondent is not on. That third category is common and invisible if you only read the ranking.

Write the answer next to each outcome before the readout, in a sentence rather than a label.

A study that arrives as twenty ranked rows will generally be debated rather than acted on. A study that arrives as twenty rows with a proposed mechanism next to each will be decided.

And carry the two input numbers into every downstream artifact. The moment the opportunity score travels on its own, someone will compare a score of 11 from this study to a score of 14 from a different study with a different outcome list, and the comparison will be meaningless in a way that is very hard to see.

Sprig's guidance on demonstrating research impact is about the step after this one, which is making the decision traceable back to the study.

Presenting the study

The readout is where a needs assessment most often fails to land, and the failure is usually structural rather than analytical.

Lead with the decision instead of the method, which means the first slide is the top tier with the proposed mechanism next to each item. The method, the sample, and the exclusion rules belong in an appendix that you point at and do not walk through.

Show the tiers as tiers. If you present a numbered list of twenty rows, every reader will fixate on the ordering within the top five, which is precisely the ordering your confidence intervals do not support. Grouped boxes communicate the uncertainty without anyone having to read an interval.

Name what the study cannot answer. It cannot tell you whether any particular solution will close a gap, it cannot price anything, and it cannot tell you the sequence in which to work. Saying this explicitly protects the study from being cited later for a claim it never made.

Finally, agree the refresh date in the room. A needs assessment with no scheduled second wave becomes a static artifact that gets quoted for years, long after its satisfaction numbers stopped being true.

Running the analysis with Claude

The calculation is simple enough to do in a spreadsheet and easy enough to get wrong that it is worth automating with a check on top. This chapter gives you a working sequence you can run end to end.

You will need the exported response file from Sprig, which you can pull through the data export or through the Claude integration if you want the data to arrive in the conversation directly.

Everything below assumes one row per respondent and one column per outcome per block.

Replace the bracketed placeholders before running each prompt, and note that defaults are given wherever a sensible one exists.

Step 1. Load and describe the file

You are analyzing a customer needs assessment export.

WHAT THIS STEP DOES: reads the file, identifies which columns are importance
ratings and which are satisfaction ratings, and reports the structure back to
me so I can confirm it before any calculation runs.
WHAT IT RETURNS: a table of column name, inferred block (importance or
satisfaction), inferred outcome ID, and count of non-blank values.

File: [PATH TO EXPORT FILE]
Importance columns are identified by: [COLUMN PATTERN, default "imp_"]
Satisfaction columns are identified by: [COLUMN PATTERN, default "sat_"]

The outcome ID is the column name with the block prefix removed. An importance
column and a satisfaction column belong to the same outcome when their IDs match
exactly. Report any ID that appears in one block and not the other.

Use code to calculate this, not estimation. Do not summarize the file by
reading a sample of rows.

Report any column you cannot confidently assign to a block. Do not guess.
Save the structure table as needs_structure.csv.

Step 2. Clean and exclude

WHAT THIS STEP DOES: removes responses that cannot contribute to the gap
calculation and reports exactly what was removed.
WHAT IT RETURNS: a cleaned dataset plus an exclusion log.

Rules:
1. Exclude a respondent from an outcome's calculation when either the
   importance or the satisfaction value for that outcome is blank. Blank means
   empty, null, or a non-response code, and blank is NOT zero. The scale is
   [SCALE MINIMUM, default 1] to [SCALE MAXIMUM, default 10]. Any value outside
   that range is a coding error, so report it and do not treat it as a rating.
   The scale minimum is a real rating and counts as a value below any threshold
   I set later.
2. Flag respondents whose within-block standard deviation is under
   [STRAIGHTLINE SD FLOOR, default 0.5] in either block.
3. Flag respondents who completed in under [SPEED FLOOR IN SECONDS]. Derive this
   yourself if I leave it empty: time the instrument at one judgment per three
   seconds, take a third of that, and tell me the number you used.

Use code to calculate this, not estimation.

Do not delete flagged respondents. Write two datasets, one including them and
one excluding them, and carry both through every later step so I can see whether
the ranking changes.

Save as needs_clean_all.csv, needs_clean_strict.csv, and needs_exclusions.csv.

Step 3. Compute the opportunity scores

WHAT THIS STEP DOES: computes mean importance, mean satisfaction, and the
opportunity score for every outcome.
WHAT IT RETURNS: one row per outcome with all three numbers and the response
count behind each.

Formula: Opportunity = Importance + max(Importance - Satisfaction, 0)

Use code to calculate this, not estimation. Do not compute any value by
reasoning about it in prose.

Use the per-outcome mean of importance and the per-outcome mean of
satisfaction, computed over respondents who answered both for that outcome.
Report the n behind each outcome separately, because it will differ.

Run this on both datasets from Step 2 and save needs_scores_all.csv and
needs_scores_strict.csv, each sorted by opportunity descending. Report any
outcome whose rank differs by more than three positions between the two.

Step 4. Recompute independently

WHAT THIS STEP DOES: two separate checks on Step 3. A transcription check, and
a check on which estimand Step 3 actually computed.
WHAT IT RETURNS: a comparison table and a verdict.

Check one, transcription. Recompute using the two-case reduction instead of the
max() form:
  If Importance > Satisfaction:  Opportunity = 2 * Importance - Satisfaction
  If Satisfaction >= Importance: Opportunity = Importance
These two forms are algebraically identical, so any difference at all is a
coding error. Report exact equality or name the rows that differ.

Check two, estimand. Compute each respondent's opportunity score for each
outcome first, then average those per-respondent scores. This is a different
quantity, not a second opinion on the same one.

Use code to calculate this, not estimation.

Because max() is convex, the respondent-level average is always greater than or
equal to the Step 3 figure. The two are exactly equal only when every respondent
in that outcome has importance above satisfaction. A negative difference is
therefore an error and not a finding, so if you see one, say so and stop.

Report the size of the difference per outcome. Where it exceeds
[DIVERGENCE TOLERANCE, default 0.5 points], name the outcomes, because a large
gap means the population is split between people the outcome is failing and
people it is not.

If check one disagrees at all, or if check two produces a negative difference,
say that you cannot reconcile the numbers and show me the specific rows. Do not
resolve a mismatch by picking the answer that looks more reasonable.

Save as needs_verification.csv.

Step 5. Rank with uncertainty

WHAT THIS STEP DOES: tells me which parts of the ranking are real and which
are noise.
WHAT IT RETURNS: the ranked list grouped into tiers.

Resample respondent rows [BOOTSTRAP ITERATIONS, default 2000] times. Resample
whole rows, so that every outcome in a replicate comes from the same respondents
and the correlation between outcomes is preserved. Recompute the full opportunity
score inside each replicate from that replicate's own means. Do not build an
interval on the opportunity score by combining separate intervals on importance
and satisfaction, which would be wrong.

Use code to calculate this, not estimation.

Then tier as follows, and use exactly this algorithm rather than eyeballing
interval overlap, which is not a valid test and does not produce a partition.

1. Sort outcomes by opportunity score, highest first.
2. For each adjacent pair in that sorted order, compute the bootstrap
   distribution of the difference between their opportunity scores, taking the
   difference within each replicate so the pairing is preserved.
3. Cut between the pair when the 95 percent interval on that difference excludes
   zero. Otherwise keep them in the same tier.
4. Report the number of tiers and the interval for every cut you made and every
   cut you declined to make.

If the top ten outcomes collapse into one or two tiers, say plainly that the
sample cannot rank them and that the ranked order must not be presented as an
ordering.

Save as needs_tiers.csv.

Step 6. Cut by segment

WHAT THIS STEP DOES: repeats the scoring within each segment.
WHAT IT RETURNS: a per-segment ranking and a divergence report.

Segment column: [SEGMENT COLUMN NAME]
Minimum n to rank a segment: [SEGMENT RANKING MINIMUM, default 300]
Minimum n to report a segment directionally: [SEGMENT FLOOR, default 100]

A segment between the floor and the ranking minimum gets its scores reported
with an explicit note that the sample cannot support an ordering. A segment
below the floor is named and not scored.

Use code to calculate this, not estimation.

Do not rank a segment below the ranking minimum, and do not score one below the
floor. Name it and state its n instead.

Report the outcomes whose rank differs by more than [RANK SHIFT THRESHOLD,
default 5 positions] between any two segments. Those are the outcomes where a
single company-wide ranking is hiding a real disagreement.

Save as needs_by_segment.csv.

Step 7. Write the readout table

WHAT THIS STEP DOES: assembles the presentation table.
WHAT IT RETURNS: the top [TOP N, default 20] outcomes with everything a reader
needs to interpret them.

Columns, in this order: outcome ID, outcome statement, mean importance, mean
satisfaction, opportunity score, tier, n.

Use code to calculate this, not estimation.

Do not include the opportunity score in any output that omits the importance
and satisfaction columns.

Round all three numbers to one decimal place. Do not report two.

Save as needs_readout.csv.

Step 8. Stress-test the interpretation

WHAT THIS STEP DOES: attacks my own reading of the results before anyone else
does.
WHAT IT RETURNS: a list of specific objections with the rows that support them.

Here is the readout table and the tier assignments. Here is the interpretation
I plan to present: [PASTE YOUR INTERPRETATION].

Use code to calculate this, not estimation, wherever a claim can be checked
numerically.

Find every claim in my interpretation that the data does not support. In
particular, check:
1. Any ranking claim between outcomes in the same tier.
2. Any ranking claim about a segment below the ranking minimum used in Step 6.
3. Any comparison of an opportunity score to a number from outside this study.
4. Any outcome whose n is more than [BLANK RATE THRESHOLD, default 15 percent]
   below the study base, where I have not mentioned the blank rate.

For each problem, quote the specific claim and give the rows that contradict it.

If my interpretation is well supported, say so plainly and say which parts you
checked. Do not manufacture objections to appear thorough, and do not soften a
real objection to be agreeable.

Save as needs_challenge.md.

Sprig's guidance on prompting for research covers why the preamble and the escape hatch matter as much as the instruction itself, and the cross-tab walkthrough is the companion piece if the segment step is where most of your interest sits.

The critique you should know before you present this

Someone in the room will object to the arithmetic, and they will be substantially right. Every composite metric attracts this eventually, and the case against NPS is the version most teams have already argued. Knowing the objection in advance is the difference between defending the method and losing the meeting.

The unit-mixing objection

Gerry Katz of Applied Marketing Science published a detailed critique of Outcome-Driven Innovation that goes directly at the opportunity formula. Note the interest before you cite it, because Applied Marketing Science is a voice-of-customer consultancy that sells competing methods. The argument stands on its own terms regardless.

His core point is structural rather than empirical: "Importance and performance are two entirely separate constructs, and should not be used in the same equation."

He puts the mechanics more bluntly. Katz identifies two problems with the formula, and names the first as the fact that "satisfaction is subtracted from importance", an operation he describes as "like subtracting apples from broccoli."

Katz also reports John Hauser, one of the most cited people in quantitative marketing, saying of the same measure: "This measure is pseudo-scientific. It mixes units of measure. No self-respecting engineer would ever do such a thing."

That is a hard quote to argue with, and you should not try. Conceding it early usually costs you nothing.

The weighting objection

The weighting objection follows from the algebra of the opportunity formula, and it is one you can verify yourself in a minute.

The formula weights importance at two and satisfaction at minus one. Nothing in the method's published rationale establishes why that ratio is the right one, or why it should be the same ratio in every market.

A team that believes satisfaction should count equally would write importance minus satisfaction and get a different ranking from the same data.

The choice of weighting is a judgment embedded in the formula and presented as arithmetic.

The ordinal objection

Martilla and James flagged the underlying measurement issue in 1977, in the paper that introduced importance-performance analysis. Their caution was that "Median values as a measure of central tendency are theoretically preferable to means because a true interval scale may not exist."

Rating scales are ordinal. Means of ordinal data, and differences between those means, are conventions the field accepts because they are useful and not because they are justified.

Note where Martilla and James landed on their own caveat, though. They raised it and then used means, judging the violation immaterial in the absence of significance testing. The objection is real and it is not fatal, which is roughly where this whole method sits.

The critiques aimed at Kano and the quadrant method

Importance-performance analysis and the Kano model each attract their own objections, and they are different objections.

For importance-performance analysis, the sharpest issue is stated importance itself. The method asks respondents how important an attribute is and treats the answer as the importance weight, which assumes people can introspect accurately on what drives their own satisfaction. Oh's 2001 review of the method, in Tourism Management volume 22 issue 6 at pages 617 to 627, set out its conceptual and methodological problems at length, with the absence of any clear definition of attribute importance at the center of them.

The practical version of that objection is easy to test in your own data. Derive importance statistically, by regressing overall satisfaction on the per-attribute performance ratings, and compare the derived weights to the stated ones. When the two orderings diverge substantially, and they frequently do, you have learned that your respondents' stated priorities and their revealed ones are not the same list.

For Kano, the objection is about the forced categorization. Each feature gets assigned to a single category based on the modal answer pair, and the mode discards the distribution behind it. A feature where 40 percent say attractive and 35 percent say must-be is reported as attractive, which erases a third of the population who will be actively dissatisfied without it.

Report the full category distribution per feature, not just the winning label. The same applies to any single-number summary, including whether to track NPS at all. It costs one extra column and it prevents the most common misreading of a Kano table.

What survives the objections

What survives the three objections above is narrower than the method's advocates usually claim.

The opportunity score is not a measurement of anything in the world.

It is a sorting heuristic, and it is a good one for a specific purpose, which is surfacing outcomes where high importance coincides with low satisfaction in a list too long to eyeball.

That purpose does not require interval scales or commensurate units. It requires only that the ordering the heuristic produces is more useful than the ordering you would have produced by argument, and in most product organizations it plainly is.

So use it, and use it the way you would use any heuristic. Do not present the score as a measured quantity, however often you see it done. Do not compare scores across studies with different outcome lists.

Do not set numeric thresholds on it, whatever you find published elsewhere. And when someone raises the unit-mixing objection, agree with them, then point at the importance and satisfaction columns you kept next to the score for exactly this reason.

If the room will not accept a heuristic, the alternative is MaxDiff for the ranking and a separate satisfaction measure alongside it. That combination costs more and answers a narrower question, and it does not mix units.

Common mistakes

Eight failures that show up repeatedly, roughly in the order they occur during a study.

Writing the outcome list from the inside

The team drafts forty outcome statements in a workshop, everyone recognizes them, and the list feels complete. It is complete with respect to what the team already believes, which is exactly the thing the study was supposed to test.

The tell is that no statement in the list surprises anyone, which is typically the clearest early warning available. The three rules for writing survey questions are a hygiene layer, and no hygiene saves a list drafted from the wrong source.

A list derived from customer interviews will usually contain three or four statements that make somebody in the room uncomfortable, and those are frequently the ones that rank highest.

Leaving solutions in the statements

"Reduce the time it takes to export a report to CSV" is a feature request with a verb in front of it.

The customer who does the job by copying rows into another system has no way to rate it, and the outcome disappears from your data because the statement assumed the mechanism.

Strip every mechanism, including the ones that seem generally harmless. If the statement still makes sense when the product does not exist, it is an outcome.

Fielding both blocks in one grid

Asking importance and satisfaction side by side for each outcome, in a single two-column grid, looks efficient and destroys the measurement.

Respondents anchor the second rating on the first, and the gap you are measuring shrinks toward zero for reasons that have nothing to do with your product.

Separate blocks, importance first. This is typically the single cheapest fix on the list.

Treating non-response as zero

Coding a blank satisfaction rating as zero silently inverts findings. A respondent who has never attempted an outcome leaves satisfaction blank, and a blank coded as zero produces maximum dissatisfaction from someone with no experience at all.

The outcomes most likely to collect blanks are the novel ones, which means the error concentrates precisely on the outcomes you are most interested in.

Applying a borrowed threshold

Applying a published opportunity-score cutoff to your own study is the single most common analytical error in this method. A 15 cutoff you found in a blog post is not a decision rule for your data, because it was derived from nothing and calibrated against no distribution.

Reporting the score without its inputs

Once an opportunity score travels alone it becomes uninterpretable, and it becomes uninterpretable in a way that looks fine. Two outcomes at 10 look like the same finding, and the chapter on the core calculation shows why they usually are not.

Ranking a list the sample cannot rank

If the top ten outcomes sit inside each other's confidence intervals, presenting them in order is presenting noise with a number attached. Tier them and say so, which is often an easier conversation than it looks.

The pressure not to do this is real, because tiers look less decisive than a ranked list. A ranked list that reverses itself on the next wave looks considerably worse.

Running the study with the roadmap already locked

The study lands, the ranking is clear, and the next two quarters were committed before fielding started. What happens next is predictable to anyone who has watched it happen once.

The findings get cited where they agree with the existing plan and set aside where they do not, and the organization concludes that research does not change anything. Read the four lessons on evangelizing research before it happens to you.

Check the decision window before you field, because it is commonly the constraint nobody raised. If nothing can move for six months, field in five.

A ninth failure sits underneath several of the others, and it is worth naming on its own. Teams treat the study as a one-time event rather than as the first wave of a series.

Almost everything valuable about this method compounds across waves. The threshold question resolves itself once you have your own distribution, the satisfaction half becomes a real measure of whether your releases worked, and Kano migration becomes visible rather than theoretical.

A single wave gives you a ranking. Two waves give you a direction, and direction is generally what a roadmap argument actually needs.

Synthetic respondents and where they fit

Simulated respondents are increasingly offered as a shortcut for exactly this kind of study, and the offer deserves a direct answer.

Do not field an outcome list to synthetic respondents and report the resulting ranking as customer needs. The model has no unmet needs, and what it produces is a plausible ordering derived from the language in the statements and from patterns in its training data.

That ordering will look reasonable, which is the problem, because a reasonable-looking ranking is indistinguishable from a real one on the page.

The specific failure mode here matters more than the general objection. Opportunity scores depend on the gap between two ratings, and a model asked to supply both will generally produce a coherent pair.

Real populations frequently produce incoherent pairs, and that incoherence is often where the signal lives.

Where simulation does help

Instrument pretesting is generally the honest use, and it is typically a ten-minute job. Ask a model to answer your outcome statements and then to explain what it understood each one to mean.

Statements it interprets differently from your intent are statements your respondents will also misread, and you have found that in ten minutes instead of after fielding.

Sprig's own position on where human judgment stays non-negotiable is the relevant frame. The model is typically useful on the instrument and consistently unreliable on the population.

How Sprig supports a customer needs assessment

A candid inventory of what the platform does for this study and what it leaves to you.

Sprig is an AI-native research platform built around in-product measurement, and a needs assessment sits slightly outside its center of gravity, so the boundaries are worth being precise about.

What Sprig runs natively

The instrument builds directly. Matrix questions carry the outcome grid, and Rating Scale questions carry the per-outcome scoring. Matrix rendering is supported on Desktop Web and on Mobile Web and Mobile Apps.

Delivery covers the three channels a study like this generally needs. In-product for the screening intercept, email for the main instrument, and link surveys for anyone outside your product.

Panels reaches non-customers, with over 300 available filters for constructing the frame. This is how you score outcomes for the people who chose a competitor.

Attributes carry your segment definitions into every response row, and cross-tab analysis will cut results by those attribute values.

MaxDiff is available on Enterprise plans in the Standard Format, for item lists between 4 and 24. Scores run from 100 to negative 100.

Open-text analysis handles the qualitative follow-up, grouping responses into themes. AI Study Reports will summarize a study once it has at least ten responses.

What you do outside the platform

The opportunity score itself is calculated outside Sprig, though it is commonly assumed to be built in. There is no built-in field for it. You export and calculate, generally in a notebook or a spreadsheet.

Kano categorization, which is often the bigger surprise. The two-question structure typically fields perfectly well, and the mapping from answer pairs to categories is yours to run.

Any needs-based segmentation, of any kind. Sprig documents no clustering, latent class, or k-means function. What the open-text tooling calls clustering is thematic grouping of text, and what cross-tab calls segments are the attribute values you supplied. Neither derives a segmentation from the rating data.

Bootstrapping, confidence intervals, and tiering are typically the largest piece of that outside work. All export-side.

A combined view of MaxDiff scores next to rating-scale results. There is no single report that shows both.

The practical shape of the workflow

Build and field in Sprig, export the response file, calculate in your own environment, and bring the readout back to the team.

The Claude integration removes the manual file handoff, because the data lands in the same place the analysis runs.

One constraint to plan around. MaxDiff item lists lock at launch. You can edit the wording of existing items afterward, and you cannot add or remove any.

A needs assessment that discovers a missing outcome mid-field has to be refielded rather than amended.

Alternatives and adjacent methods

What to reach for when a needs assessment is not quite the instrument.

MaxDiff when you need a defensible ordering and nothing else. It forces trade-offs, it produces one clean scale, and it does not mix units, which is generally why quantitative reviewers prefer it. What it will not tell you is whether the top-ranked item is currently well served.

Conjoint analysis when the decision involves bundles or trade-offs against price. Conjoint is the only method here that lets you simulate a configuration you have not built, and it is substantially more work to design.

Concept testing when you already have candidate solutions. A needs assessment sizes the space and concept testing evaluates your answer to it, and running them in the wrong order is a common and expensive mistake.

Satisfaction metrics sit adjacent rather than in competition. NPS, CSAT and customer effort each measure how you are doing on a relationship or an interaction.

None of them tells you what to build, and all of them are useful for detecting that a gap you closed actually closed.

Session replay and heatmaps operate two altitudes below this method. When the question is why a specific flow fails, behavioral data answers it faster than any survey will.

For the list-building stage, three Sprig templates map onto the qualitative half of this work. Identify user goals for the outcome discovery, and understand feature demand for the demand read.

Frequently asked questions

How many outcome statements should I field?

Field up to about forty outcome statements. Each is rated twice, so forty statements is eighty judgments, which is roughly where a respondent stops reading rows and starts pattern-filling them.

If your full inventory runs to a hundred or more, split it across two or three fieldings with a shared anchor block of eight to ten statements.

Can I use a five-point scale instead of ten?

A five-point scale works, and the ranking it produces will generally be coarser. The opportunity score is a difference between two ratings, so a five-point scale gives the difference only nine possible values and produces large numbers of ties.

Use ten points when the study's purpose is ranking, which it usually is, and check the question-writing guidance first.

What is a good opportunity score?

No published threshold exists, from either the method's author or his firm, so any source that hands you one without showing its derivation is repeating a convention.

The only defensible cut is a relative one taken inside your own study, fixed before you see the ranking.

Is "Jobs to Be Done" a trademark?

The phrase "Jobs to Be Done" is not currently a registered trademark, and the history is more interesting than the answer. An early United States application received a final refusal on 3 July 2014 and was abandoned for failure to respond.

A later application drew two oppositions, from Jarvin Ventures LLC and from The Rewired Group LLC. Both were sustained by default, which means the applications failed without any ruling on the merits of the underlying claim. The phrase is now in common general use.

What about "Outcome-Driven Innovation"?

Outcome-Driven Innovation is a registered trademark. The mark is owned by Strategyn, LLC, and its second renewal was recorded on 11 June 2025.

Strategyn marks it with the registered symbol in its own material and uses its other named concepts without one. No federal registration for "Kano Model" turned up on the register. Strategyn uses Opportunity Algorithm, Universal Job Map, and Outcome-Based Segmentation in its own material, and at least one of them has appeared elsewhere with a registered symbol attached, so treat the status of those three as unsettled rather than as clear either way.

Can I run Kano and outcome scoring in the same study?

You can run Kano and outcome scoring in one study, and it is generally a mistake. Kano needs two questions per item and outcome scoring needs two ratings per item, so combining them doubles an instrument that was already long.

The deeper problem is that the two methods want different item types. Kano wants discrete features and outcome scoring wants solution-free statements.

How large a sample do I need?

Around 300 for outcome scoring against a single population, which gives you roughly plus or minus 0.35 points of precision on the difference between two outcomes.

Around 200 per segment for Kano, because the categorical output needs a leading category clearly ahead of the runner-up. Segment-level reporting requires the full sample in each segment and not a share of the total.

How do I handle "not applicable" responses?

Treat a not-applicable response as a genuine missing value, excluded from that outcome's calculation, and never as a zero. Report the differential response rate across outcomes, because an outcome with unusually high non-response is telling you something about who has attempted the job.

Should I survey non-customers?

If your frame is the job rather than your product, yes. People who evaluated you and chose something else frequently hold the largest unmet gaps, and they are invisible to in-product fielding. Market and consumer insights is the shape of it, and Panels is the route.

Can Sprig calculate the opportunity score for me?

No, Sprig does not compute the opportunity score. Build and field the instrument in Sprig, export the responses, and calculate outside. The analysis chapter above gives a prompt sequence that does the calculation and then checks itself against a second method.

How often should I refield?

Once or twice a year for a given market, against the guidance on how often to survey. Outcome importance moves slowly and satisfaction moves faster, so refielding the satisfaction block alone is a reasonable mid-year check.

Kano categories migrate as expectations rise, so date every Kano study and avoid comparing categories more than eighteen months apart.

Two segments produced opposite rankings. Which one is right?

Both segment rankings are typically right, and the single company-wide ranking you were about to present is the artifact.

When two segments disagree by more than a handful of rank positions on an important outcome, that is a finding about your market structure and it usually deserves more attention than the pooled ranking does.

The bottom line

A customer needs assessment is a gap measurement dressed as a survey. Importance on one side, satisfaction on the other, and the distance between them as the reason to act.

The arithmetic that combines the two is a heuristic and not a measurement, and the people who object to it on unit-mixing grounds are correct.

That objection does not make the method useless. It makes the score a sorting device that should always travel with the two numbers it came from.

Do the qualitative work first. Write statements with a direction, a unit, an object, and a context.

Field importance before satisfaction, in separate blocks, to at least three hundred people. Exclude blanks rather than zeroing them. Tier the ranking whenever the sample cannot support an order.

And set your own threshold, from your own distribution, before you see the results.

The teams that get value from this method are not the ones with the cleanest arithmetic, or the ones with the best continuous insight system. They are the ones who arrive at the readout with a proposed mechanism written next to each of the top ten gaps.

Back to top
Solutions
Experience measurementStrategic & foundational discoveryJourney & behavioral researchMarket & consumer insightsConcept & prototype testing
Agents
DesignFieldAnalyzeSynthesize
Deploy
EmailPanelsWeb apps and websitesMobile app
Pricing
Community
EventsBlogGuides
CustomersIntegrationsCompare
Company
About usCareersService agreementPrivacy policyData addendumSystem status
Socials
LinkedInX