Solutions
Experience measurement
Track sentiment and KPIs with AI-driven gap analysis
Strategic & foundational discovery
Uncover market whitespace with AI-led foundational studies
Journey & behavioral research
Connect user actions to motivations across the lifecycle
Market & consumer Insights
Understanding markets, audiences, & opportunity
Concept & prototype testing
Test designs and prototypes with rapid feedback
Agents
Design
Structure rigorous studies
Field
Run adaptive studies at scale
Analyze
Get statistical analysis you can trust
Synthesize
Turn results into research reports
Deploy
Email
Reach external audiences with native deliverability
Panels
Recruit from 300K+ verified participants
Web apps and websites
Embed studies in web experiences
Mobile apps
Run studies in iOS and Android apps
Customers
Community
Events
Join curated gatherings shaping the future of research
Blog
Insights on integrating AI into research craft
Book icon
Guides
Ultimate playbooks for enterprise survey research
Pricing
Sign in
Book a demo
Sign in
Book a demo
Guide

How to Run: MaxDiff Study

October 2, 2026

By The Sprig Team

Example H2
Example H3
Example H4
Example H5
Example H6

Introduction

MaxDiff shows a respondent a handful of items at a time and asks which is best and which is worst. Repeat that a dozen times with different handfuls, and the pattern of choices produces a ranking of the whole list.

The method exists because rating scales often fail to discriminate on a long list. Ask someone to rate twenty features on importance and the ratings tend to bunch near the top, because nothing costs anything. That is the standard argument for forced choice rather than a measured finding, and the critique chapter returns to how well it holds.

A forced choice makes the respondent give something up, and the giving up is what produces the signal.

Use it when you have a long list and need to know what to build first. Eight to twenty-four items is the workable range, and Sprig caps the question at twenty-four.

Run it unanchored by default, and reach for an anchored design only for the specific reason below. The scores are purely relative, which means the top item in a list of twenty bad items still ranks first.

If you need to know whether anything on the list is good in absolute terms, that is a different design and it takes an extra question.

The two things most likely to sink a study are both upstream of the analysis. A list that was never reduced properly, so the items overlap and split each other's votes.

And items written at different levels of specificity, so respondents pick the concrete one over the abstract one regardless of what they actually want.

This guide covers the design, the fielding and the sample. The MaxDiff analysis guide covers what happens to the data afterwards.

What MaxDiff actually is

MaxDiff is a form of best-worst scaling, and the two names are used interchangeably in the literature. The respondent sees a subset of your items, picks the most and least of whatever you asked about, and the exercise repeats with a different subset until every item has been seen several times.

Because every choice is a comparison within a set, the method never asks anyone for an absolute judgement, which is often misread as a shortcoming. It only ever asks which of these is more, which is both the method's constraint and the reason it works.

That is the whole mechanism, and it is why the output is a ranking rather than a set of scores you can interpret on their own.

The origin, which most write-ups get wrong

The method comes from Jordan Louviere and George Woodworth, in a 1990 University of Alberta working paper.

The paper is unpublished and does not appear to be retrievable anywhere, and its title is rendered slightly differently across the secondary citations that reference it, so treat it as an unpublished working paper rather than as a document you can go and read.

The first published application is Finn and Louviere in 1992, in the Journal of Public Policy and Marketing, volume 11 issue 2, pages 12 to 25, applying best-worst scaling to public concern about food safety.

A number of write-ups commonly call that paper the origin of the method, which overstates it. The Cambridge volume credits both papers together with developing best-worst scaling, so the honest framing is that the working paper came first and the 1992 paper is where the method reached print.

The canonical modern treatment is Louviere, Flynn and Marley's 2015 book, Best-Worst Scaling: Theory, Methods and Applications, from Cambridge University Press.

That book describes the two earlier papers as having together "developed the best-worst scaling method and developed and evaluated a probabilistic model of their Case 1 data."

Case 1 is the version this guide teaches, meaning a list of items with no attributes attached. Cases 2 and 3 handle profiles and attribute levels and start to overlap with conjoint.

When to run a MaxDiff

The clearest case is generally a prioritisation argument that a rating scale has already failed to settle. Everything scored important, the list did not move, and the team is back where it started.

It typically fits a feature backlog, a set of candidate messages, a list of benefits, a set of use cases, or anything else where the question is which of these matters most and the list is longer than a person can hold in their head.

It fits well ahead of a roadmap cycle, where the deliverable is an order rather than a decision about one thing.

And it fits when you need a single number per item that different teams will read the same way. A ranking on one scale settles more arguments than twenty separate ratings do.

Where it does not fit is anywhere the answer needs a magnitude rather than an order. MaxDiff will tell you that onboarding speed beats reporting depth.

It will not tell you by how much in any unit anyone outside the study can interpret.

When not to run a MaxDiff

Four situations, each with the instrument that is actually right.

You need to know whether anything on the list is any good. Unanchored MaxDiff scores are purely relative, and they are relative to the specific list you fielded. Run twenty mediocre items and one of them still comes first, which is frequently how a weak list survives review. This is the limitation most likely to change what a reader does, because a ranked list looks like a verdict and is not one. The fix is an anchored design, which adds a separate question asking whether each item is a must-have, a nice-to-have, or not important, or a plain absolute rating alongside the MaxDiff. Note that Sprig has no anchored MaxDiff question type, so the anchor is a separate question you build, and Sprig's own MaxDiff analysis documentation describes that pattern.

The four limits

You need to value a bundle or a configuration. MaxDiff scores items one at a time and has no way to represent a combination. Two items that are each mid-ranked can often be worth a great deal together, and the method cannot see it. Run choice-based conjoint instead, which is built to value configurations and can simulate one you have not shipped.

You need to know what people will pay. MaxDiff produces preference, not willingness to pay, and a price item dropped into an item list produces a nonsense comparison between a price and a feature. Gabor-Granger is the instrument for a price point, and Van Westendorp for a range.

You need to know why an item ranks where it does. A forced choice records the choice and nothing else. There is no reasoning in the data because none was collected. Add an open-text probe on the best and worst picks and theme it, which Sprig's open-text analysis does natively, or run depth interviews.

There is a fifth case that is a scoping problem rather than a method problem.

A very short list can be ranked directly, and below roughly eight items the design overhead generally stops paying for itself. That figure is a working convention rather than a sourced threshold.

Design, and the anchored fork

One decision shapes everything else, and it is worth taking deliberately rather than by default.

Unanchored MaxDiff produces a relative ordering and nothing more, which is enough for most decisions and not for all of them. Anchored MaxDiff adds a reference point so you can say which items clear a bar rather than only which items beat which.

Default to unanchored. Most prioritisation questions are generally relative. You have capacity for three things and you want to know which three, and whether the fourth is objectively good does not change the decision.

Anchor when the answer might be none of them. If the plausible conclusion is that the whole list is weak and the team should go back to discovery, an unanchored design cannot reach that conclusion and will hand you a confident ranking of things nobody wants.

The analysis guide covers anchored designs and how the estimation changes.

What matters at design time is that there is no anchored MaxDiff question type, so anchoring means adding your own question, typically a three-point must-have, nice-to-have, not-important item per item on the list, which is a substantial addition to the instrument for a list of any size.

There is an intermediate option worth knowing, and it costs considerably less than a full anchor. Run the MaxDiff unanchored, then ask a single absolute question about the top few items only, once the MaxDiff has told you which those are.

It will not anchor the whole scale, and it will tell you whether the winner is worth building.

What anchoring actually costs

Worth being concrete about the trade, because the design chapters on most vendor pages present anchoring as a free upgrade.

An anchored design on a twenty-three item list means twenty-three additional judgements, one per item, on top of the fourteen MaxDiff screens that list needs at five per screen.

That roughly doubles the length of the instrument for a study that was already long, and the added questions are the flat, undifferentiated kind that MaxDiff exists to avoid.

The compensation is that you can then say which items clear a bar rather than only which beat which, and you can compare across waves even when the list changes, because an absolute reading does not depend on what else was on the list.

Decide by asking yourself one question before you commit to the longer instrument. If every item on your list turned out to be a must-have, or none of them did, would the team do something different?

If yes, anchor, and budget for the extra length. If the decision is which three of these, the anchor is paying for an answer nobody will use.

Building the item list

This is typically where MaxDiff studies fail, and it is the part the analysis cannot repair.

Sprig's MaxDiff question accepts a minimum of four and a maximum of twenty-four items. Most real prioritisation lists commonly arrive longer than that, so reduction is a design step rather than an afterthought.

The reduction protocol

The list itself should come from customer language rather than a planning document, which is what a voice of customer study or a round of user goal interviews produces.

Start by writing down what the list is a list of. One sentence. Candidate features for the next two quarters. Reasons a trial converts. Benefits worth putting in the headline. If two items belong to different sentences, they belong in different studies, and combining them produces a ranking nobody can act on because the items are not alternatives to each other.

Remove anything already decided. Items the team will build regardless do not need ranking, and leaving them in wastes exposure on a question nobody is asking. They will also usually rank first, which makes the study look like it confirmed the plan.

Collapse near-duplicates, and be aggressive. Two items describing the same underlying thing split the votes between them and both rank lower than the concept deserves. This is frequently the single most common way a real priority disappears from a MaxDiff result. When in doubt, merge the pair and use the broader wording of the two.

Cut anything that is a solution to another item. If one item is an outcome and another is a mechanism for reaching it, the two are not alternatives and the respondent is being asked an incoherent question.

Split anything compound. An item joined by and is two items. Respondents who feel differently about the halves have no way to express it, and the item's score means nothing.

If you are still over twenty-four, run a pre-study. An open-text or multi-select screening question on a smaller sample will tell you which items nobody mentions unprompted, and Sprig's understand feature demand template is a reasonable shape for it. Drop those. Do not cut by internal vote, which reintroduces exactly the opinion the study was meant to replace.

A worked reduction

A product team arrives with thirty-eight candidate features for the next two quarters. Here is what the protocol typically removes.

Four items fail the sentence test, because they are infrastructure work with no customer-visible outcome. They are real work and they are not alternatives to the other thirty-four, so they leave.

Seven are already committed. They stay on the roadmap and they leave the study without argument.

Nine items collapse into four, and this is where most of the reduction happens. Three separate items about export, one about scheduled exports, one about export formats and one about exporting filtered views, become a single item about getting data out on a schedule in the format the reader needs.

Two separate items about notification timing become a single item about when alerts arrive. Four items about permissions become two, because the underlying distinction between sharing and administration turned out to be real.

Two are mechanisms for other items on the list, and they leave.

Three are compound and split into six, which is the one step that adds items rather than removing them.

That arrives at twenty-three, inside the cap, with the concepts intact. The important thing about that sequence is that the only judgement call was the permissions merge, and everything else followed a rule.

What to do with the overflow

A list that genuinely will not reduce below twenty-four is telling you something. Usually it is two studies wearing one coat, and the split is often obvious once named.

If the overflow is real, split by the sentence test above and run two studies against separate samples, with a few shared items appearing in both.

The shared items let you see whether the two samples rank common ground the same way. They do not merge the two scales into one, and you should not report them as though they do.

Writing the items

Four failure modes make a MaxDiff unreadable, and all four are invisible until the data comes back looking strange.

Mismatched specificity. "Faster performance" and "Reports that export to CSV in under five seconds" are not comparable objects. Respondents typically pick the concrete one, because it is easier to evaluate, and you learn about item construction rather than preference. Write every item at the same altitude, and check the list end to end for the one that drifted.

Compound items. Anything with an and in it. Split it into two items or cut it, and do not leave it as written.

Negations. "Fewer errors during checkout" asks the respondent to evaluate an absence, and a best-worst choice between a presence and an absence is cognitively harder than it looks. Rewrite the item in the positive, describing the state the customer wants rather than the failure they avoid.

Unparallel grammar. A list mixing noun phrases, verb phrases and full sentences is generally harder to scan, and scanning is what respondents do on the fifth screen. Pick one grammatical form and hold it across every item on the list.

Four rewrites, one per failure mode:

| Problem | Before | After | |:----------------------:|:------------------------------------:|:------------------------------------------:| | Mismatched specificity | Better performance | Dashboards that load in under two seconds | | Compound | Faster exports and better formatting | Exports that finish without waiting | | Negation | Fewer errors when importing a file | Imports that succeed on the first attempt | | Unparallel grammar | Scheduling | Reports that send themselves on a schedule |

Note what the specificity rewrite does. It does not make the item longer for its own sake, it puts the item at the same altitude as its neighbours, and if the rest of the list is written at the level of "better performance" then the fix runs the other way.

Beyond those, keep items short enough to read at a glance. A respondent seeing five items per screen across fifteen screens reads your item list many times over, and length compounds.

Avoid brand and vendor names inside items unless the study is about brands. An item naming a competitor becomes an item about that competitor.

And test the list on a handful of people before fielding. Ask them to say each item back in their own words.

Anything that comes back changed will often be read differently by different respondents, which is a measurement problem rather than a matter of taste. The general guidance on writing survey questions applies on top of this.

Items per screen, screens per respondent, and exposure

Three numbers determine the design, and they are not independent. The item count, the items shown per screen, and the number of screens together fix how many times each item is seen.

That exposure count is generally the quantity that matters, because an item seen once has been compared against four others and an item seen four times has been compared against sixteen.

The arithmetic that connects the three is one line, and it is worth computing before you build anything.

exposure = (screens x items per screen) / total items

screens  = ceil(target exposure x total items / items per screen)

What the sources actually recommend

Sawtooth Software publishes the only concrete numbers in general circulation, and they are vendor technical guidance rather than peer-reviewed findings, which is worth saying plainly.

Their design documentation recommends "asking as many sets (questions) per respondent such that each item has the opportunity to appear from three to five times per respondent," and gives the rule with its governing condition attached: "For best results (under the default HB estimation used in MaxDiff), we suggest at least the following number of sets: 3K/k where K is the total number of items in the study, and k is the number of items displayed per set."

Keep that parenthesis in view. It is an estimation-specific recommendation, and Sprig does not use that estimation method in product, so treat 3K/k as a floor for the design rather than as a property of Sprig's own score.

On items per screen they recommend "displaying either four or five items at a time," and they are specific about why not more: "Research using synthetic data suggests that asking respondents to evaluate more than about five items at a time within each set may not be very useful in MaxDiff studies.

The gains in precision of the estimates are minimal when using more than five items at a time per set."

Sawtooth's MaxDiff System technical paper states the two balance properties a design should hold. "Frequency balance. Each item appears an equal number of times," and "Orthogonality. Each item is shown in the same sets with each other item an equal number of times." Both matter later, because a balanced design is what makes a simple count of best and worst picks behave.

What Sprig does by default, and why you should probably override it

Sprig's MaxDiff question offers two configuration modes. In Recommended mode, Sprig "will use half the total possible items as the number of items per set (up to a maximum of 8)" and then applies the formula "(total # items/items per set) * 3" to set the number of sets.

Custom mode lets you "configure the number of items per set and the number of times each item should be shown to each respondent."

Sprig's own MaxDiff documentation asks for these three numbers at analysis time, which is a good reason to record them at design time. Run the arithmetic on the Recommended mode and it clears the Sawtooth floor of three, landing on exactly three when the item count divides evenly by the items per screen and on up to about 3.4 otherwise.

It gets there by showing up to eight items per screen, which is above what Sawtooth recommends and past the point where they say the precision gains stop.

| Items | Sprig recommended | Exposure | Custom at 5 per screen | Exposure | |:-----:|:-----------------------:|:--------:|:------------------------:|:--------:| | 12 | 6 per screen, 6 screens | 3.0 | 5 per screen, 8 screens | 3.3 | | 16 | 8 per screen, 6 screens | 3.0 | 5 per screen, 10 screens | 3.1 | | 20 | 8 per screen, 8 screens | 3.2 | 5 per screen, 12 screens | 3.0 | | 24 | 8 per screen, 9 screens | 3.0 | 5 per screen, 15 screens | 3.1 |
/

Read the twenty-four item row, which is the one most studies land on. Sprig's default gives nine screens of eight items and eighteen clicks. Five items per screen gives fifteen screens and thirty clicks, at a comparable exposure of 3.1.

The default gives you fewer and harder screens, and the override gives you more and easier ones.

Use Custom mode and set items per screen to four or five. Be clear about what that rests on, though, because it is thinner than it looks. Sawtooth's case against more than five is statistical, that precision stops improving, rather than a finding about burden, and no study in the next chapter manipulates items per screen at all. What the next chapter does establish is that extra screens are cheap, which is the half of the trade with evidence behind it.

One constraint to plan around before you launch. Sprig documents that "you cannot change the configuration of MaxDiff item sets once the survey has been launched," so this decision is not revisable in flight.

Sample size

No source publishes a MaxDiff-specific sample-size rule with a stated derivation. Not the originating literature, not the Cambridge volume, and not the vendor technical papers, which give design minimums rather than respondent counts.

So derive it instead of copying one, and the exposure count from the previous chapter is what you derive from.

The quantity that determines precision is total observations per item, which is respondents multiplied by that item's exposure count.

Two hundred respondents at three exposures gives six hundred observations per item. One hundred respondents at six exposures gives the same six hundred, from half the sample and twice the burden.

That trade is the real decision, and it is why sample size and design cannot be settled separately.

For a single population and a straightforward ranking, a few hundred respondents at an exposure of three or more is typically where studies sit comfortably. Treat that as the working range rather than as a rule, because nothing published derives it.

Segment reporting is commonly where teams get into trouble. A segment ranking is its own estimation problem and needs its own sample, not a share of the total.

Four segments reported separately need four full samples, and quotas are how the small ones actually fill.

One numeric floor does appear in Sprig's own analysis guidance, and it is worth reading precisely: the analysis guide carries a fill-in-the-blank threshold in its prompt, with thirty offered as the example, for flagging segments too small to model separately.

That is a placeholder in a template rather than a recommendation, and a modelling-stability guardrail in any case. Thirty respondents is far below what you would want before reporting a segment ranking.

Audience and screening

Screen on the behaviour the items are about, not on product usage, since the two increasingly diverge.

If the item list is candidate features for a reporting workflow, the population is people who do that workflow, including the ones doing it in a spreadsheet or in a competitor's product.

People who happen to have opened your app are generally a different and less useful frame.

A screening question asking whether the respondent has performed the relevant task recently will generally produce a cleaner frame than any behavioural trigger will.

Carry your segment definitions in as attributes before you launch, because attributes not attached at response time cannot be reconstructed from the export afterwards. Those attribute values are what cross-tab analysis will cut the results by.

If the frame extends past your own customers, Panels is the route. Non-customers frequently rank a list differently from people already invested in your product, and that difference is often the finding.

Fielding

A MaxDiff is not a contextual microsurvey, and the delivery choice should reflect that.

Fifteen screens of forced choices fielded as an in-product intercept will typically produce heavy abandonment and a completed sample skewed toward unusually patient people. Email delivery or a link survey sets the right expectation, and both are documented delivery options.

Use the in-product channel for recruitment instead. A short intercept that screens on the task and routes qualified respondents into the longer instrument generally completes better.

One fielding detail specific to this question type: Sprig advances automatically once both picks are made, with "no 'Next' button between sets."

That makes the exercise faster and it also means a misclick cannot be corrected, which is worth a line in the instructions.

Quality control

Forced-choice data fails differently from rating-scale data, and the usual checks do not transfer.

What to check

Straight-lining has no obvious analogue, because there is no scale to flatten. What replaces it is typically positional response, meaning a respondent who picks the top item as best and the bottom item as worst on screen after screen regardless of content.

Check the distribution of chosen positions per respondent against chance, and flag anyone whose picks cluster on position.

Speed is the other cheap check, and it is the one most likely to be worth automating. Set the floor from your own pilot rather than from a rule, taking roughly a third of the pilot median as the cutoff, and review everyone below it rather than deleting automatically. That fraction is a convention, so check what it actually excludes before applying it.

The more interesting check is internal consistency, and there is a published tool for it. Sawtooth's guidance on identifying poor respondents observes that "for MaxDiff studies in which each item appears at least three times per respondent, it's very difficult to obtain a high RLH fit statistic for MaxDiff surveys via simplification strategies."

Read that carefully, because the clause at the end is the point. A respondent answering by a shortcut cannot fake a high fit score, which is exactly what makes the statistic usable as a screen.

Sawtooth publishes cutoffs for designs at three appearances per item, and they vary by items per screen. At four items per screen the threshold is 0.336, and at five it is 0.269. They report those cutoffs catching roughly 80 percent of genuinely random responders while wrongly excluding under 2 percent of real ones.

Use the cutoff for your own items-per-screen setting. A threshold borrowed from a design with a different screen size will be wrong in one direction or the other.

Report your exclusions alongside the result, every time, rather than in an appendix nobody opens. State the starting count, the number removed and the rule, before anyone sees a ranking.

Bot detection covers the automated end, which matters more for link and panel fielding than for a recruited sample.

The pre-launch checklist

Almost nothing about a MaxDiff can be changed after launch, which makes the checklist unusually load-bearing. Work through it in this order, because several of the steps depend on the ones above them.

Confirm the list passes the sentence test, meaning every item is an alternative to every other item and they all belong to one question.

Confirm no two items describe the same underlying thing. This is typically the last merge you will regret not making.

Confirm every item is written at the same altitude, with the same grammatical shape, with no compounds and no negations.

Read the full list aloud in one pass, which surfaces problems that silent reading misses. Items that are hard to say are generally hard to scan, and respondents scan.

Set the mode to Custom and set items per screen to four or five. Do not accept the Recommended default without checking what it produces.

Compute the exposure count and confirm it is at least three. Screens times items per screen, divided by the item count.

Count the clicks you are asking of each respondent, which is screens times two, and know the number before a stakeholder asks for it.

Confirm the left label is positive and the right label is negative, because Sprig's calculation depends on it and reversing them inverts the result silently.

Decide now which score scale you will report, the in-product count score or a model-based share, and write it in the study plan.

Set the attributes you will cut by, because they cannot be attached retrospectively.

Set quotas if you intend to report any segment separately, sized for a full sample per segment rather than a share of the total.

Field to a handful of colleagues first and time them. Their median completion time is what sets your speed floor for quality control.

Then launch, because after this point the item list, the set configuration and the labels are all fixed.

Respondent burden, and what the evidence actually says

The intuition is that forcing many repeated choices tires people out and degrades the later screens. It is the question behind most survey design advice about length, and the evidence on it is not what the advice assumes. The evidence is more interesting than that, and it does not support a hard screen limit.

Where the fatigue concern came from

The original finding is Bradley and Daly in 1994, in Transportation, volume 21 issue 2, pages 167 to 184.

Studying stated-preference ranking and repeated pairwise choice, they reported that "the amount of unexplained variance is shown to increase as rankings become lower, and as the number of pairwise choices completed becomes greater."

That result set the expectation for two decades. The three best modern tests all push back on it.

What the modern tests found

Hess, Hensher and Daly revisited the question directly in 2012, in Transportation Research Part A, volume 46 issue 3, pages 626 to 644.

Their conclusion: "we provide strong evidence that the concerns about fatigue in the literature are possibly overstated, with no clear decreasing trend in scale across choice tasks in any of our studies." They also note in passing that "it is not uncommon to see surveys with up to twenty tasks per respondent in some areas."

Bansak and colleagues ran respondents through as many as thirty conjoint tasks, in Political Analysis volume 26 issue 1, pages 112 to 119, and found "detectable but quite limited increases in survey satisficing as the number of tasks increases," concluding that "researchers can assign dozens of tasks without substantial declines in response quality."

Bech, Kjaer and Lauridsen randomised respondents to five, nine or seventeen choice sets, in Health Economics volume 20 issue 3, pages 273 to 286.

They reported somewhat higher response variance in the seventeen-set group than in the five-set group, indicating that cognitive burden may increase beyond a certain threshold, alongside no differences in response rate or model fit. Their own conclusion is the one worth carrying: respondents proved capable of managing seventeen choice sets without problems, though the resulting preference estimates were context-dependent to some degree.

The one finding that should worry you

Savage and Waldman compared the same choice experiment fielded by mail and online, in the Journal of Applied Econometrics volume 23 issue 3, pages 351 to 371.

Their result: "mail respondents answer questions consistently throughout a series of choice experiments, but the quality of the online respondents' answers declines."

That is the finding most relevant to anyone fielding a MaxDiff on the web, which is everyone reading this. The reassuring results above come substantially from contexts with more respondent commitment than a link survey has.

What follows

Two things, and the second is a negative worth stating.

No methodological guideline and no vendor specification publishes a maximum number of screens. The ISPOR task force on choice-experiment design, in Value in Health volume 16 issue 1, states the trade-off without a number: "Statistical efficiency is improved by asking a large number of difficult trade-off questions, while response efficiency is improved by asking a smaller number of easier trade-off questions."

If you find a page asserting a hard screen cap, it does not have a source.

And every study above is choice-based conjoint or a discrete choice experiment. No peer-reviewed study appears to manipulate the number of MaxDiff screens specifically and measure what happens.

The evidence transfers by analogy, which is reasonable and is not the same as direct evidence, and a guide that presented it as direct would be overclaiming.

Timing and cadence

Run a MaxDiff when a prioritisation decision is actually open, and not on a schedule.

Preference across a feature list typically moves more slowly than teams expect, and refielding an unchanged list against the same population two quarters later will mostly return the same ordering with noise on top. The noise will nonetheless get read as movement by somebody in the room.

What does justify a refield is a changed list. A new candidate feature, a shipped item that leaves the list, or a strategic shift that makes a different set of items relevant.

That is also the awkward part, because changing the list breaks comparability with the previous wave.

The two positions cannot both be held: either the list is frozen and the tracker is valid, or the list evolves with the roadmap and each wave stands alone.

For most product teams the second is generally the honest choice.

Treat each MaxDiff as a decision input for one planning cycle rather than as a tracked metric, and use the general guidance on survey frequency to keep it from colliding with your continuous measurement programs.

The core calculation

Scores are estimated from the pattern of repeated best and worst choices across screens.

Two approaches to turning those choices into scores are in common use. A count-based score simply tallies how often each item was chosen best against how often it was chosen worst, relative to how many times it appeared.

A model-based score fits a choice model to the same data and returns utilities, which are usually converted into preference shares that sum to 100 across the list.

The model-based route is usually more informative, because it uses the full pattern of which items beat which rather than only the totals, and it supports confidence intervals. The counting route is faster and needs no estimation step.

That is as far as this guide goes on estimation. The MaxDiff analysis guide covers the rest.

Two score scales, and why your numbers will not match

This catches people often, it is entirely avoidable, and neither of the two pages a Sprig user is likely to read currently connects them.

The two quantities

Sprig's in-product MaxDiff report uses a count-based score. The question documentation gives the formula and the range: the score "ranges from 100 to -100: (# Best - # Worst) * 100 / total number of appearances."

The analysis guide, running the data through a fitted choice model, produces preference shares where each item's share is the exponentiated utility divided by the sum across items, expressed as a percentage that sums to 100 across the list.

These are two different quantities and no transformation maps one onto the other. The two scales have different origins and different geometry. A count score sits symmetrically around zero, where zero marks an item picked best as often as worst, and it runs to plus and minus 100 at the extremes. A share is bounded below at zero, where zero means nobody ever picks the item, and the shares across the list are forced to sum to 100.

So the average item lands at a count of zero on one scale and at 100 divided by the item count on the other, which is about 4.2 percent on a twenty-four item list. Those two markers describe the same item and look nothing alike, and every other point on the two scales is related by no fixed rule.

Three rules follow from that, and all three are easier to hold than to recover from.

Pick one scale at the start of the analysis and report only that one throughout. A deck carrying both will get a question you cannot answer.

Never subtract one from the other, and never place them in adjacent columns. The difference between a count score and a share is not a number that means anything.

Expect the two orderings to agree closely. Under a balanced design, where each item appears equally often and against every other item equally often, they generally will, and Sprig's own analysis guide reports exactly that in its worked test.

Which means the reason to fit a model is not a better ranking. It is that a fitted share is provably a probability, it can carry a confidence interval, and it can be estimated per respondent, so you can look at whether the list splits the audience.

Benchmarks

There is no cross-industry benchmark for MaxDiff scores, and there could not be one.

Scores are rescaled within a study to sum to a constant, or bounded within a fixed range in the counting case, so an item's score is a function of what else was on your list.

Add two strong items and everything else drops, with no change in what anyone wants.

That means two studies' MaxDiff numbers are not comparable even within the same company, let alone across an industry. A competitor's reported score for a feature is a property of their item list, not of their market.

The consequence for tracking is worth stating, because teams try it.

You can only compare MaxDiff results across waves if the item list is identical, and Sprig's post-launch lock on adding and removing items means a tracked MaxDiff has to be designed as a tracker from the first wave.

If you want an absolute reading that survives a list change, that is the anchored design from the design chapter, and it is a different study.

Reading the results

Read the cliffs

Start with the ordering and the gaps, not the scores.

What you are looking for is where the list breaks. A ranked list of that length usually has a cliff or two in it, and the cliffs are generally the finding.

The exact position of item nine against item ten is almost never the decision.

Report tiers rather than ranks wherever the gaps are small, and right-size the report around the decision it feeds. A ranked list invites a reader to fixate on adjacent positions the data cannot distinguish, and grouped bands communicate the uncertainty without anyone having to read an interval.

Cut by segment before you commit to a single ordering, which is what cross-tab analysis is for. A company-wide ranking that hides two segments wanting opposite things is worse than no ranking, because it looks decisive.

Where two segments disagree by more than a few positions on a high-ranked item, that disagreement is usually the more valuable result.

Then do the thing that makes the study actionable. For the top handful of items, write down what would have to be true to ship it.

Some will be build work, some will be things that already exist and are not discoverable, and some will be packaging. That third category is invisible in a ranking and frequently common in the answers.

And remember what the ranking is relative to. If the list was weak, the winner is the best of a weak list, and nothing in an unanchored result will tell you that.

What a result looks like

A twenty-three item study fielded to four hundred respondents, reported as the in-product count score, might open like this.

| Rank | Item | Score | Tier | |:----:|:------------------------------------------:|:-----:|:----:| | 1 | Imports that succeed on the first attempt | 38 | 1 | | 2 | Dashboards that load in under two seconds | 34 | 1 | | 3 | Exports that finish without waiting | 19 | 2 | | 4 | Reports that send themselves on a schedule | 16 | 2 | | 5 | Permissions that can be set per workspace | 14 | 2 | | 6 | A mobile view of the main dashboard | -4 | 3 |

The useful structure is the cliff between rank two and rank three, not the gap between one and two. Nineteen points separate the top pair from everything else, and four points separate rank three from rank five.

That shape says something a ranked list alone does not: there are two priorities, then a band of five or six roughly equivalent items, then a tail.

A roadmap conversation about the top two is productive, and a conversation about whether rank four beats rank five is not.

Note rank six sitting below zero. On the count scale a negative score means the item was picked worst more often than best, which is a real signal and a different thing from ranking last.

Handing the analysis off

The MaxDiff analysis guide is the companion to this page and it owns everything downstream of a fielded study.

It covers what a naive prompt gets wrong on MaxDiff data, the engineered prompt to use instead, choice-model estimation against raw counting, bootstrapped confidence intervals, anchored designs, why a z-score comparison is the wrong test for this data, and how large a gap has to be before it means anything.

Use it with the Claude integration reading directly from the study, or against a CSV export if you prefer to work locally.

Two things to carry across the handoff. Tell the model your actual design, meaning the item count, the items per screen, the number of screens and the resulting exposure count, because the estimation depends on all four and the prompt asks for them.

And decide which score scale you are reporting before you start, per the chapter above, because that decision is much more annoying to reverse after a deck exists.

What this guide owns and that one does not: the item list, the item wording, the design parameters, the sample, the screening and the burden question.

Those are all settled before a single response exists, and none of them can be repaired in analysis.

The critique you should know

Two objections, and the second one is the more serious. Every widely used research instrument eventually collects a literature like this, and the case against NPS is the version most product teams have already argued through.

Hypothetical bias

Respondents in a survey are not spending anything, and the literature on stated preference finds that people overstate what they want when the choice is free.

The two meta-analyses most often cited disagree instructively about the size. List and Gallet, in Environmental and Resource Economics volume 20 issue 3, found subjects "overstate their preferences by a factor of about 3."

Murphy and colleagues, in the same journal at volume 30 issue 3, found "the median ratio of hypothetical to actual value of only 1.35, and the distribution has severe positive skewness."

Both are right about different statistics, and the skew is the interesting part: most studies typically show modest bias and a few show enormous bias, so a mean overstates the typical case.

State the scope limit honestly, because it matters. Both of those are contingent-valuation economics, measuring willingness to pay against real payment. Neither studied MaxDiff. The concern transfers by analogy and the magnitudes do not.

There is also a structural reason the bias bites less here.

MaxDiff asks which of these rather than how much would you pay, and a relative comparison between two hypothetical items is generally less exposed to inflation than an absolute valuation of one is.

The item list is the study

The sharper objection is the one this guide has generally spent two chapters on. MaxDiff can only rank the items you wrote, and it produces a confident, precise-looking ordering whether or not the list was any good.

There is no diagnostic in the output that tells you an important item was missing. A ranking of twenty-four items looks exactly the same whether the twenty-fifth item, the one people actually wanted, was omitted or never existed.

That is not a flaw you can fix with a better estimator. It is an argument for doing qualitative work before the list is frozen, and for the reduction protocol rather than an internal vote.

What survives

The counter-argument is strong and it comes from a hostile source.

Chapman and Callegaro reviewed the Kano method critically at the 2022 Sawtooth Software Conference, and their assessment is blunt. Kano, they write, "gives a 'compelling' answer to questions about features, but it is impossible to know whether it is a correct answer."

It will tell a story, they add, and quite possibly an incorrect one. Their objection is to the quality of Kano's survey items and to the theory behind its scoring.

What they recommend instead is conditional and specific. Where a team has two dimensions of interest, importance and satisfaction being their first example, they "often recommend a MaxDiff for preference plus a Likert type rating scale for the other dimension."

Weigh that correctly rather than overselling it. The paper appeared at Sawtooth Software's own conference, and Sawtooth is the leading vendor of MaxDiff software, so this is not an independent endorsement. It is conference proceedings rather than peer-reviewed journal work, and Kano is an academic method from 1984 rather than a vendor invention.

What it does establish is narrower and still useful. Two survey methodologists auditing a widely used method for reliability reached for forced choice as the more dependable alternative, and their complaint about Kano concerns item quality and scoring rather than anything MaxDiff shares. It does not make MaxDiff immune to the item-list problem above.

Common mistakes

Leaving near-duplicates in the list. Two items describing the same thing split the vote and both rank below where the concept belongs. The finding vanishes and nothing in the output flags it.

Mixing levels of specificity. Concrete items beat abstract ones on readability alone. You measure item construction and report it as preference.

Accepting the default items per screen. Sprig's Recommended mode goes up to eight items per screen, above what the design literature supports. Use Custom mode and set four or five.

Flipping the labels. Sprig's documentation is explicit that "the left label must be a positive descriptor for Sprig's Best-Worst calculation to be accurate." Swap them for a design reason and the scores invert silently, which is the worst kind of error because the output still looks plausible.

Launching before the list is final. After launch "you can only edit the existing items in a MaxDiff question. You cannot add or remove items," and the set configuration is locked too. A forgotten item means refielding.

Reporting the score as a magnitude. An item at 42 is not twice as wanted as an item at 21. The scale supports an ordering and the gaps between neighbours, not ratios.

Comparing across studies or waves with different lists. Covered in the benchmarks chapter and repeated here because it is the most common misuse after the duplicate problem.

Treating a relative ranking as a verdict on the list. An unanchored MaxDiff cannot tell you that everything on the list is weak. If that is a live possibility, anchor.

Synthetic respondents

Do not field an item list to simulated respondents and report the ranking as preference.

The specific failure here is that a model asked to pick best and worst from a list will produce a coherent, plausible ordering derived from the language of the items and from patterns in its training data.

Real respondents typically produce messier data, and the mess is where the segment structure lives.

Where simulation genuinely helps is on the instrument. Ask a model to state what it understands each item to mean, and any item it reads differently from your intent is an item your respondents will also misread.

That check costs ten minutes and generally catches the specificity problem.

How Sprig supports MaxDiff

MaxDiff is a native question type in Sprig. The constraints are real and worth knowing before you design.

It is "only available on the Enterprise plan and for Surveys," with a documented minimum of four and maximum of twenty-four items. There is no in-product MaxDiff.

The item set configuration supports both a Recommended default and a Custom mode where you set items per screen and the number of appearances per item directly. The design chapter above argues for Custom.

Sprig computes a count-based score in product, ranging from 100 to negative 100.

It does not fit a choice model, and the in-product report carries no significance testing, so any estimation, interval or test happens in the connected AI client or on the export. Attributing the model fit to Sprig itself would be a factual error.

There is no anchored MaxDiff question type, so an anchor is a separate question you build alongside it.

Open-text analysis on a best-and-worst probe is the strongest complement here, because it converts the reasoning MaxDiff cannot capture into themed output inside the same study.

Export is CSV only. The Claude integration reads the study directly, which skips the export step.

Alternatives

Conjoint analysis when the decision involves bundles, configurations or price trade-offs. Conjoint can simulate something you have not built, and MaxDiff cannot.

Gabor-Granger when the question is a price point rather than a preference ordering.

A rating scale when you need absolute levels and the list is short enough that discrimination is not the problem. Rating scales are frequently maligned in MaxDiff write-ups, including this one, and they answer a question MaxDiff cannot.

Concept testing when you have candidate solutions rather than a list of attributes to rank.

And a needs assessment when the question is not which of these but how well are we currently doing on each, which is a different comparison entirely.

Frequently asked questions

How many items should a MaxDiff have?

Between about eight and twenty-four, once the list has been reduced against something like the prioritise feature development framing. Below eight, a direct ranking is generally simpler and cheaper. Above twenty-four you cannot field it in Sprig, and the exposure arithmetic means each item is seen fewer times for a given number of screens.

How many items should appear on each screen?

Four or five. Sawtooth's design guidance recommends that range and notes that "the gains in precision of the estimates are minimal when using more than five items at a time per set."

Sprig's Recommended mode will go to eight, so use Custom mode to override it.

How many screens should each respondent see?

Enough for each item to appear at least three times. The formula is three times the item count divided by the items per screen, so twenty-four items at five per screen needs fifteen screens.

Is there a maximum number of screens?

No published source gives one. The choice-experiment literature tests up to thirty tasks without substantial quality loss, though the one study comparing modes found online respondents degrade where mail respondents do not.

How many respondents do I need?

No source publishes a MaxDiff-specific rule. Derive it from observations per item, which is respondents times the exposure count, and treat a few hundred respondents at an exposure of three or more as a working range rather than a threshold. Segments each need their own full sample, which is what quotas are for.

What is a good MaxDiff score?

The question has no answer. Scores are relative to the list you fielded, so there is no external benchmark and there could not be one.

Why do Sprig's scores not match the ones from the analysis guide?

Because they are different quantities. Sprig reports a count-based score from 100 to negative 100. A fitted choice model reports preference shares summing to 100. Both are correct and neither converts into the other, so pick one and report only that one.

Can I run a MaxDiff as a tracker?

Only with an identical item list across waves, because a score depends on what else was on the list. Sprig locks the item list at launch, so a tracker has to be designed as one from wave one and cannot be converted into one later.

Can I add an item after launch?

No. After launch you can edit existing item text but cannot add or remove items, and the set configuration is locked as well.

The bottom line

MaxDiff works because it takes something away. A rating scale lets a respondent call everything important, and a forced choice does not.

That mechanism is reliable, and it is endorsed by people who typically spend their time dismantling other survey methods. What the mechanism cannot do is rescue a list that was badly built.

So spend your effort upstream. Reduce the list properly, write the items at one altitude, override the default and show four or five per screen, and check the exposure count before you launch, because none of it can be changed afterwards.

Run it unanchored unless the honest answer might be that none of these are good enough. If that is live, anchor, and accept the extra instrument.

Then read the result as an ordering with cliffs in it, not as a set of magnitudes, and hand the estimation to the analysis guide.

Back to top
Solutions
Experience measurementStrategic & foundational discoveryJourney & behavioral researchMarket & consumer insightsConcept & prototype testing
Agents
DesignFieldAnalyzeSynthesize
Deploy
EmailPanelsWeb apps and websitesMobile app
Pricing
Community
EventsBlogGuides
CustomersIntegrationsCompare
Company
About usCareersService agreementPrivacy policyData addendumSystem status
Socials
LinkedInX