Solutions
Experience measurement
Track sentiment and KPIs with AI-driven gap analysis
Strategic & foundational discovery
Uncover market whitespace with AI-led foundational studies
Journey & behavioral research
Connect user actions to motivations across the lifecycle
Market & consumer Insights
Understanding markets, audiences, & opportunity
Concept & prototype testing
Test designs and prototypes with rapid feedback
Agents
Design
Structure rigorous studies
Field
Run adaptive studies at scale
Synthesize
Turn results into research reports
Deploy
Email
Reach external audiences with native deliverability
Panels
Recruit from 300K+ verified participants
Web apps and websites
Embed studies in web experiences
Mobile apps
Run studies in iOS and Android apps
Customers
Community
Events
Join curated gatherings shaping the future of research
Blog
Insights on integrating AI into research craft
Book icon
Guides
Ultimate playbooks for enterprise survey research
Pricing
Sign in
Book a demo
Sign in
Book a demo
Guide

How to Run a Conjoint Analysis: Step By Step

September 21, 2026

By The Sprig Team

Example H2
Example H3
Example H4
Example H5
Example H6

Introduction

Running a conjoint analysis takes five decisions in order. Choose choice-based conjoint with a none option, because asking people to pick between whole configurations is generally closer to how they actually buy than asking them to rate features one at a time.

Pick three to six attributes with three to five levels each, and spend most of your design time here, since attribute selection is where conjoint studies fail rather than the analysis.

Set the task count from respondent burden and the number of parameters you need to estimate. Derive the sample from a power calculation instead of copying a threshold.

Then read the output as importance conditioned on the levels you tested, never as a property of the market.

Most conjoint content circulates a sample-size rule of thumb without saying it is one. The rule most often quoted comes from Sawtooth Software, and Sawtooth's own chapter on the subject calls it a rule of thumb in those words.

On Sprig specifically, conjoint is an Enterprise question type on link surveys only, under caps the design chapter sets out.

This guide covers:

  • Pick attributes and levels first, because that decision constrains everything downstream
  • Use choice-based conjoint with a none option, and word the none option as a real alternative
  • Derive sample size from power rather than from a circulated threshold
  • Read importance as a function of your design, never as a market fact
  • Hand the estimation to a connected AI client, because Sprig does not compute utilities

This guide covers the design and the fielding. Its companion, Conjoint Analysis with Claude, works a real study through the estimation, and the two are meant to be read together.

Platform capabilities and documentation cited here were verified against sprig.com and docs.sprig.com on September 17, 2026. No prices appear anywhere in this guide.

What conjoint analysis is

Conjoint analysis is a survey method that estimates how much each feature of a product contributes to a buying decision, by showing people whole configurations and recording which ones they choose.

The method decomposes a stated choice into part-worth utilities, one for every level of every attribute. Those utilities are what the analysis produces, and everything else is derived from them.

Where it comes from

Green, P. E. and Rao, V. R. (1971), "Conjoint Measurement for Quantifying Judgmental Data," Journal of Marketing Research 8(3), 355 to 363, introduced the approach to marketing.

The technique it borrowed from mathematical psychology was built to recover the contribution of individual components from judgments about combinations, which is exactly the problem a product team faces when it wants to know what a feature is worth.

What it is not

Conjoint is not a feature-importance survey. Asking people to rate how important each feature is produces a list where everything is important, because rating carries no cost.

The method exists to remove that problem, and a study that reintroduces it by also asking direct importance questions has spent its budget twice.

The one sentence that matters

A conjoint study measures trade-offs, not preferences. A respondent who says every attribute is important has generally told you nothing, and a respondent forced to give one up to get another has told you a great deal.

When to use conjoint analysis

Use conjoint when the decision in front of you involves choosing between configurations, when you can name three or more attributes that plausibly move the choice, and when your respondents can reasonably be asked to evaluate a bundle rather than a single item.

Three conditions that make it the right instrument

The first is a genuine trade-off. Conjoint typically earns its cost where adding one thing means giving up another, such as a price increase paying for a capability or a faster tier costing storage.

The second is a bounded attribute set. Three to six attributes is the working range, and a study that wants to rank twenty items is a different method rather than a larger conjoint.

The third is a decision that will actually be made. A conjoint nobody will act on is an expensive way to produce a chart, and the design effort is hard to justify against a simpler instrument.

The three questions it answers well

Which configuration wins against a named set of competitors. What a given attribute is worth relative to the others in the set you tested. How preference shifts when one level changes and everything else holds.

Those three are the honest claims. Anything broader is frequently an overreach, and the next section names which.

When not to use conjoint analysis

Four things a conjoint cannot tell you. Each names the instrument that answers the question instead.

It cannot tell you whether anyone wants the category at all

Conjoint measures trade-offs among the configurations you show. A respondent who dislikes every option typically still picks the least bad one, and that pick enters the model as a preference.

A none option detects this partially. It catches the respondent who refuses outright, and it commonly misses the one who is mildly unenthusiastic about the whole set and picks anyway.

The instrument that answers this is a concept test with an absolute purchase-intent read, or a behavioural test with a real action attached such as a signup or a deposit.

It cannot tell you what a single attribute is worth when attributes interact

Standard additive models assume attributes are separable, meaning the value of a fast tier does not depend on which storage level it comes with. Real bundles do not always cooperate.

Where two attributes plausibly interact, an additive model leaves the interaction unmodelled. In a balanced design the main effects usually survive that omission, and what fails is the prediction for the specific bundles where the interaction lives.

The instrument that answers this is a design with explicit interaction terms, which costs tasks and sample, or a smaller study scoped to the two attributes in question.

It cannot tell you which of many items matters most

Conjoint's attribute count is bounded by respondent burden, and in Sprig the fifteen-task maximum bounds it further. A list of twenty candidate features does not fit.

The instrument that answers this is MaxDiff, which is built for exactly this problem and scales to far longer item lists than a conjoint can carry.

It cannot tell you what people will actually pay

A price attribute inside a conjoint produces a price sensitivity estimate conditioned on the configurations you showed. That is not a willingness-to-pay curve, and reporting it as one is the most common misreading of a conjoint output.

The difference matters commercially. It tells you how price traded against the other attributes you tested, and nothing about what a customer would pay for a product configured differently.

The instrument that answers this is Gabor-Granger, which reads acceptance across a price ladder directly, or Van Westendorp where the question is about a price range rather than a point.

The test before you commission one

Write down the decision the study will inform and the options it will choose between. A conjoint commissioned without that sentence typically arrives as a chart in search of a meeting.

What kind of evidence this produces

A conjoint produces stated-preference evidence about a hypothetical choice. It does not produce revealed-preference evidence about a real purchase, and the gap between the two is the subject of the critique section.

Treating a conjoint simulation as a revenue forecast is where most of the disappointment with this method comes from.

Study design, and the fork that decides everything else

Run choice-based conjoint with a none option unless you have a specific reason not to. The alternative designs are legitimate and they answer narrower questions.

The fork

Choice-based conjoint shows the respondent a small set of complete configurations and asks which one they would choose. It mirrors a purchase, and it is typically the design most commonly used in commercial work.

Ratings-based conjoint, sometimes called traditional or full-profile conjoint, asks the respondent to rate or rank profiles one at a time.

It extracts more information per profile, and it asks the respondent to do something they never do in a store.

Adaptive designs adjust the profiles shown based on earlier answers. They are efficient with respondent time, and they complicate the analysis in ways that are hard to explain to a stakeholder who wants one number.

Choose choice-based unless you are estimating many attributes on a small sample, where the extra information per profile in a ratings design earns its awkwardness.

Why the none option is not optional

Without a none alternative, every task is forced, and the consequence is often invisible in the output. The model then estimates preferences conditional on buying something, and it cannot distinguish a strong preference from a reluctant least-bad pick.

Word it as a real alternative rather than as a refusal. "I would keep my current solution" is a choice a respondent recognises.

"None of these" reads as an opt-out from the survey, and it often attracts respondents who want the task to end.

The cost is that a none option consumes choice share and reduces the effective information in each task. Budget for it in the task count rather than dropping it.

Attribute independence, and the prohibited pair problem

Conjoint assumes the levels of one attribute can combine with the levels of another. Where a combination is impossible or absurd, the design has to exclude it, and every exclusion costs statistical efficiency.

A profile offering enterprise single sign-on at the free tier is not a trade-off, it is an error the respondent notices.

It also teaches them the profiles are not real, and that lesson generally carries into the rest of the survey.

Prohibited pairs are the standard remedy, and they should be rare. A design needing many of them is generally describing attributes that are not independent, and the fix is to restructure them.

What a choice task actually looks like

The respondent sees a question stem and a small set of complete profiles side by side, each one a combination of one level from every attribute, plus the none alternative.

A task built from the four-attribute example later in this guide reads like this.

Which of these would you choose for your team?

| | Option A | Option B | Option C | |:----------------------:|:--------------:|:---------------------:|:-------------:| | Price per seat | Middle tier | Lowest tier | Highest tier | | Storage per workspace | 1 TB | 100 GB | Unlimited | | Support response | Within 4 hours | Within 1 business day | Within 1 hour | | Integrations supported | 50 | 10 | 200 |

I would keep my current solution.

Two notes on that illustration. The price levels appear as tier labels only because this guide prints no prices, and a real instrument carries the figures. The none alternative is drawn outside the columns as a presentation choice, not as an answer to the open slot question above.

The respondent picks one and moves to the next task, where the levels are recombined. Fifteen of those is the platform maximum, and reading one takes longer than a stakeholder usually assumes.

Level balance and design efficiency

A design where every level appears roughly equally often, and where every pair of levels across attributes appears roughly equally often, extracts the most information from each task. That property is what design efficiency measures.

Perfect balance is rarely achievable once prohibited pairs enter, and it is typically unnecessary, and near-balance is generally enough. What matters is that no level is rare, because a level seen by few respondents is generally estimated poorly and its utility will move on very little evidence.

Task count, and where the ceiling actually sits

More tasks produce more choice observations, and each one costs attention. The published work here is unusually practical.

Orme's chapter reports that Johnson and Orme (1996), a Sawtooth Software paper that re-analysed 21 commercial choice-based datasets, "determined that having each respondent complete ten tasks is about as good at reducing error as having ten times as many respondents complete one task."

Read that as a strong argument for asking more of fewer people, and not as a clean ten-to-one exchange rate. The information in later tasks declines, and the paper is about how many tasks to ask rather than about how large a sample to field.

The limit is fatigue rather than arithmetic. Later tasks are typically answered more quickly and less consistently, so treat the documented platform maximum as a ceiling rather than a target.

What is fieldable in Sprig, and what is not

The Sprig constraints are documented and they are specific. A conjoint question is available to Enterprise teams on link surveys only. It requires at least three features, each with at least three levels.

Options per task caps at four. Tasks per respondent caps at fifteen.

| Design you want | Fieldable in Sprig | What to do instead | |:---------------------------------------------------------------------------------:|:------------------------------------------:|:-----------------------------------------------------------------------------------:| | Two attributes, or two levels on any attribute | No, three feature and three level minimums | Add a third, or switch to an A/B test | | Five or more options per task | No, four option maximum | Reduce to four, or split into two studies | | More than fifteen tasks | No, fifteen task maximum | Reduce attributes, or accept a less precise design | | Conjoint inside your product | No, link surveys only | Field the conjoint by link and keep the in-product channel for shorter instruments | | Conjoint on a non-Enterprise plan | No, Enterprise only | Use a simpler trade-off instrument, or MaxDiff where the question is prioritisation | | Three to six attributes, three to five levels, four options, ten to fifteen tasks | Yes | Field it |

One thing the documentation does not settle. It is not stated whether a none alternative consumes one of the four option slots or sits outside the count, and the answer changes whether a four-option task means four real profiles or three.

Confirm it in the editor first, because the wrong reading costs a profile per task across the whole study.

The link-survey restriction is the one that changes the most plans. A team that runs everything in-product will need a different distribution route for this study, and that generally means a panel or an emailed link to a known list.

Writing the instrument: attributes and levels

This is where conjoint studies fail. The analysis is mechanical and the estimation is solved, and a study built on the wrong attributes produces a clean model of the wrong question.

What makes an attribute real

An attribute belongs in the study when it is something you could actually change, when the respondent can evaluate it without being told what it means, and when a plausible person would trade something else to get it.

Attributes that fail the first test are the most common waste. "Brand reputation" is not a lever a product team pulls next quarter, and a study spending a third of its design space on it has bought an importance score nobody can act on.

Attributes that fail the second test are subtler. A capability the respondent has to have explained gets evaluated on the explanation, and the utility you recover belongs to your copywriting.

What makes a level real

Levels have to be concrete, mutually exclusive, and spaced across a range the business would genuinely consider.

Vague levels are typically the standard failure. "Fast support" against "faster support" asks the respondent to invent the difference, and each one will invent a different difference.

"Response within four hours" against "response within one hour" asks a question with one meaning.

Range matters more than most teams expect. A narrow range makes an attribute look unimportant and an implausibly wide one makes it look dominant, and neither result is about the market.

The number-of-levels effect, and why it is a trap

An attribute given more levels tends to show higher measured importance than the same attribute given fewer, independent of how much respondents actually care about it.

This is a documented artifact of the design rather than a finding about preference.

The practical defence is level balance. Keep level counts roughly equal, and where one attribute genuinely needs five against another's three, say so when you report importance.

The attribute and level selection worksheet

Run every candidate attribute through this before the design is locked. An attribute that fails any of the first four tests should be cut, and the study is usually better for it.

| Test | Question to ask | Cut if | |:------------:|:-------------------------------------------------------------------------------------------------------:|:----------------------------------------------------------------:| | Actionable | Could we change this within the planning horizon this study feeds? | The answer is no, or it needs another team's roadmap | | Self-evident | Can a target respondent evaluate this level without an explanation? | It needs a sentence of setup to be understood | | Tradeable | Would a plausible buyer give up something else to get the better level? | Everyone wants the better level and nobody would pay for it | | Independent | Can every level of this attribute combine with every level of the others? | More than one or two combinations are impossible | | Level count | Does this attribute carry a similar number of levels to the others? | It carries noticeably more, and the importance read will inflate | | Range honest | Is the best level something we would actually ship, and the worst something we would actually tolerate? | Either end is there to make the other look good |

The sixth test is the one teams argue about. Stretching a range to produce a clear result is a way of designing the answer, and a reader of the output has no way to detect it.

A worked attribute set

A team pricing a collaboration product might start with eight candidate attributes and end with four.

Price per seat survives, because it is actionable, self-evident, tradeable and independent. Storage per workspace survives for the same reasons, with levels set at three points the business would genuinely ship.

Support response time survives once the levels are made concrete, moving from fast against faster to within one business day, four hours, and one hour.

Integration depth survives after restructuring. Naming specific partner tools failed the independence test because several combinations were impossible, so it became a count of supported integrations.

Four were cut. Brand trust failed actionable. Security posture failed tradeable in this particular case, since the team's candidate levels were ones every respondent wanted and none would trade for. Where security features are genuinely gated behind a paid tier, as enterprise single sign-on often is, it passes the test and belongs in the study. Onboarding quality failed self-evident. Roadmap velocity failed actionable, self-evident and tradeable together.

That last group is the one teams fight hardest for, and cutting it is often what makes the study work.

How many attributes

Three to six is the working range for a choice-based design, and the Sprig minimum of three is a floor and not a recommendation.

Studies do run wider. The worked example in the analysis guide carries seven attributes, which is workable with enough tasks and sample and is more than most teams should attempt on a first study.

Below three the study is usually answering a question a simpler instrument answers faster. Above six the profiles get long, respondents start simplifying, and they commonly settle on one or two attributes and ignore the rest.

That simplification is not visible in the output. The model will still return utilities for every attribute, and some of them will describe a respondent who stopped reading.

Scale and scoring choices

A choice-based conjoint has no rating scale. The respondent picks, and the pick is the data, which removes a whole class of measurement problems that rating-based instruments carry.

What the output actually is

Part-worth utilities are estimated on an arbitrary scale with no natural zero. A utility of 1.4 for a level means nothing on its own, and it means something only relative to the other levels of the same attribute.

Importance is derived rather than measured. It is each attribute's utility range divided by the sum of all attribute utility ranges, expressed as a percentage.

Utilities are not comparable across attributes

A utility of 1.4 on price and a utility of 1.4 on storage do not mean the same thing, because each attribute's utilities are centred within that attribute.

What compares across attributes is the range, which is why importance is built from ranges rather than from raw utility values.

Why importance percentages travel badly

An importance score is a property of your design. It depends on which attributes you tested, which levels you chose, and how wide each range was.

Two studies of the same market with different level ranges will report different importances, and both can be correct. A number that behaves this way should never be quoted without the design that produced it, and it should never be compared to a competitor's.

The none option in scoring

The none alternative's utility is estimated alongside the others and sets the reference point for whether a configuration beats the status quo. A simulation that reports share among the configurations shown while ignoring the none share generally overstates demand.

Sample size and precision

Derive the sample from a power calculation. The thresholds circulating in practitioner content are conventions, and at least one of the most widely quoted is called a rule of thumb by the organisation that publishes it.

What the circulated rule actually says

The rule most often repeated comes from Sawtooth Software. In "Chapter 7: Sample Size Issues for Conjoint Analysis," Bryan Orme writes that Rich Johnson "has recommended a rule-of-thumb when determining minimum sample sizes for aggregate-level full-profile CBC modeling: set nta/c >= 500 where n is the number of respondents, t is the number of tasks, a is number of alternatives per task (not including the none alternative), and c is the number of analysis cells."

The same chapter adds that "it would be better, when possible, to have 1,000 or more representations per main-effect level."

One thing about that rule deserves saying plainly. Its own source characterises it as a rule of thumb, in those words, and the chapter presents no derivation for it.

On the definition of c, Orme is specific. For main effects it is the largest number of levels on any one attribute, and for two-way interactions it is the largest product of levels across any two attributes.

The worked example below uses the main-effects reading.

The derivation that does exist

de Bekker-Grob, E. W., Donkers, B., Jonker, M. F. and Stolk, E. A. (2015), "Sample Size Requirements for Discrete-Choice Experiments in Healthcare: a Practical Guide," The Patient 8(5), 373 to 384, DOI 10.1007/s40271-015-0118-z, gives a power-based sample-size approach for choice experiments.

The paper's formula is built from the significance level, the desired power, the asymptotic variance-covariance matrix of the parameter estimates, and the smallest effect size you need to detect.

That is a derivation in place of a convention, and it is the right basis for a defensible number.

The paper is direct about the alternatives. It states that "the disadvantage of using one of the rules of thumb mentioned in paragraph 2.2 is that such rules are not intended to be strictly accurate or reliable."

A worked example, and what the rule of thumb actually returns

Take a design with four attributes carrying four, three, three and three levels. Three options per task plus a none alternative. Twelve tasks per respondent.

The main effects to estimate are the sum across attributes of levels minus one, which is three plus two plus two plus two, so nine parameters.

Now run the circulated rule. With n times t times a divided by c set at 500 or more, where t is 12, a is 3 excluding the none alternative, and c is the largest number of levels on any attribute at 4, the arithmetic is 500 times 4 divided by 36, which gives 55.6 and rounds up to 56 respondents.

Take the stronger version of the same rule, the one asking for 1,000 or more representations per main-effect level, and it becomes 1,000 times 4 divided by 36, which gives 111.1 and rounds up to 112.

Fifty-six respondents is obviously not a study. That is the tell. The rule is a floor for aggregate-level estimation on a well-behaved design rather than a recommendation for how many people to field to.

It also says nothing about segments. Three customer segments each have to clear the floor independently, which puts the same design at 168 under the 500 version and 336 under the 1,000 version, before any allowance for screen-outs or quality exclusions.

And none of that arithmetic references the effect size you care about. That is the gap the power-based approach fills, and it is why the derivation is worth the extra effort.

What to do in practice

Name the smallest utility difference that would change your decision. That is the effect size, and without it no sample-size method can give you an answer.

For a sense of what this produces in practice, the worked study in Conjoint Analysis with Claude validates its segments against an 84-respondent floor derived for that specific design, which is a long way from a single number applied to every study.

Then set your significance level and power, count the parameters your design has to estimate, which is the sum across attributes of levels minus one, and work the calculation from there.

Where you lack the variance inputs, a pilot of 50 to 100 respondents supplies them. That is a different exercise from the 20 to 30 respondent soft launch described under fielding, which checks the instrument and not the variance.

Report the derivation alongside the number. A sample size a stakeholder can interrogate survives a challenge, and a threshold quoted from a blog post does not.

The honest position on precision

No conjoint-specific sample-size rule with a published derivation was located for this guide outside the discrete-choice-experiment literature above.

That absence is worth stating, because the figures circulating in vendor content are conventions repeated until they sounded official.

Audience, targeting and screening

A conjoint is only as good as the population that answered it, and the trade-offs a non-buyer reports are not the trade-offs that produce revenue.

Who should be in the sample

Restrict the sample to people who could plausibly make or influence the purchase. A conjoint fielded to a general audience returns a price sensitivity belonging to people who were never going to buy.

Where the decision is made by a committee, decide before fielding whether you are modelling the individual or the committee.

Those are different studies, and a sample that mixes evaluators and approvers produces utilities that average across two decision processes.

Screening without teaching the answer

Screen on category behaviour instead of stated interest. "Have you purchased in this category in the last twelve months" beats "are you interested in a product like this," which tells the respondent what you hope to hear.

Keep the screener short and keep the conjoint's attributes out of it. A screener naming the features you are about to test primes the respondent to notice them, and priming shows up in the utilities.

Sample size for the segments you will report

Decide the reporting cut before fielding. A study powered for an overall read and then sliced three ways at analysis time produces segment estimates that are frequently too imprecise to act on.

Each reported segment needs to clear the sample requirement on its own, which in practice means the segment plan drives the total rather than the other way round.

Where the sample comes from

A study about your own customers can be fielded to a list you already hold. A study about a market you do not yet serve needs a frame you do not own, which in practice generally means a panel.

Panel sampling brings its own problems. Professional respondents often complete a great many surveys, and the speed and consistency that makes them attractive to a provider is what makes their choice patterns unlike a first-time respondent's.

Set quotas on characteristics that plausibly move the trade-off, typically segment, company size or role, and not on demographics with no connection to the decision.

Fielding and delivery

Conjoint in Sprig is fielded by link. The documentation states that conjoint questions are available to Enterprise teams on link surveys only, which forecloses the in-product route that most product teams default to.

What the link constraint means in practice

A link survey has to be distributed rather than triggered. The routes are an email to a known list, a panel, or a link placed where your target population already is.

Each route changes who answers. An emailed link reaches your most engaged users first, a panel reaches people with no relationship to your product, and the same conjoint fielded both ways will not return the same utilities.

Decide which population the decision needs before choosing the route, instead of choosing whichever route is easiest to arrange and inheriting whichever population it produces.

Order effects and what the platform handles

Position bias is real in choice tasks. A respondent shown the same attribute in the same position in every task will commonly weight it differently than one who sees it move.

Sprig documents a conjoint setting to randomize the order in which features are displayed across respondents, which the documentation describes as mitigating position bias. Turn it on.

Randomization of page and question order, and of answer options in several question types, is documented separately. What is not documented anywhere is random assignment of a respondent to one of several conditions, so a design that needs each respondent to see exactly one of N versions has to be built another way.

A pilot is not optional here

Field 20 to 30 respondents before the full launch and read their records individually. A conjoint with a broken profile, an impossible combination, or a mislabelled level is often invisible in the editor and very obvious in the data.

The pilot also gives you a real completion time, which is the number to put in the invitation.

Length and completion

A conjoint is often longer than the respondent expects. Fifteen tasks at the four-option maximum is up to sixty profiles to read, and the respondent agreed to a survey and not to an exercise.

Say how long it will take, and be accurate. Understating it is the most reliable way to produce a break-off mid-task, which costs the whole respondent.

Timing and cadence

Conjoint is episodic. It answers a design question at a decision point, and re-running it on a schedule generally produces a chart that moves for reasons nobody can attribute.

Run one when the configuration decision is live, which generally means before a pricing change, a packaging revision, or a roadmap commitment large enough to be worth de-risking.

Re-run when the market or your offer has changed enough that the attribute set is stale. A conjoint whose attributes no longer describe the product measures a configuration you no longer sell.

Where a re-run is genuinely needed for a later decision, hold the attribute set and the levels fixed. Changing either makes the second study a new study, and comparing the two importance charts will produce a difference that is about the design rather than about the market.

Do not run one to track anything. Utilities are scaled within a study, so two conjoints fielded six months apart are not on the same scale and the difference between them is not a trend.

Quality control and data hygiene

Choice data hides bad responses better than rating data does. A respondent clicking the first option every time frequently produces a complete, plausible-looking record.

What to check before analysing

Straight-lining by position is generally the first check, and it depends on whether option position was randomized. Sprig documents randomization of feature order within a conjoint question, and documents nothing about randomizing which position a profile appears in, so confirm the behaviour before relying on this check.

Where position is randomized, count how often each respondent picked each position and flag anyone far from what random choice would produce. Where it is not, a respondent with a genuine preference will look like a straight-liner, and the check produces false positives.

Speed is the second. Compute the median time per task and flag the bottom few percent, then read a handful of those records before deciding what to do with them.

The none rate is the third and it runs both ways. A respondent who picked none in every task has told you something real about the category, and a respondent who never picked it may simply be avoiding the extra reading.

The check almost nobody runs

Include one repeated task, usually placed near the end and identical to an earlier one. A respondent who answers it differently has either changed their mind or was typically not attending, and the rate across the sample is a direct read on data quality.

Report that rate. A high inconsistency rate calls for a more cautious reading, and saying so is more useful than presenting the utilities as if they were clean.

Budget for it. One repeated task plus one or two holdout tasks consumes two or three of the fifteen the platform allows, which is a trade worth making deliberately instead of discovering at analysis.

What to do with flagged records

Decide the exclusion rule before you look at the results, and write it down. A rule applied after seeing who produced an inconvenient answer is not quality control.

Analysis: the core calculation

Part-worth utilities are estimated from the pattern of choices across tasks, typically by multinomial logit or a hierarchical Bayesian model that produces individual-level estimates.

Attribute importance is then derived. Each attribute's utility range, meaning its highest level's utility minus its lowest, is divided by the sum of those ranges across all attributes, and the result is expressed as a percentage.

Share of preference for a set of configurations is simulated from the utilities, most commonly by a logit rule that converts utilities into predicted choice probabilities.

One design consequence belongs here rather than in the analysis guide. A simple logit simulator carries the independence of irrelevant alternatives assumption, which means adding a near-duplicate configuration to a simulation takes share proportionally from everything else instead of mostly from the option it resembles, so the two similar options together end up with more share than the single option had.

That is the red bus and blue bus problem, and it is the reason a simulation loaded with similar configurations overstates them as a group. Keep simulated sets distinct, and treat share estimates for near-identical options with suspicion.

That is the whole calculation. The analysis guide covers execution.

Benchmarks and what a good result looks like

There is no cross-industry benchmark for conjoint utilities or attribute importances, and there could not be one. Utilities are scaled within a study, and importances are a function of the levels you chose to test.

A competitor's importance score for price is a property of their design, not of their market. Reproducing an uncited importance table from a vendor deck imports someone else's attribute ranges into your decision.

What to compare against instead

Compare configurations within your own study, commonly the only fair comparison. The simulator exists for exactly this, and the comparison it supports is the one the study was designed to answer.

Compare against the none option. Whether your best configuration beats the status quo is a more useful question than whether its share is high in absolute terms.

Compare against a holdout task. Set aside one or two choice tasks, exclude them from estimation, and check whether the model predicts them. The analysis guide does not cover holdout validation, so this one is yours to run.

That is an internal validity check you can actually run, and it is worth more than any external benchmark.

What good looks like

A good conjoint has a clear winner among the configurations the business would ship, a none share low enough that the category is viable, and a holdout prediction that beats chance comfortably.

A conjoint where every configuration is close is not a failed study. It usually means the attributes you tested do not drive the decision, which is worth knowing before you build them.

Interpreting and acting on the result

Read the utilities first and the importance chart second. The chart compresses the study into one picture, and the picture loses the level detail that the decision actually needs.

Reading level detail

The gap between adjacent levels is where the decision lives. An attribute with high importance whose top two levels are nearly tied is typically telling you the cheaper of the two is enough.

That pattern is increasingly common with price and with performance tiers. The importance chart will show the attribute as dominant, and the level detail will show that most of the dominance comes from avoiding the worst level instead of from reaching the best one.

Running the simulator honestly

Share-of-preference simulation is the one step neither Sprig nor the analysis guide performs for you, so budget for building it from the utilities yourself.

Simulate configurations you would actually ship, against competitors that actually exist. A simulation populated with implausible configurations returns a share number for a market that is not there.

Include the none option in every simulation and report its share. A configuration with 60 percent share among four options looks different when a third of respondents chose none of them.

A worked read

Suppose price comes back at 41 percent importance, support response at 24, storage at 20, and integrations at 15. The obvious reading is that this is a price-driven market.

Then look at the levels, because two different patterns produce the same 41 percent.

Where the utility gap between the middle and the lowest price is large and the gap between the middle and the highest is small, respondents want the cheapest option and are near-indifferent between the middle and top tiers. That reading says discount.

Where the gap between the middle and the highest is large and the gap between the middle and the lowest is small, respondents are punishing the top tier instead of chasing the bottom one. That reading says hold the middle tier and make the top tier optional.

The importance chart cannot distinguish them, which is why the level detail is the part worth reading twice.

What to hand the decision-maker

Give them the two or three configurations worth building, each with its predicted share, the none share alongside, and the level detail behind the ranking.

State the scope limit in the same document. The study says which configuration wins among those tested, under the ranges you chose, for the population that answered.

Running the analysis with Claude or ChatGPT

Sprig does not estimate conjoint utilities. The analysis runs in a connected AI client, and Sprig documents the recipe instead of performing the calculation.

The walkthrough lives in Conjoint Analysis with Claude, which works a real study end to end. Use it instead of rebuilding the prompt here.

What that guide covers

It estimates utilities with a discrete-choice model using dummy-coded attributes against explicit reference levels. It computes importance as each attribute's utility range divided by the sum of all ranges, in code at full precision.

It segments, in its worked case by purchase authority, and it validates the segment sizes before reporting them. It tests differences formally, dividing the gap by the combined standard error and reading 1.96 as the threshold.

And it requires two independent implementations to agree before a number is trusted, which is the verification step this guide's quality section keeps pointing at.

What it does not cover, and what that means for you

It does not cover share-of-preference simulation, and it does not cover holdout validation. Both are recommended earlier in this guide, so plan to run them yourself rather than expecting the walkthrough to carry you through them.

That gap is worth knowing before you promise a simulator to a stakeholder. Hold out your tasks at design time anyway, because a holdout you did not reserve cannot be recovered at analysis.

Two planning details

Importance is computed in code instead of in prose, and code execution has to be enabled for the verification step to run, which the documentation notes is on by default for Team and Enterprise accounts.

You can reach the data two ways. Export the responses as CSV, or connect the survey directly through the Sprig MCP so the analysis runs against live data. The MCP retrieves up to 1,000 responses per call, so a large sample arrives in pages.

The division of labour

This guide owns the design, the instrument, the fielding and the sample. The analysis guide owns the estimation, the importance calculation, the segmentation and the significance testing.

Where the two overlap, on what conjoint is and when to use it, read this one for the design decision and that one for the worked execution.

The critique you should know

The standing objection to conjoint is hypothetical bias. People often choose differently when the choice costs them nothing, and a stated choice in a survey costs nothing.

The two meta-analyses, and why they disagree

The most-cited evidence comes from contingent valuation in environmental economics rather than from conjoint, and that scope limit belongs in the same breath as the numbers.

List, J. A. and Gallet, C. A. (2001), Environmental and Resource Economics 20(3), 241 to 254, DOI 10.1023/A:1012791822804, report that subjects overstate their preferences by a factor of about 3.

Murphy, J. J., Allen, P. G., Stevens, T. H. and Weatherhead, D. (2005), in the same journal, 30(3), 313 to 325, DOI 10.1007/s10640-004-3332-z, report a median ratio of hypothetical to actual value of only 1.35, and note that the distribution has severe positive skewness.

Both are right, and the disagreement is instructive rather than embarrassing.

Why the skew explains the gap

The two figures come from different meta-analyses over overlapping but not identical study sets, so they are not two statistics of one distribution. Skew still explains much of the gap, because a small number of studies with extreme overstatement pull a mean well above a median.

That has a practical reading. For a typical study the bias is real and moderate, and the tail cases are where a stated-preference result goes badly wrong.

Which tail you are in is not random. Overstatement is generally worst where the good is unfamiliar, where the stated choice carries social desirability, and where no budget constraint is present in the task.

Murphy and colleagues also report a finding that matters directly to this guide's recommendation. They conclude that a choice-based elicitation mechanism is important in reducing bias, which is an argument for the design this guide recommends rather than against it.

What this means for your design

Include price in the attribute set, often the most effective correction. A trade-off that costs the respondent nothing invites the overstatement the literature documents, and a price attribute reintroduces the constraint.

Include a none option. It gives the respondent a costless way to decline, which is closer to their real alternative than a forced pick.

And discount the absolute numbers. A conjoint ranks configurations reliably and forecasts demand levels poorly, so the ranking is the output to act on.

The counter-critique

Conjoint's defenders have a case, and the review of record is Green, P. E. and Srinivasan, V. (1990), "Conjoint Analysis in Marketing: New Developments with Implications for Research and Practice," Journal of Marketing 54(4), 3 to 19, DOI 10.1177/002224299005400402.

Their summary judgment is that "in sum, the empirical evidence points to the validity of conjoint analysis as a predictive technique."

The paper carries one hard predictive number, and it is a single case. Reporting Benbenisty (1983) on AT&T's entry into the data terminal market, Green and Srinivasan write that "the simulator forecasted a share of 8% for AT&T four years after launch. The actual share was just under 8%."

That is one study, reported second-hand inside a review, and it should be presented as what it is. It is also worth noting that the review dates from 1990, which predates both the rise of choice-based designs and hierarchical Bayesian estimation.

The reliability number that is not a validity number

Green and Srinivasan, summarising a review by Bateson, Reibstein and Boulding, report that "the median reliability correlation is about .75."

That figure circulates widely as evidence that conjoint is accurate. It is not. Reliability is the stability of the estimates on repeat measurement, and validity is whether they predict anything.

A method can often be highly reliable and consistently wrong. Anyone quoting .75 as a predictive accuracy figure has confused the two, and the confusion is frequently common enough to be worth catching before it reaches a slide.

Where this leaves the method

Conjoint is good at ranking and poor at forecasting levels. Treat the simulator's shares as relative comparisons, not as demand estimates, and the method does what it is actually good at.

Common mistakes

Six that recur, in rough order of how much damage they do.

Designing the attribute set from what the team wants to hear rather than from what the business could change. Every downstream number inherits that choice, and no amount of analytical care fixes it.

Stretching an attribute range to produce a clear winner. The result is a designed answer, and a reader of the output has no way to see it.

Giving one attribute more levels than the others and then reading its higher importance as a finding. That gap is at least partly an artifact of the design.

Reporting importance percentages as facts about the market. They are facts about your design, and they do not transfer to another study or another company.

Dropping the none option to get cleaner data. It is cleaner because it no longer says whether anyone wants the category.

Reading a simulated share as a demand forecast. The share is relative to the configurations simulated, and the critique section explains why the absolute level is least trustworthy.

The mistake that is hardest to see

Running the study on a population that cannot buy. Every check above is about the instrument, and this one is about who answered.

A conjoint fielded through a panel for a product sold to enterprise IT will return trade-offs from people who have never held that budget.

The model fits, the chart renders, and the utilities describe a decision nobody in the sample makes.

The defence is a screen on category purchase behaviour and a hard look at the completed sample before analysis instead of after.

Pre-launch checklist

Run this in order before the study goes live.

  • Confirm every attribute passes the six selection tests
  • Confirm the level count is balanced across attributes, or note the imbalance in the reporting plan
  • Confirm the prohibited pairs are few, and that none of them is hiding an attribute that should be restructured
  • Confirm the none option is worded as a real alternative rather than as a refusal
  • Confirm the design fits the platform caps, four options and fifteen tasks
  • Confirm feature order randomization is enabled
  • Confirm a repeated task is placed for the consistency check
  • Confirm the exclusion rule for flagged respondents is written down
  • Confirm the sample size traces to an effect size and a power calculation
  • Confirm the screener does not name any attribute the conjoint tests
  • Confirm the stated survey length matches a timed pilot run

Synthetic respondents

A synthetic respondent can produce a plausible set of conjoint choices. Whether those choices carry the trade-off structure of a real population is a different question, and the answer is not yet settled.

The specific risk for this method is variance. A model asked to simulate many respondents typically produces choices that are more internally consistent than human ones, and conjoint estimation reads consistency as strength of preference.

That inflates utilities and narrows the apparent spread between segments. The output commonly looks more decisive than a human sample would, which is the direction of error most likely to be believed.

Use synthetic responses to pressure-test an instrument before fielding it. Do not use them to produce the utilities a pricing decision rests on.

How Sprig supports this method

Conjoint is a native question type in Sprig, fielded under the documented caps set out in the design chapter above. Those caps are the short version of what this platform will and will not build.

Feature order can be randomized across respondents, which the documentation describes as mitigating position bias. Random assignment of respondents to conditions is not documented, so a monadic design has to be handled outside the platform's randomization settings.

Panels supply a sample frame for a market-level study rather than a customer-level one, with targeting attributes documented in the hundreds. No incidence rate, minimum or maximum sample size, or fielding turnaround time is documented, so plan those with your account team rather than from the page.

For getting data out, response export is CSV, capped at 950,000 rows, and the download link expires after one week.

The Sprig MCP connector is the alternative to an export. It exposes study configurations, responses, and themes, retrieves up to 1,000 responses at a time, and can create a draft study.

No study can be launched from the AI client, which the documentation states directly. Draft creation means the instrument can be assembled from a connected client even though launching stays in the Sprig interface.

Three honest limits belong alongside all of that. Sprig's own conjoint documentation describes the estimation as running in a connected AI client, and nothing in the product documentation describes Sprig computing utilities itself. There is no significance testing anywhere in the product, though the analysis guide shows how to run one in the AI client, dividing the gap by the combined standard error and reading 1.96 as the threshold.

And the link-survey restriction means a team that runs everything in-product cannot field this instrument the way they field the rest of their research.

Alternatives and adjacent methods

MaxDiff where the question is which of many items matters most. It scales to far longer lists than a conjoint can carry, and it is the right instrument when you have twenty candidate features rather than five attributes.

Gabor-Granger where the question is price. It reads acceptance across a price ladder directly instead of inferring price sensitivity from trade-offs against other attributes.

Van Westendorp where the question is a price range rather than a point. Note that Sprig offers this as a template and not as a full analysis guide in the way the others are.

A concept test where the question is whether anyone wants the thing at all. Conjoint assumes the category is wanted and measures configuration within it.

And once the study is fielded, Conjoint Analysis with Claude is the next document rather than an alternative to this one.

A key-driver study where you already have the product and want to know which attributes move satisfaction among current users.

Conjoint asks about a hypothetical configuration, and a driver model reads the one you already ship.

A discrete choice experiment in the health-economics tradition where you need a formally derived sample size and a defensible power calculation.

It is the same family of method with a more rigorous published methodology around design efficiency.

Frequently asked questions

What is conjoint analysis?

Conjoint analysis is a survey method that estimates how much each feature of a product contributes to a buying decision, by showing respondents complete configurations and recording which ones they choose.

The output is a set of part-worth utilities, one for every level of every attribute, from which attribute importance and simulated preference share are derived.

What is the difference between choice-based conjoint and traditional conjoint?

Choice-based conjoint shows several complete configurations and asks which one the respondent would choose, which mirrors a purchase. Traditional or ratings-based conjoint asks the respondent to rate or rank profiles one at a time, which extracts more information per profile and asks the respondent to do something they never do when buying. Choice-based is the usual default in commercial work.

How many attributes and levels should a conjoint have?

Three to six attributes with three to five levels each is the working range. Below three attributes a simpler instrument usually answers the question faster, and above six the profiles get long enough that respondents commonly start simplifying and ignoring attributes, which is generally not visible in the output.

How many respondents does a conjoint need?

Derive the sample size from a power calculation rather than from a threshold. The published derivation to work from is de Bekker-Grob et al. (2015) in The Patient, which builds sample size from the significance level, the power, the variance-covariance matrix and the smallest effect size you need to detect.

The widely quoted Sawtooth figure is described by its own source as a rule of thumb.

What is a none option and do I need one?

A none option lets the respondent decline every configuration shown. Include it. Without one the model estimates preferences conditional on buying something and cannot distinguish enthusiasm from a reluctant least-bad pick.

Word it as a real alternative, such as keeping the current solution, and not as an opt-out.

Can conjoint tell me what people will pay?

No. A price attribute inside a conjoint produces a price sensitivity conditioned on the configurations you showed, not a willingness-to-pay curve. Gabor-Granger reads price acceptance directly and is the better instrument for that question.

Is conjoint accurate?

Conjoint ranks configurations reliably and forecasts demand levels poorly. The hypothetical-bias literature, drawn from contingent valuation rather than conjoint, finds a median overstatement ratio of 1.35 with severe positive skew, meaning typical studies overstate modestly and a minority overstate badly. Treat simulated shares as relative comparisons.

What is the difference between conjoint and MaxDiff?

Conjoint measures trade-offs among attributes of a configuration. MaxDiff prioritises a list of separate items. Where you have five attributes that combine into a product, run conjoint.

Where you have twenty features and need to know which matter most, run MaxDiff.

Can I run a conjoint in-product?

Not in Sprig. Conjoint questions are documented as available to Enterprise teams on link surveys only, so the study has to be distributed by link, email or panel rather than triggered inside the product.

Where does the analysis happen?

In a connected AI client, following Conjoint Analysis with Claude. That guide covers estimation, importance, segmentation and significance testing on a worked study. It does not cover share-of-preference simulation or holdout validation, so plan to run those yourself.

Who computes the utilities?

Sprig does not. The analysis runs in a connected AI client, reached either through a CSV export or through the Sprig MCP connector against live data.

The bottom line

Conjoint analysis answers one question well. Among the configurations you could actually ship, which one do buyers prefer, and what are they trading to get it.

It answers that through a design you control completely, which is its strength and its weakness. The attributes and ranges you choose determine the importance scores, so the study is only as honest as the attribute set behind it.

If your goal is to rank configurations and understand the trade-offs inside them, this is the right instrument and the method has been in commercial use since Green and Rao introduced it in 1971.

If your goal is to forecast demand or to find a price, the evidence points elsewhere, to a behavioural test for the first and to a direct price instrument for the second.

Your next study is the analysis itself, run against your own choice data with the attribute set documented alongside the output. Sprig's conjoint analysis guide covers that step, and it is where to go once the design in this guide is fielded.

Back to top
Solutions
Experience measurementStrategic & foundational discoveryJourney & behavioral researchMarket & consumer insightsConcept & prototype testing
Agents
DesignFieldSynthesize
Deploy
EmailPanelsWeb apps and websitesMobile app
Pricing
Community
EventsBlogGuides
CustomersIntegrationsCompare
Company
About usCareersService agreementPrivacy policyData addendumSystem status
Socials
LinkedInX