Introduction
Running a customer journey survey takes five decisions in order, and the first one is definitional because the phrase covers four different instruments.
Decide what you are actually measuring, and a journey survey can validly ask whether an experience felt joined-up. It cannot reliably ask what happened, in what order, when, and how many times, outside a scaffolded recall window of about a day.
Build a touchpoint program rather than a single retrospective questionnaire. That means one common instrument fired at each touchpoint in the moment, sharing a respondent identity key, and analysed as a set.
Keep the recall window short wherever you ask retrospectively, because the evidence on event reconstruction is considerably worse than the evidence on summary judgment.
Then size the program from attrition instead of from a threshold, since every additional wave costs respondents and the loss is neither small nor random.
The four instruments are frequently treated as one, and retrospective reconstruction is commonly taught without any mention of how poorly people perform it.
In one study of a major purchase consumers made themselves, fewer than six in ten recalled the date accurately.
This guide covers:
- Separate the four instruments the phrase "journey survey" covers, because three are real and one is not
- Measure at the touchpoint, in the moment, and join the responses afterwards on a respondent key
- Ask retrospectively only for perceptions of design, never for order, dates, counts or durations
- Budget for attrition wave over wave, since it front-loads and it is not random
- Read the category's founding statistic as consulting data, because that is what it is
What a customer journey survey is
The phrase typically covers four instruments that answer different questions. Three are defensible and one asks respondents to do something the evidence says they cannot.
The four
Touchpoint measurement in the moment. A short common instrument fired at each touchpoint as it happens, with the responses joined afterwards. This is the committed recommendation of this guide.
Retrospective journey-design measurement. A questionnaire asking whether the experience felt coherent and consistent. This is generally legitimate, it has validated scales behind it, and it measures a perception rather than a history.
Longitudinal or experience-sampling designs. Repeated measurement of the same people over time, which is the most rigorous option available and the most expensive in respondent burden.
Journey analytics. Server-side event data describing what people actually did. Not a survey at all, and frequently the right instrument for the question a journey survey gets asked.
And the one that is a category error
Retrospective event reconstruction. Asking a customer to list the touchpoints they encountered, place them in order, date them, and rate how they felt at each one.
That is four separate memory operations layered on each other, and the literature on each of them individually is typically discouraging. The critique section sets out why, and the finding is narrower and more interesting than the usual framing.
What the field's own measurement work says
Gahler, M., Klein, J. F. and Paul, M. (2023), "Customer Experience: Conceptualization, Measurement, and Application in Omnichannel Environments," Journal of Service Research 26(2), DOI 10.1177/10946705221126590, developed the field's most rigorous current scale across seven studies and 3,523 participants, arriving at six dimensions and 18 items.
Then they ruled out the journey-level reading explicitly. In their words, "the CX construct does not cover a customer's overall perception of the customer journey but instead assesses the experience of an individual customer interaction during the customer journey."
Lemon, K. N. and Verhoef, P. C. (2016), "Understanding Customer Experience Throughout the Customer Journey," Journal of Marketing 80(6), 69 to 96, DOI 10.1509/jm.15.0420, the canonical review, reach the same place from the other direction. They state that "There is not yet agreement on robust measurement approaches to evaluate all aspects of customer experience across the customer journey," and describe the multi-touchpoint contribution problem as a neglected area.
So the measure the field trusts is scoped to the interaction, and no validated instrument reconstructs the journey as a sequence of events. Two validated whole-journey scales do exist, and both ask for a perception of design rather than a history, which the critique section sets out. That distinction is the starting point for anything you build.
When to use a customer journey survey
Use touchpoint-level journey measurement when the experience spans several distinct interactions, when you can instrument at least the ones that matter, and when a decision hangs on which of them to fix first.
Three conditions that make it the right instrument
The first is genuine multiplicity. A journey program generally earns its overhead where the customer meets you repeatedly across a period, and a single interaction is typically better served by a single transactional measure.
The second is instrumentation. You have to be able to fire an instrument at the touchpoint, which commonly constrains the design more than most teams expect and is covered in the design chapter.
The third is a prioritisation decision. Journey measurement is generally at its best when the question is which touchpoint to fix, and at its weakest when the question is how good the journey was overall.
The question it answers well
Which touchpoint is underperforming relative to the others in the same program, for which segment, and by how much.
That is a comparative question and it is answerable. Whether the journey as a whole is good is typically a harder question with a weaker measurement basis behind it.
What kind of evidence this produces
A journey program produces stated evaluations of individual interactions, measured close to when they happened, from the people who reached those interactions.
It does not produce behavioural evidence, it does not produce causal evidence, and it does not reach the people who never arrived. Each of those has a better instrument, named in the next section.
When not to use a customer journey survey
Six things a journey survey cannot tell you, and each names the instrument that answers the question instead.
It cannot tell you what people actually did
Self-reported behaviour and logged behaviour disagree, and they disagree in a way you cannot correct for.
Parry, D. A. and colleagues (2021), "A systematic review and meta-analysis of discrepancies between logged and self-reported digital media use," Nature Human Behaviour 5(11), 1535 to 1547, DOI 10.1038/s41562-021-01117-5, meta-analysed 66 effect sizes from 44 studies totalling 52,007 participants.
The correlation between self-report and logs was r = 0.38, with a 95 percent confidence interval of 0.33 to 0.42.
The accuracy analysis is the part that matters most. Across 49 comparisons, only three, or 6.12 percent, had a mean self-reported figure within 5 percent of the logged mean. Of the rest, 46.94 percent over-reported and 46.94 percent under-reported.
Read that symmetry carefully. Self-report is not biased in a correctable direction, it is unreliable in both, so there is no adjustment factor to apply.
The instrument that answers this is server-side event analytics on your own logs.
It cannot tell you why a touchpoint failed at the level of interface detail
A score typically tells you only that checkout underperforms, and it does not tell you which control confused people.
Faulkner, L. (2003), Behavior Research Methods, Instruments, and Computers 35(3), 379 to 383, DOI 10.3758/BF03195514, tested 60 users and resampled random subsets. Five users found between 55 and 99 percent of known problems depending entirely on which five, ten users found at least 80 percent, and twenty users found at least 95 percent.
That quantifies the trade you are making, and a journey survey with 400 responses tells you checkout scored 3.2. Twenty usability sessions tell you at least 95 percent of the reasons.
The instrument that answers this is usability testing.
It cannot tell you whether a touchpoint caused churn
A journey survey is observational and carries no counterfactual. Unhappy customers rate everything lower and they also leave, which means the correlation between a low touchpoint score and churn is generally overdetermined.
The instrument that answers this is a controlled experiment, or a discrete-time hazard model fitted to your retention data.
It cannot tell you when anything happened
Morwitz, V. G. (1997), "It Seems Like Only Yesterday: The Nature and Consequences of Telescoping Errors in Marketing Research," Journal of Consumer Psychology 6(1), 1 to 29, DOI 10.1207/s15327663jcp0601_01, checked recalled purchase dates against records for a major durable good people had bought themselves.
Accuracy was 58.8 percent among 97 households in June 1988 and 55.8 percent among 215 households in June 1989. Among those who recalled inaccurately, forward telescoping dominated, at 67.5 and 64.2 percent, and accuracy declined significantly as time since purchase increased.
Roughly four in ten consumers misdate a major purchase they personally made. A journey survey asking a respondent to place four touchpoints in sequence is asking for four such judgements and compounding them.
The instrument that answers this is a timestamped event log.
It cannot tell you how long or how effortful the journey was
This is the best-evidenced limit in the list.
Redelmeier, D. A. and Kahneman, D. (1996), Pain 66(1), 3 to 8, DOI 10.1016/0304-3959(96)02994-6, tracked real-time and retrospective pain across 154 colonoscopy patients and 133 lithotripsy patients.
Retrospective evaluation correlated with peak pain at r = 0.64 and 0.63, and with end pain at r = 0.43 and 0.46.
It correlated with duration at r = 0.03 and 0.11, across a colonoscopy duration range of 4 to 67 minutes.
Fredrickson, B. L. and Kahneman, D. (1993), Journal of Personality and Social Psychology 65(1), 45 to 55, DOI 10.1037/0022-3514.65.1.45, state the mechanism directly: "Retrospective evaluations appear to be determined by a weighted average of 'snapshots' of the actual affective experience, as if duration did not matter."
So a journey survey will return a number when you ask how long or how hard something was, and the number will be roughly uncorrelated with the truth.
One distinction prevents a confusion the critique section would otherwise create. Duration is neglected when people estimate how long something took, and duration nonetheless weights how they evaluate a multi-episode sequence. Those are compatible: the evaluation tracks how long the good and bad parts lasted without the respondent being able to report those durations accurately.
The instrument that answers this is instrumented time-on-task and funnel timing.
It may be measuring the wrong construct entirely
This one is not a measurement limit, it is a design limit, and it applies before any of the others.
Siebert and colleagues argue that customer experience research converges too quickly on the smooth journey model. For a recreational rather than an instrumental product, unpredictability is often part of the appeal, and a program scoring every touchpoint on effort reduction is optimising against the thing customers came for.
Ask what your product is for before you build the instrument. The instrument that answers this is qualitative work on what customers value in the experience, and the critique section covers the evidence.
It cannot tell you what the people who abandoned experienced
Lemon and Verhoef's touchpoint taxonomy has four categories: brand-owned, partner-owned, customer-owned and social or external. Two of those four are structurally out of reach of a survey you fire inside your own product.
And an in-product survey fires at a touchpoint that an abandoner never reached, so the people whose experience most needs explaining are the people the instrument cannot see.
This is an argument from the canonical framework rather than an empirical result, and it should be read as one. The instrument that answers it is funnel analytics paired with an exit intercept.
Study design, and the fork that decides everything else
The fork is which of the four instruments you are building. Get that wrong and every later decision is downstream of a category error.
Choose by the question, not by the phrase
Read the question you were actually asked, then pick the branch. Most teams typically arrive having been asked something imprecise, and the value of this table is that it forces the imprecision into the open before anyone writes a questionnaire.
| What you were asked | What you should build | Why |
|:----------------------------------------------------:|:-----------------------------------------------:|:--------------------------------------------------------------------:|
| Which step of onboarding should we fix first? | Touchpoint measurement in the moment | Comparative, actionable, and measurable close to the event |
| Does our experience feel joined-up to customers? | Retrospective journey-design measurement | A perception of design, with validated scales behind it |
| How does sentiment change over the first 90 days? | Longitudinal or experience-sampling design | Requires the same people measured repeatedly, and costs accordingly |
| Where do people drop out, and how long does it take? | Journey analytics on event logs | Behavioural and temporal, which surveys measure badly |
| What did customers do, in what order, and when? | Journey analytics on event logs | Not a survey question. The memory evidence is against you |
| Why did they abandon at step three? | Funnel analytics plus an exit intercept | The in-product survey never reaches them |
| Which touchpoint drives renewal? | Touchpoint measurement joined to retention data | The survey supplies the ratings, your warehouse supplies the outcome |
The committed default
Build a touchpoint program. One common instrument, fired in the moment at each touchpoint you can instrument, sharing a respondent identity key, analysed as a set.
That design is what the field's own measurement work supports, since the validated construct is scoped to the interaction. It is also the design that survives the memory evidence, because it never asks anyone to reconstruct anything.
The comparison across touchpoints is where the value sits. A single touchpoint score is often nearly uninterpretable on its own, and the same score read against five sibling touchpoints measured the same way is a priority list.
It works at commercial scale, and the evidence is specific
The in-the-moment design is not only the defensible one, it has been run at size. Baxendale, S., Macdonald, E. K. and Wilson, H. N. (2015), "The Impact of Different Touchpoints on Brand Consideration," Journal of Retailing 91(2), 235 to 253, DOI 10.1016/j.jretai.2014.12.008, used contemporaneous SMS capture to record tens of thousands of touchpoint encounters across four product categories.
Two findings carry directly into a program design. Touchpoint positivity added explanatory power over touchpoint frequency alone, which means counting encounters is not enough and you have to measure how each one felt.
And the touchpoint that mattered most was one nobody instruments. Peer observation, which the authors describe as an almost entirely neglected touchpoint, was consistently significant.
That second finding is a warning about coverage. A program measuring only what it can instrument may be missing the touchpoint that moves the outcome, and the honest response is to say so rather than to report the measured set as the whole.
What has to be identical across touchpoints
The comparison only works if the instrument is genuinely common. Changing the wording between touchpoints generally means the differences you find are partly differences in the question.
Hold the core item wording identical, changing only the named object. Hold the scale identical in length, labels and direction, and hold the response options identical and in the same order.
Where a touchpoint genuinely needs a different item, add it as an extra question rather than by altering the common one.
The identity key, and where the join happens
The program depends on being able to tell that response A at touchpoint one and response B at touchpoint four came from the same person.
Sprig documents that identifier on the response record itself. The CSV export carries a visitorId and a nullable userId on every row, alongside a responseGroupUid that reassembles the answers from one sitting.
The read API returns visitorId, visitorUuid and externalUserId on each response object.
The join happens outside the platform. Sprig documents no cross-study respondent view, no linked-response timeline and no numeric comparison across studies, so the program is assembled in a warehouse or a notebook from exported responses rather than in the interface.
Two practical consequences. Set your own user ID through the SDK or API wherever you can, because the userId field is nullable and an anonymous respondent falls back to the visitor identifier.
And plan the join as a real work item with an owner, since it is the step that turns a set of separate studies into a program.
The touchpoint map comes before the instrument
If you have not drawn one yet, mapping the journey is the step before this one, and distinguishing a user flow from a user journey is the step before that.
List the touchpoints first, and mark each one as instrumentable or not. Lemon and Verhoef's four categories are the right frame: brand-owned, partner-owned, customer-owned, and social or external.
An in-product survey reaches brand-owned digital touchpoints. That is a fact about the instrument class rather than about any platform, and it means a journey program built this way is measuring a subset of the journey by construction.
Say so in the reporting. A program that measures four brand-owned touchpoints and reports a journey score is overclaiming, and the correction is a sentence rather than a redesign.
Writing the instrument
A touchpoint instrument is short by design. Two or three questions is typically the working ceiling, because you are going to fire it several times at the same person.
The core item
Ask about the interaction that just happened, naming it explicitly. "How was checking out?" is a better item than "How was your experience?" because the second one invites the respondent to answer about something else.
Keep it evaluative rather than reconstructive. The question is how this felt, not what happened, and the distinction is the whole basis of the design.
Example wording
A touchpoint instrument asks two questions. Written out, at a checkout touchpoint, they look like this.
How was checking out? (Very difficult / Difficult / Neither / Easy / Very easy)
What would have made that easier?
At an onboarding touchpoint, the same instrument reads:
How was setting up your account? (Very difficult / Difficult / Neither / Easy / Very easy)
What would have made that easier?
Only the named object typically changes. That is the whole discipline, and it is what makes the two scores comparable.
The follow-up
One open-ended probe, worded to ask for a specific change instead of a justification. "What would have made that easier?" generally returns codeable causes, and "Why did you give that score?" typically returns defensiveness.
Where the platform supports an AI follow-up on a low score, consider it. Probing in the moment can recover a reason that a rating alone does not carry, though no published accuracy or reproducibility claim accompanies the feature and the output should be read as a lead rather than as a finding.
The touchpoint program specification
Write this once and apply it to every touchpoint. The specification is the artifact that makes the set comparable, and a program without one commonly drifts within a quarter.
| Element | Specification |
|:-----------:|:--------------------------------------------------------------------------------------------------:|
| Core item | One evaluative question, identical wording except for the named touchpoint |
| Scale | Identical length, labels and direction at every touchpoint, frozen for the life of the program |
| Probe | One open-ended item, identical wording, asking for a change rather than a justification |
| Trigger | The behavioural event that marks the touchpoint, defined once per touchpoint and documented |
| Timing | Fired on completion of the interaction, not at session end and not the next day |
| Identity | The respondent key present on every response, with your own user ID set wherever possible |
| Suppression | A recontact window long enough that a multi-touchpoint program does not exhaust the same people |
| Attributes | The segment fields you will cut by, set before delivery, since they are snapshotted at that moment |
| Extra items | Permitted per touchpoint, never replacing or altering the common core |
| Change log | Every wording, scale or trigger change, dated, because it breaks comparability from that date |
The two fields people forget
Attributes are captured as of delivery. Sprig documents that a response carries attribute values only where the respondent had those values set when the study was delivered, which means a segment you define next quarter does not retroactively appear on last quarter's responses.
So decide your reporting cuts before you field, not after. This is the same discipline the sampling chapter asks for and it is forgotten more often.
The second is the change log, and a journey program runs for quarters and somebody will improve the wording. That improvement breaks the comparison silently unless the date is recorded.
What not to put in the instrument
Do not ask which other touchpoints the respondent has encountered. Do not ask them to place this interaction in a sequence. Do not ask how long the overall journey has taken so far.
Each of those is event reconstruction, each has evidence against it in the critique section, and each will return a confident number that does not correspond to anything.
Scale and scoring choices
Pick one scale, freeze it, and use it at every touchpoint. Which scale you pick matters considerably less than whether it is the same one everywhere.
Which scale to use
The guide declines to name a universal winner because the evidence does not support one, and a reader still has to choose something today. Here is a defensible default.
Use a five-point labelled scale with every point labelled, running from a clearly negative anchor to a clearly positive one, and keep the same five labels at every touchpoint.
Five points is generally enough to detect the differences a journey program acts on, every point labelled removes the interpretation the respondent would otherwise supply, and an odd number gives a genuine midpoint for an interaction someone found unremarkable.
If your organisation already runs a standard satisfaction or effort item, use that instead. Consistency with your existing reporting is usually worth more than any property of a better scale, because the comparison people will actually make is against numbers they already have.
What matters far more than the choice is that you freeze it. A scale changed mid-program breaks every comparison before and after the change.
What to report
Report the distribution alongside the summary. A touchpoint scoring 3.4 with a bimodal distribution is a different problem from one scoring 3.4 with everyone clustered in the middle, and the mean hides the difference.
Report the count with every score. Touchpoint programs produce wildly unequal response volumes, because some touchpoints are typically hit by everyone and some by a tenth of the base, and a comparison that ignores this will rank noise.
Comparing across touchpoints, carefully
The comparison is the point of the program and it is also where the errors live.
Different touchpoints generally attract different populations. The people who reach cancellation are not the people who reach onboarding step two, so a difference in score is partly a difference in who answered.
Say that in the reporting rather than adjusting it away. The honest framing is that touchpoint A scored lower than touchpoint B among the people who reached each, which is a weaker claim and a true one.
On significance
Sprig documents no significance testing anywhere in the product, so any formal comparison between touchpoint scores happens in your own analysis after export.
That is a real gap rather than an inconvenience. A program that ranks six touchpoints by mean score without any interval around them will produce a priority order that reshuffles next month for no reason, and nothing in the product will stop it.
Sample size and precision
No source publishes a defensible sample-size rule for journey studies. Not the canonical reviews, not the scale-development papers, and not any of the pages currently ranking for this query.
Do not let a plan smuggle in a figure like a hundred responses per touchpoint. Nobody derived it.
Why the scale papers do not give you a threshold
Kuehnl, Gahler and Jaakkola all report the samples they used, at 4,814, 3,523, and 278 with 239 for the nomological test. Those are typically scale-development samples sized for factor analysis and validation.
Presenting a scale-development sample as a fielding threshold would be a fabrication. They answer a different question about a different kind of study.
What does exist is a longitudinal procedure
Elkasabi, M., Suzer-Gurtekin, Z. T. and Chen, Y. (2023), "A framework for sample size calculations in longitudinal surveys to measure net and gross changes," PLOS ONE 18(9), e0291449, DOI 10.1371/journal.pone.0291449, gives a sequential algorithm rather than a closed form, implemented in the nchange R package.
Its instruction for the part that matters here is to inflate sample sizes to account for panel attrition and nonresponse, using wave-specific completion rates.
The catch is in the inputs. The procedure needs the between-wave correlation and the per-wave completion rates, and you will generally not have either before your first study.
So run wave one as the pilot
This is the honest instruction and it is genuinely useful. Field the first wave unpowered, treat it as the pilot that estimates your own between-wave correlation and completion rate, then power the later waves from your own numbers rather than from anyone's rule of thumb.
That is typically slower than copying a threshold and it produces a number you can defend.
What attrition actually does
Two facts change how a multi-wave program should be planned, and both come from a large general-population panel.
Cabrera-Álvarez, P., James, N. and Lynn, P. (2023), Understanding Society Working Paper 2023-03, Institute for Social and Economic Research, University of Essex, report General Population Sample retention against wave one of 64.1 percent at wave four, 51.0 percent at wave seven, 42.7 percent at wave ten, and 39.9 percent at wave eleven.
The shape matters more than the levels. In their words, "Most dropouts occurred in the first four waves (35.9%)," after which the attrition rate decreased significantly.
Attrition front-loads. A three-wave journey program sits entirely inside the steepest part of that curve, so it commonly bleeds at close to the worst rate throughout in place of settling down.
And attrition is not random. The paper documents higher dropout among younger respondents, ethnic minorities, people in poorer health, people on lower incomes, students and the unemployed.
Read the consequence carefully, because it is the one that produces false findings. A journey program showing satisfaction rising across waves may be watching its own sample gentrify.
This is a UK household panel and the magnitudes do not transfer to a product onboarding flow. The direction and the shape do.
What compliance looks like when people are trying
Vachon, H., Viechtbauer, W., Rintala, A. and Myin-Germeys, I. (2019), Journal of Medical Internet Research 21(12), e14475, DOI 10.2196/14475, meta-analysed 79 experience-sampling studies covering 8,013 participants.
Pooled compliance was 79.7 percent, with a 95 percent confidence interval of 77.5 to 81.8, and pooled retention was 94.0 percent, with an interval of 92.0 to 95.7. Both figures follow the exclusion of influential outliers.
The mandatory caveat belongs in the same breath as the number. These are research studies with consented, often incentivised participants, frequently in clinical populations.
Treat 79.7 percent as a ceiling rather than as a forecast for a commercial program where nobody signed up for anything.
Three moderators are directly actionable, and fixed sampling schedules beat semi-random ones by 6.7 percentage points. Compliance falls by roughly one percentage point for each additional daily prompt, which works out to about eight points between two and ten prompts a day.
And longer gaps between prompts help, with a 10.8-point difference between hourly and four-hourly intervals.
The burden budget
Work this out before you field, not after. The point of the worksheet is that a six-touchpoint program typically looks free on a whiteboard and costs most of its sample in practice.
| Input | Where it comes from | Effect on completed responses |
|:------------------------------:|:-----------------------------------------------------------:|:------------------------------------------------------:|
| Touchpoints instrumented | Your touchpoint map | More touchpoints means more prompts per person |
| Prompts per person per period | Trigger frequency against your recontact window | Each additional prompt reduces compliance |
| Recontact window | Your suppression setting | A longer window protects compliance and thins coverage |
| Expected compliance per prompt | Your own pilot, with the research ceiling as an upper bound | Multiplies through every wave |
| Waves | Program duration | Attrition compounds, and front-loads in the first few |
| Reporting segments | Your analysis plan | Each cut divides the surviving sample again |
Multiply and not estimate. Start from eligible respondents, apply expected compliance at each prompt, apply attrition wave over wave, then divide by the number of segments you intend to report.
The number at the bottom is what you will actually analyse, and it is usually a surprise.
The recontact window is the control you have
A recontact waiting period is the documented mechanism for stopping a multi-touchpoint program from exhausting the same people. It is also the setting that determines how much of the journey any one respondent sees.
One documented detail matters for testing. Sprig states that the recontact window is only enforced in production environments, so a development environment will show the same visitor study after study and will give you a misleading picture of the program's real burden.
Audience, targeting and screening
The population for a journey program is defined by the touchpoint rather than by the customer. Each instrument has its own eligible population, and the differences between them are a finding instead of a nuisance.
Who is eligible at each touchpoint
Everyone who reached the touchpoint, and that is the whole rule, and it has a consequence people miss.
The populations at different touchpoints are generally nested, not parallel. Everyone who reaches step four also reached step one, and the people who dropped out between them are absent from the later measurement by construction.
So a later touchpoint scoring higher than an earlier one may be telling you that the dissatisfied left rather than that the experience improved. Check the drop-off between the two before reporting the difference as progress.
Sampling within a touchpoint
Sample a fraction in place of everyone, and set the fraction from the burden budget instead of from what the trigger allows.
A program that surveys every event at every touchpoint typically trains customers to dismiss the widget, and the dismissal rate is not random either.
Screening, and why journey programs need less of it
The behavioural trigger is the screen. Someone who just completed checkout has been screened by having completed checkout, which is generally a more reliable qualification than anything they would self-report.
Add an explicit screener only where the trigger cannot distinguish a case that matters, such as separating first-time from repeat buyers when the event looks identical.
Segments, decided in advance
Set the attributes you will cut by before the study is delivered, because attribute values are captured at delivery. A segment defined later does not appear on earlier responses.
Keep the segment list short. Every cut divides a sample that the burden budget has already thinned, and a program reporting six touchpoints across five segments is reporting thirty cells that mostly contain too few responses to read.
Fielding and delivery
Fire the instrument at the touchpoint, in the moment. That is the single most consequential delivery decision in the design and it is the one that distinguishes this program from a retrospective questionnaire.
Behavioural triggering is the mechanism
The instrument should be triggered by the event that constitutes the touchpoint instead of by a schedule or a page view. Event triggering is what makes in-the-moment measurement possible, and it is the capability the whole design rests on.
Define the triggering event once per touchpoint and write the definition down. A trigger that drifts, or that two people define differently, produces a series that is not comparable with itself.
What in the moment actually means
Immediately after the interaction completes, not at the end of the session and not the following day.
The recall-window section explains why the tolerance here is tighter than intuition suggests. The short version is that the memory evidence degrades quickly, and a survey fired a day later is effectively a retrospective instrument wearing an in-product costume.
Firing a survey is itself an intervention
This belongs in the design and not in a footnote, because it is a real cost of the in-the-moment approach and the retrospective alternatives do not carry it.
Chandon, P., Morwitz, V. G. and Reinartz, W. J. (2005), Journal of Marketing 69(2), 1 to 14, DOI 10.1509/jmkg.69.2.1.60755, report that "On average, the correlation between latent intentions and purchase behavior is 58% greater among surveyed consumers than it is among similar nonsurveyed consumers," and state the model applies to intentions, attitude or satisfaction data.
Measuring changes the measured. A journey program fires at the same person repeatedly, which is a repeated intervention, though the study does not address repeated measurement and you should not claim it compounds across waves on this evidence.
The practical response is a holdout. Keep a fraction of eligible users unsurveyed and compare their behavioural outcomes against the surveyed group, which turns an unmeasurable worry into a number.
Channel constrains coverage
An in-product instrument reaches brand-owned digital touchpoints and nothing else. Partner-owned, customer-owned and social touchpoints need a different instrument or go unmeasured.
Decide which it is and record it. A program that quietly measures four of eleven touchpoints and reports on the journey is making a claim its coverage does not support.
Timing and cadence
Two clocks run in a journey program. One is how soon after the touchpoint you ask, and the other is how often you are willing to ask the same person.
The recall-window rule
This table is the one to check a draft questionnaire against. It comes from combining what the reconstruction evidence rules out with what the Day Reconstruction Method shows is recoverable when recall is scaffolded.
| You may ask retrospectively | You may not ask retrospectively |
|:-------------------------------------------------------------------------:|:------------------------------------------------------------:|
| Whether the experience felt coherent or consistent | Which touchpoints they encountered |
| Whether it felt personalised or thematically joined-up | What order those touchpoints came in |
| How they feel about the relationship now | When any of them happened |
| Yesterday's episodes, if you scaffold the reconstruction the way DRM does | How many times anything occurred |
| Whether they would repeat or recommend | How long the journey took, or how much effort it accumulated |
The left column is a perception of design, held now, and asking for it is legitimate. The right column is event reconstruction, and the critique section sets out the evidence against each row.
On the scaffolded exception
Kahneman, D., Krueger, A. B., Schkade, D. A., Schwarz, N. and Stone, A. A. (2004), "A Survey Method for Characterizing Daily Life Experience: The Day Reconstruction Method," Science 306(5702), 1776 to 1780, DOI 10.1126/science.1103572, validated a structured reconstruction procedure against experience sampling on 909 employed women, documenting close correspondence between the two.
The constraint is the point. DRM works over yesterday, with a structured episode-by-episode protocol, and it is not a licence to ask someone about a six-week onboarding.
Cadence
The recontact window sets how often one person can be asked. Set it from the burden budget rather than from the default, and set it once for the whole program so that every touchpoint competes fairly for the same respondent.
A short window gives you coverage of more touchpoints per person and costs compliance. A long window protects compliance and means most respondents contribute to only one or two touchpoints, which weakens the within-person comparison.
Quality control and data hygiene
Journey programs typically fail quietly. Each individual study looks fine and the set does not add up, usually for one of four reasons.
The join is where the program breaks
Check it first, because everything downstream depends on it. Export two studies, join them on the respondent key, and count how many respondents appear in both.
A low overlap is not necessarily a bug. It may mean your recontact window is doing its job, or that few people reach both touchpoints. But it is typically the number that determines whether you have a program or a collection of separate studies, and most teams commonly discover it at analysis time.
Where your own user ID is set, join on that. Where it is not, the visitor identifier is the fallback and it is weaker, since a person on two devices is two visitors.
Drift in the common instrument
Compare the live wording of the core item across every touchpoint, monthly. Someone will usually improve one of them.
The change log is the control. Without it a wording change is often invisible in the data and shows up as a step change in one touchpoint's score that everyone tries to explain behaviourally.
Response volume asymmetry
Check the counts per touchpoint before reading any comparison. Touchpoints hit by everyone and touchpoints hit by a tenth of the base produce wildly different precision, and a ranked list treats them as equals.
Trigger correctness
Verify that each instrument is firing where it claims to. Fire the trigger yourself, in production instead of in development, and confirm the response lands with the right study, the right touchpoint label and the right attributes.
Development environments do not enforce the recontact window, which makes them useless for testing the program's real behaviour and fine for testing one instrument in isolation.
Analysis: the core calculation
The analysis is a comparison across touchpoints, within segments, with intervals. It is arithmetically simple and frequently easy to do misleadingly.
The calculation
For each touchpoint, compute the summary score, the full response distribution, and the count. Then compute an interval around each score, because the whole exercise is a comparison and comparisons without intervals typically reshuffle monthly.
Rank by score, and read the ranking against the intervals rather than against the point estimates.
One statistical caution, because the intuitive version of this check is wrong. Two confidence intervals that do not overlap do imply a difference, and two that overlap do not imply no difference. To compare two touchpoints properly, compute an interval around the difference between them instead of eyeballing whether their separate intervals touch.
The within-person view, where the join supports it
Where enough respondents appear at more than one touchpoint, the within-person comparison is stronger than the between-touchpoint one, since it removes the differences between populations that the sampling chapter warned about.
Compute it as the difference between a respondent's own scores at two touchpoints, averaged across respondents who answered both. That number is harder to produce and much harder to argue with.
Joining to outcomes
The program's highest-value analysis is joining touchpoint scores to what happened next. Retention, renewal, expansion, or whatever outcome your business actually tracks.
Do that in your warehouse on the respondent key, and read it as association rather than cause.
One design change follows from the counter-critique, and it is the concession made concrete. If the outcome you care about is renewal, repurchase or referral, add a periodic remembered-evaluation item to the program alongside the in-the-moment touchpoint measures.
The evidence in the critique section is that remembered experience predicts the desire to repeat better than experience measured as it happens. A program built only on in-the-moment measurement is using the weaker predictor for precisely that question, and the fix commonly costs one recurring item. Unhappy customers rate everything lower and also leave, which is the confound named earlier and it does not go away because the join is technically clean.
Before you join to outcomes, check what you are allowed to join
The analysis above links individual survey responses to individual commercial outcomes, which turns anonymous feedback into identified records about named customers.
Confirm the basis for that before you build it. Whether your privacy notice covers linking survey responses to account data, whether respondents were told, and what your retention period is for the joined table are questions that typically have owners outside the research team.
Keep the joined dataset minimal and time-bounded. A program that generally needs the link only for a quarterly analysis does not need to retain identified response-level records indefinitely, and the difference is a retention rule somebody writes down.
What the platform computes and what it does not
Sprig documents no significance testing anywhere in the product, and no cross-study numeric comparison. The per-study views summarise one study at a time.
There is an AI-generated insights feed that aggregates across studies and produces narrative output, gated on response thresholds. Read it as a prompt for investigation rather than as a computed statistic, and note that a narrative describing a correlation is not a correlation coefficient.
Benchmarks and what a good result looks like
There is no sourced cross-industry benchmark for journey-level satisfaction or for touchpoint scores, and the structure of the measure means there could not be one.
A touchpoint score depends on which touchpoint, measured how, on what scale, to whoever reached it. None of those travel between companies.
The vendor benchmark tables
Several vendors publish industry benchmark tables for customer experience. The pattern across them is consistent: no disclosed sample, no field dates, no sampling frame, and no statement of whether the unit is a transaction, a relationship or a journey.
None of them is journey-level. The closest to documented is a large vendor's pre-made industry benchmark set, and even there the public documentation does not disclose sample sizes, computational method or field dates, and no benchmark is labelled for journey-level scoring.
The one real satisfaction index, the American Customer Satisfaction Index, is measured at company level. It is neither journey-level nor touchpoint-level and using it as either is a category error.
The founding statistic, sourced and not debunked
The claim that journeys matter more than touchpoints has a single origin and it is worth knowing exactly what it is.
Rawson, A., Duncan, E. and Jones, C. (2013), "The Truth About Customer Experience," Harvard Business Review 91(9), reprint R1309G, report that "performance on journeys is 30% to 40% more strongly correlated with customer satisfaction than performance on touchpoints is," and separately that it is "20% to 30% more strongly correlated with business outcomes."
They also report that "the gap between the top- and bottom-quartile companies on journey performance was 50% wider than the gap between top- and bottom-quartile companies on touchpoint performance."
The evidence base is McKinsey's own cross-industry customer experience surveys, with an exhibit covering seven companies in each of two industries. There is no methodology section, no sampling design, no statistical method and no confidence intervals.
That is not a reason to dismiss it. It is a reason to cite it as what it is, which is practitioner research from a consulting firm, and to stop presenting it as an established empirical result.
The entire category rests on this article, and naming its basis precisely is more useful than either repeating it or attacking it.
So what does good look like
Compare against your own prior periods on a frozen instrument, and against your own other touchpoints measured the same way.
A good result is a clear ranking with intervals that separate, a worst touchpoint that somebody owns, and a change in that touchpoint's score after somebody fixed it.
That last one is the only real validation a journey program gets.
Interpreting and acting on the result
The output is a ranked list of touchpoints with intervals. Turning it into action takes three reads, in order.
Read the ranking against the intervals
Start with which touchpoints are reliably worse than the others, not with which is lowest. A point estimate at the bottom of a list with heavily overlapping intervals is a candidate instead of a finding.
Where nothing separates, the honest report is that no touchpoint stands out at current volume, and the action is to run longer instead of to pick one.
Read the distribution behind each score
A touchpoint with a bimodal distribution is a different problem from one with everyone bunched in the middle, and the mean is identical.
Bimodality at a touchpoint usually means two populations are hitting the same interaction with different needs. That is a segmentation finding, and it is generally more actionable than the score.
Read the open text against the score, not instead of it
Code the probe responses into themes and attach the score to each theme. Read the theme table by mean score rather than by count, because a small theme with a very low mean is frequently a broken experience affecting one segment and it will not surface in a count-ranked list.
Check the unclear bucket first. A large unclear bucket means the theme list does not fit the data, and every downstream number inherits the problem.
What to hand the decision-maker
One touchpoint, named, with its score, its interval, its count, the distribution behind it, and the three themes from the open text with their own mean scores.
Then the scope limit in the same document. The program measured the touchpoints it could instrument, among the people who reached them, and the ranking holds within that.
And what to do next
Fix the touchpoint, then measure it again on the frozen instrument. Keeping that loop running across an organisation is its own discipline, described as journey management. A journey program's only real validation is that a score moved after somebody changed something, and programs that never close that loop commonly become reporting furniture within a year.
Running the analysis with Claude or ChatGPT
The comparison this program needs is not available inside the platform, so the analysis happens in a connected AI client against exported data. This section covers the whole workflow.
What the analysis produces
Three artifacts, and a touchpoint comparison table with scores, counts and intervals. A within-person comparison for the respondents who answered at more than one touchpoint.
And a theme table from the open-text probe with the score attached to each theme.
The first is the priority list, and the second is the stronger version of the first. The third is why.
Setup and getting your data in
Export each touchpoint study as CSV, with responses and the user attributes you intend to cut by, and keep the label export as well as the values so the codebook travels with the data.
Every exported row carries visitorId and, where you have set it, userId, plus responseGroupUid for the answers from a single sitting. Those are the fields the join runs on.
The alternative to exporting is the MCP connector, which reaches studies and responses directly from a supported client. One caution for this particular workflow: the connector's output schema is not documented, so for a join that depends on specific identifier fields the CSV export or the read API is the safer route. The 1,000-response page size is documented for the read API rather than for the connector.
Prompt one: the touchpoint comparison
What this prompt does: computes the touchpoint comparison table from a set of
exported study files.
What it returns: one table of touchpoints ranked by score, each with n, the
summary score, the full distribution and a confidence interval.
You have [NUMBER, default 6] CSV files, one per touchpoint, exported from a
customer journey program. The score column is [COLUMN NAME, default "score"] on a
[SCALE, default 1 to 5] scale, where top-box means [TOP BOX, default 4 or 5].
Use code to calculate this, not estimation.
Rules:
1. Exclude blank responses and non-responses from all calculations. Report how
many you excluded per touchpoint.
2. Any touchpoint with fewer than [THRESHOLD, default 30] valid responses is
reported as "insufficient responses" instead of as a number. A touchpoint
with zero valid responses is also reported as "insufficient responses", not
omitted and not counted separately.
3. Compute the mean, the full response distribution, and the top-box share.
Put a 95 percent interval around the mean using a t-interval, and a 95
percent interval around the top-box share using the adjusted Wald method.
Name both methods in the output and say which quantity each interval is
around.
4. Recompute every interval a second time by a genuinely different route, for
example a bootstrap over the response values, and compare the two.
5. If the two routes disagree by more than [TOLERANCE, default 0.5] scale
points on any interval bound, report the disagreement explicitly. Do not
reconcile it silently and do not pick one.
6. Output one markdown table named "Touchpoint comparison", sorted by score
ascending, plus a short list of every exclusion and every disagreement.
Prompt two: the within-person comparison
What this prompt does: finds respondents who answered at more than one
touchpoint and compares their own scores against each other.
What it returns: a paired comparison table plus the overlap counts.
Using the same files, join on [ID COLUMN, default userId, falling back to
visitorId where userId is blank].
Use code to calculate this, not estimation.
Rules:
1. Report the overlap first: how many distinct respondents appear at each pair
of touchpoints. Do this before any comparison.
2. Exclude blanks and non-responses. Where a respondent has more than one
response at the same touchpoint, keep the earliest and say how many
duplicates you dropped.
3. Any touchpoint pair with fewer than [THRESHOLD, default 30] respondents in
common is reported as "insufficient overlap" instead of as a number. A pair
with zero respondents in common is also reported as "insufficient overlap".
4. For each pair above the threshold, compute the mean within-person
difference and a 95 percent interval around that difference using a paired
t-interval. Name the method in the output.
5. Recompute the mean difference a second time by a genuinely different route,
for example a bootstrap over respondent-level differences, and compare.
6. If the two routes disagree by more than [TOLERANCE, default 0.2] scale
points on the mean difference or either bound, report the disagreement
explicitly. Do not reconcile it silently and do not pick one.
7. Output one markdown table named "Within-person touchpoint comparison" and
state plainly which pairs had insufficient overlap.
Reading the output
Read the overlap counts before the comparison table. A within-person analysis resting on eleven respondents is arithmetic instead of evidence, and the threshold rule is there to stop it being reported.
Read the two tables against each other, and where the between-touchpoint ranking and the within-person ranking agree, the finding is solid.
Where they disagree, the difference is population rather than experience, and the within-person version is the one to trust.
What to verify before reporting
Four checks, and the third catches the most embarrassing error available.
Confirm the exclusion counts are plausible against the raw response totals. Confirm every reported touchpoint clears the stated threshold. Confirm the scale direction is what you think it is, because if the verbatims attached to your highest scores describe frustration then the scale has been inverted somewhere.
And confirm the interval method is named in the output rather than assumed.
Pitfalls
Models asked for a confidence interval without a named method will usually produce the plain Wald interval, which performs poorly at the proportions journey programs produce. Name the method.
Models asked to double-check their work will typically re-run the same route and agree with themselves. Requiring a genuinely different computation is what surfaces an error, which is why both prompts ask for one.
And models are generally agreeable. A prompt that implies which touchpoint you expect to be worst will frequently get you that answer, so ask for the ranking without indicating what you expect to see.
The critique you should know
The standard objection to journey surveys is that memory distorts summary judgments, usually invoked as peak-end and duration neglect. That framing is half right, and a guide that states it plainly is wrong in a way a knowledgeable reader will catch.
What peak-end actually establishes
For a single continuous episode, the evidence is strong. Kahneman, D., Fredrickson, B. L., Schreiber, C. A. and Redelmeier, D. A. (1993), Psychological Science 4(6), 401 to 405, DOI 10.1111/j.1467-9280.1993.tb00589.x, found that 22 of 32 participants, or 69 percent, chose to repeat the longer trial containing more total discomfort, z = 2.15, p < .05.
Redelmeier and Kahneman (1996), cited earlier for duration, found the same shape in real medical procedures. Peak and end predicted retrospective evaluation and duration did not.
A support call, a checkout flow, an onboarding screen: those are single continuous episodes and this evidence applies to them.
What it does not establish
A journey is not a single episode. It is a sequence of discrete episodes separated in time, and for that structure the finding reverses.
Miron-Shatz, T. (2009), "Evaluating multiepisode events: Boundary conditions for the peak-end rule," Emotion 9(2), 206 to 213, DOI 10.1037/a0015295, studied days composed of many discrete episodes, across samples of 810 participants in the United States, 820 in France and 805 in Denmark.
In her words, "The duration-weighted average of these feelings represented the normative approach to evaluation, and, contrary to the predictions of the peak-end rule, the average was the best predictor of retrospective evaluations of the day."
The duration-weighted average predicted overall evaluation at beta = .84 in the United States and Denmark and .82 in France. And "The end episode did not add to the participants' overall evaluations of the previous day beyond its contribution to the duration-weighted net affect."
The peak-end effect is real, and it does not beat the average
Alaybek, B. and colleagues (2022), "Meta-analytic evidence for the peak-end rule," Organizational Behavior and Human Decision Processes 170, 104149, DOI 10.1016/j.obhdp.2022.104149, meta-analysed 174 effect sizes and found the peak-end effect at r = 0.581, with a 95 percent confidence interval of 0.487 to 0.661.
Their headline conclusion supports the peak-end rule. The effect is substantial and it is well established, and any argument that leans on this paper has to start there.
One comparative finding inside the same analysis is what matters here. The peak-end effect was comparable to the effect of the overall average score, while being stronger than the effects of the first episode, the lowest-intensity episode, and the trend and variability across episodes.
So peak and end carry real information, and so does the simple average. Designing a journey program around the assumption that only the peak and the end are remembered is not supported, and neither is the claim that peak-end makes retrospective judgment worthless.
One citation note. A corrigendum, DOI 10.1016/j.obhdp.2023.104278, published in January 2024, re-ran the analyses under a uniform sign-reversal rule and reports that virtually all substantive conclusions were unchanged, though two external commentators disagree with the recoding. Cite the pair instead of the original alone.
So the real problem is different, and it is worse
The retrospective journey survey does not fail because people distort their summary. On the evidence, their summary is generally roughly normative.
It fails at event reconstruction. Ordering, dating, attributing and enumerating touchpoints are the operations the memory literature says people cannot perform reliably, and they are exactly what a retrospective journey questionnaire asks for.
Morwitz's finding sits here rather than in the peak-end literature. Roughly four in ten consumers misdate a major durable purchase they made themselves, accuracy commonly decays with elapsed time, and a journey questionnaire compounds several such judgements in a single grid.
That distinction is the part most treatments of this topic leave out, and it is the one that changes what you should write on a questionnaire.
The counter-critique, and it is strong
Four arguments run the other way and the guide is weaker without them.
Remembered experience predicts choice better than experience does. Wirtz, D., Kruger, J., Scollon, C. N. and Diener, E. (2003), Psychological Science 14(5), 520 to 524, DOI 10.1111/1467-9280.03455, found that remembered experience, but neither online nor anticipated experience, directly predicted the desire to repeat the experience.
Their conclusion, verbatim: "although on-line measures may be superior to retrospective measures for approximating objective experience, retrospective measures may be superior for predicting choice."
If your business question is whether they will renew, repurchase or refer, the remembered evaluation is the causally relevant variable and the in-the-moment measure is often answering a different question.
For multi-episode events, retrospective global evaluation is approximately normative. That is Miron-Shatz again, and it cuts both ways. It undermines the peak-end objection and it supports asking for a global judgment.
Two validated whole-journey scales exist, and they predict loyalty. Kuehnl, C., Jozic, D. and Homburg, C. (2019), Journal of the Academy of Marketing Science 47(3), 551 to 568, DOI 10.1007/s11747-018-00625-7, define effective customer journey design as "the extent to which consumers perceive multiple brand-owned touchpoints as designed in a thematically cohesive, consistent, and context-sensitive way," validated across 4,814 consumers in two countries, and find it raises loyalty over and above brand experience.
Jaakkola, E. and Terho, H. (2021), Journal of Service Management 32(6), 1 to 27, DOI 10.1108/JOSM-06-2020-0233, define service journey quality as the degree to which customers perceive touchpoints "functioning as a (1) seamless, (2) coherent and (3) personalized whole," validated on 278 respondents with a further 239 for the nomological test, and report service journey quality predicting loyalty intentions at beta = 0.76.
Read what those scales actually ask, and they ask for a perception of design, whether it felt coherent, consistent, personalised. They do not ask for a reconstruction of events.
That is precisely the line this guide draws, and these papers sit on the permitted side of it.
And seamlessness is not universally the goal. Siebert, A., Gopaldas, A., Lindridge, A. and Simões, C. (2020), Journal of Marketing 84(4), DOI 10.1177/0022242920920262, argue that "CXM research is too quickly converging on the smooth journey model, without recognizing legitimate alternatives."
Against the loyalty loop they set the involvement spiral, "a cyclical pattern of unpredictable experiences that motivates increasing experiential involvement over time," drawn from ethnographic work with 40 informants across 43 journeys.
For a recreational rather than an instrumental product, a journey program scoring every touchpoint on effort reduction is measuring the wrong thing.
Their study is retrospective and interview-based instead of longitudinal, and should be described that way.
Common mistakes
Seven that recur, in rough order of how much damage they do. Each one has a tell, which is the thing to look for when you suspect it rather than the thing that proves it.
These are the failures that survive a careful design, because every one of them happens after the instrument is built and while the program is running.
Asking respondents to reconstruct the journey. The single biggest one, and the most commonly taught. Ordering, dating and enumerating touchpoints is the operation the evidence says people cannot perform, and the survey will return a confident grid of answers regardless.
The tell is that the data looks unusually clean. Reconstructed sequences are tidier than real ones because respondents produce a plausible story rather than a recalled one.
Firing the instrument a day later and calling it in-the-moment. The tolerance is typically tighter than intuition suggests. A survey that arrives the next morning is a retrospective instrument, and it should be designed as one or moved earlier.
The tell is in the trigger configuration instead of in the data. Check what actually fires the instrument, because a daily batch send is easy to inherit and invisible afterwards.
Letting the common instrument drift. Somebody improves the wording at one touchpoint and the comparison quietly stops meaning anything. The change log is cheap and nobody keeps one.
The tell is a step change in one touchpoint that nobody can attribute to a release. Before building a theory, diff the live question text against what the specification says it should be.
Reading a later touchpoint's higher score as improvement. The populations are nested. The people who were unhappy at step one may simply not be present at step four, and the score rose because they left.
The tell is the drop-off rate between the two touchpoints. Where it is large, survivorship is the first explanation to rule out and it usually cannot be ruled out.
Ranking touchpoints without intervals. Journey programs produce very unequal response counts across touchpoints, so a ranking by point estimate reshuffles month to month for no reason and everyone then tries to explain the movement.
The tell is that the bottom of the ranking changes more often than anything is being shipped. A ranking that moves faster than the product is measuring noise.
Reporting a journey score from four brand-owned touchpoints. An in-product instrument reaches one of the four touchpoint categories. A program that measures a subset and reports on the whole is overclaiming, and the fix is a sentence about coverage.
The tell is the absence of any coverage statement at all. Ask what fraction of the mapped touchpoints the program instruments, and if nobody knows, that is the answer.
Skipping the holdout. Firing surveys changes behaviour, and a repeated program is a repeated intervention. A fraction of eligible users left unsurveyed turns that from a worry into a number.
The tell is that nobody can say what the program costs. A holdout is the only way to answer whether the measurement is paying for the behaviour it changes.
Synthetic respondents
A synthetic respondent can produce plausible touchpoint ratings. Whether those ratings carry the structure of a real population is unsettled, and for this method there is a specific reason for caution.
The whole value of a journey program is the comparison between touchpoints. That comparison depends on real differences in real experiences at different points in a product, which is exactly the kind of grounded specificity a model simulating a persona does not have.
A synthetic panel asked to rate six touchpoints will typically return six plausible numbers, and the differences between them will be generated rather than observed.
Because the program's output is the ranking, a generated ranking is not a weaker version of the finding, it is a different object entirely.
Use synthetic responses to pressure-test the instrument before fielding, which is a genuine use. Do not use them to produce the touchpoint ranking a roadmap decision rests on.
How Sprig supports this method
Four documented capabilities do real work in a journey program, and each has a constraint worth planning around.
Behavioural event triggering is the on-method capability. Firing the instrument at the event that constitutes the touchpoint is what makes in-the-moment measurement possible, and it is the mechanism the whole design depends on.
Constraint: it reaches brand-owned digital touchpoints only, which is a property of the instrument class rather than of the platform.
The recontact waiting period is the documented fatigue control for a multi-touchpoint program, which is the direct answer to the burden evidence.
Constraint: it is suppression only, it operates on the visitor record, and it is enforced only in production environments.
The respondent identifier on every response is what makes the program a program. The export carries visitorId and a nullable userId, and the read API returns visitorId, visitorUuid and externalUserId.
Constraint: the join happens outside the platform, since no cross-study respondent view is documented.
AI follow-ups probe a low score in the moment, which recovers the reason a rating alone cannot carry. Constraint: no published accuracy or reproducibility claim accompanies them.
Five absences belong beside those. There is no cross-study trend chart and no numeric comparison across studies. There is no significance testing anywhere in the product.
Export is CSV only, capped at 950,000 rows with a one-week link expiry, and there is no response import. And the cross-study insight surface is AI narrative instead of computed statistics.
One more worth naming plainly, because a reader will meet it. Sprig's marketing describes tracking experience continuously and segmenting by lifecycle stage. Neither a cross-study trend view nor an in-product respondent timeline is documented, so read those as descriptions of what the program lets you build rather than as features you will find in the interface.
Alternatives and adjacent methods
Journey analytics on event logs where the question is what people did, in what order, and how long it took. Surveys measure all three badly and logs measure all three exactly.
Usability testing where the question is why a specific touchpoint fails. Twenty sessions find at least 95 percent of the known problems, which is a better return than 400 survey responses for that question.
A validated retrospective journey-design scale where the question is whether the experience felt coherent. Kuehnl's effective customer journey design and Jaakkola and Terho's service journey quality are the two with published validation, and both ask for a perception rather than a reconstruction.
Experience sampling or a diary study where the question is how sentiment moves over time within a person. It is the most rigorous option and the compliance evidence sets expectations honestly.
The Day Reconstruction Method where you genuinely need retrospective detail and the window is about a day. It is validated, it is structured, and it does not scale to a six-week onboarding.
A transactional satisfaction or effort measure where there is really only one interaction that matters. A journey program is usually overhead if the journey is one step.
Frequently asked questions
What is a customer journey survey?
A customer journey survey measures how customers experience the interactions that make up their relationship with a product or service. The phrase covers four different instruments: touchpoint measurement fired in the moment, retrospective measurement of whether the journey felt well designed, longitudinal or experience-sampling designs, and journey analytics on behavioural logs.
The defensible default is touchpoint measurement in the moment, with responses joined afterwards on a respondent identifier.
Can I just ask customers about their whole journey afterwards?
You can ask whether the journey felt coherent and personalised, and two validated scales do exactly that. You cannot reliably ask which touchpoints they encountered, in what order, when, how many times, or how long it all took.
Those are event-reconstruction questions, and the evidence is that people generally answer them confidently and inaccurately.
How many responses does a journey survey need?
No source publishes a defensible sample-size rule for journey studies, and the figures circulating are not derived from anything. The published procedure to work from is Elkasabi and colleagues (2023) in PLOS ONE, which inflates sample size using wave-specific completion rates.
It needs your between-wave correlation and completion rate as inputs, so run the first wave as an unpowered pilot that estimates them.
What is the difference between a touchpoint survey and a journey survey?
A touchpoint survey measures one interaction. A journey survey, done defensibly, is a set of touchpoint surveys using a common instrument, fired at several interactions, and joined afterwards on a respondent identifier so the results can be compared. The two are not alternatives. A journey program is built out of touchpoint surveys, and what makes it a journey program is the shared instrument and the join instead of a different questionnaire.
How is a journey survey different from a transactional survey?
A transactional survey measures one interaction. A journey program measures several interactions with a common instrument and compares them against each other, which is where its value sits.
A single touchpoint score is nearly uninterpretable alone, and the same score read against five siblings measured identically is a priority list.
Does peak-end bias ruin journey surveys?
Peak-end bias does not ruin journey surveys in the way it is usually claimed. Peak-end is established for single continuous episodes, and for multi-episode sequences a duration-weighted average predicts retrospective evaluation better, at beta above .8 in three national samples.
A meta-analysis of 174 effect sizes also found the peak-end effect comparable to the effect of the simple average. The real failure of a retrospective journey survey is event reconstruction rather than summary distortion.
Can I link one person's answers across several studies?
Yes, outside the platform. Sprig's response export carries a visitorId and a nullable userId on every row, and the read API returns visitorId, visitorUuid and externalUserId. Join the exports on your own user ID where you set it, falling back to the visitor identifier where you do not. Sprig documents no in-product view that does this for you.
How often can I survey the same person?
That is set by the recontact waiting period, and it is the main lever you have over fatigue in a multi-touchpoint program. Experience-sampling research finds compliance falling by roughly one percentage point per additional daily prompt, with about eight points separating two prompts a day from ten.
Those figures come from consented research participants, so treat them as a ceiling rather than a forecast.
What is a good benchmark for a journey score?
There is no sourced cross-industry benchmark for journey-level or touchpoint scores, and the structure of the measure means one is not possible. A score depends on which touchpoint, measured on what scale, among whoever reached it.
Compare against your own prior periods on a frozen instrument and against your own other touchpoints.
Where did "journeys matter more than touchpoints" come from?
From Rawson, Duncan and Jones (2013) in Harvard Business Review, reporting McKinsey cross-industry surveys covering seven companies per industry. The article contains no methodology section, no sampling design and no confidence intervals.
It is practitioner research, it founded the category, and it is worth citing accurately rather than either repeating or attacking.
The bottom line
A journey survey answers one question well. Among the touchpoints you can instrument, which one is underperforming, for whom, and by how much.
It answers that by measuring each interaction close to when it happened and comparing the results, which is what the field's own validated measure supports and what the memory evidence permits.
If your goal is to find the weakest step and fix it, this is the right instrument, though the field itself has not settled on an agreed measurement approach for the journey as a whole and this guide does not pretend otherwise.
If your goal is to reconstruct what customers did and when, the evidence points at your event logs instead, and no amount of careful question wording will recover it from memory.
Your next study is the cheapest test of everything above. Instrument one touchpoint in the moment, and compare what it says against what your retrospective survey said about that same touchpoint.
The gap between the two is the guide's argument, measured on your own data.