Solutions
Experience measurement
Track sentiment and KPIs with AI-driven gap analysis
Strategic & foundational discovery
Uncover market whitespace with AI-led foundational studies
Journey & behavioral research
Connect user actions to motivations across the lifecycle
Market & consumer Insights
Understanding markets, audiences, & opportunity
Concept & prototype testing
Test designs and prototypes with rapid feedback
Agents
Design
Structure rigorous studies
Field
Run adaptive studies at scale
Analyze
Get statistical analysis you can trust
Synthesize
Turn results into research reports
Deploy
Email
Reach external audiences with native deliverability
Panels
Recruit from 300K+ verified participants
Web apps and websites
Embed studies in web experiences
Mobile apps
Run studies in iOS and Android apps
Customers
Community
Events
Join curated gatherings shaping the future of research
Blog
Insights on integrating AI into research craft
Book icon
Guides
Ultimate playbooks for enterprise survey research
Pricing
Sign in
Book a demo
Sign in
Book a demo
Guide

When to Run: AI-Moderated Interviews

October 2, 2026

By James Villacci

Example H2
Example H3
Example H4
Example H5
Example H6

Introduction

Run an AI-moderated interview when the answer you need is a reason rather than a number. AI interviewers are well suited to questions about why people behave as they do, how they describe a problem, and what they were trying to accomplish, particularly when the population is small or unknown and no figure will leave the room.

Run a survey when the answer has to hold for people who were not in the study. A share, a ranking, a trend over time, or a comparison between segments needs a fixed instrument, a stated base, and a margin of error.

Most research programs need both, and they often run in that order: interviews to find the reasons, a survey to size them. A human moderator remains the stronger choice for deep discovery in a specialist domain, so the routing decision is typically three-way rather than two-way.

The randomized evidence published in 2026 supports the depth claim and complicates the coverage claim. In one experiment with 3,160 panelists, AI interviews produced 4.8 times the words per assigned respondent, while completion fell from 99.4 percent to 40.5 percent.

What an AI-Moderated Interview Is, and What a Survey Is

The two instruments differ in who asks the questions and whether every respondent hears the same ones, and that difference typically decides which one a question needs.

What an AI-moderated interview is

An AI-moderated interview is a one-to-one conversation in which a language model asks the questions, listens to the answer, and decides what to ask next.

The researcher writes a discussion guide, and the AI interviewer follows it, probes vague answers, and asks for examples, typically over voice, video, or text.

The output is a set of transcripts, typically with summaries and themes attached. Every participant has a different conversation, and that is the point of the method.

What a survey is

A survey is a fixed instrument: every respondent sees the same questions, in the same wording, with the same answer options, in an order the researcher controls.

Open-ended and closed-ended questions can sit in the same instrument, but the stimulus stays identical.

Rather than producing an account from each participant, a survey produces a distribution across all of them. That distribution is what supports a percentage, a ranking, or a tracked metric.

A third instrument matters here too: the human-moderated interview, in which a trained researcher runs the conversation. This guide treats it separately, because the evidence shows it is not interchangeable with an AI interviewer.

AI-Moderated Interviews vs. Surveys: The Core Difference

AI-moderated interviews answer reasoning questions from relatively few participants. Surveys answer incidence questions from a known population, with a defensible estimate attached. Rather than the channel or the interface, the difference is the kind of evidence each instrument can produce, and that decides which one a research question needs.

| | Survey | AI-moderated interview | |--------------------------|:------------------------------------------------:|:------------------------------------------------------------------------:| | Question it answers | How many, how often, which wins, by how much | Why, under what conditions, what the person was trying to do | | Evidence produced | An estimate for a population, with a stated base | An account from a named set of participants | | Who asks | A fixed instrument, identical for everyone | An adaptive interviewer, different for everyone | | Participant count | Sized in advance to a margin of error | Sized by saturation, with no published benchmark | | What the output supports | Tracking, segment cuts, a business case | Explaining a number, generating hypotheses, finding the unasked question |

The table compresses the argument, and every row has the same root. Surveys trade depth for comparability, and AI interviews trade comparability for depth.

Why identical stimulus matters

Comparability across respondents comes from identical stimulus. If one person is asked a follow-up about pricing and another is asked about onboarding, their answers cannot be counted into the same percentage.

An adaptive interviewer deliberately gives up identical stimulus in exchange for depth.

That is a correct trade for a qualitative instrument and a disqualifying one for population estimation. It is a property of the method, not a gap in any vendor's roadmap.

A question type is not a methodology

Several AI interview platforms now ship structured question types, including ranking, multiple choice, and MaxDiff, and at least one publishes its estimation method. Those features are real and worth conceding.

But a question type is the visible half of a method. The defensible half is the sampling, completion, and inference apparatus around it: a stated base, a sample sized in advance, and a check on who finished.

Rather than asking whether a tool can show a ranking question, ask whether it can defend the ranking.

The decision rule in one sentence

If the answer needs to hold for people who were not in the room, it is a survey question. If the answer is about what happens in the room, it is an interview question.

Rather than asking which tool is better, ask which kind of evidence the decision requires.

When to Run an AI-Moderated Interview

Run an AI-moderated interview when depth is the deliverable and a population estimate is not. Six situations meet that bar.

Before an instrument exists

Exploratory work before a survey is written is where interviews most often win outright. A survey written too early measures the researcher's assumptions rather than the customer's.

For example, a team planning a persona study can use 30 AI interviews to learn the vocabulary customers actually use, then write answer options from that vocabulary.

When the population is small

A small, known population changes the math. With 30 enterprise accounts or a 12-person internal team, a percentage moves several points with each answer, so the deeper instrument is often the better trade.

Rather than forcing a survey onto a group that cannot produce a stable percentage, interview them.

When the reasoning is the deliverable

Some decisions turn on how people think, not how many think it. Concept reaction is the common case. A product team testing a new pricing page often needs to hear how a buyer explains the tiers back, more than it needs a purchase-intent score.

A concept test that reports only a top-two-box number frequently misses the confusion that explains the number.

When a free-text box is going unread

This is the one real displacement of survey budget, and it should be conceded plainly. A free-text question at the end of a survey that nobody probes and nobody reads produces thin answers.

In the Verasight experiment described below, written open-ends averaged 27 words per assigned respondent, against 128 for AI interviews.

Where an open-end typically returns a sentence, an AI interview can often return a paragraph with an example in it. But the same experiment shows the trade: the interview reached far fewer of the people assigned to it.

When you need to test survey wording before launch

AI interviews can serve as a form of cognitive pretesting. Asking a small group from the target audience to explain a draft question back in their own words often surfaces ambiguous terms before launch.

Instead of discovering after fieldwork that "workspace" meant three different things, a team can fix the wording before the survey reaches a thousand people. The interview here serves the survey rather than replacing it.

When reach across languages and time zones matters

Nielsen Norman Group lists multilingual studies among the stronger use cases for AI moderators. A team that needs reasoning from customers in Brazil, Germany, and Japan in the same week often cannot staff human moderators for all three, and an AI interviewer typically can run the same guide in each language.

When Not to Run an AI-Moderated Interview

An AI-moderated interview is the wrong instrument whenever the reader of the result will ask for a base size. Six situations call for something else, and each one names the better instrument.

When you need a population share

"What percentage of admins want single sign-on" is a survey question. An AI interview can tell you why admins want it, but more interviews are still not an estimate.

Use a survey with a stated base and a sample sized to the precision the decision needs.

Rather than inference to a population, scale in the AI interview category typically means more interviews.

A stakeholder who hears that pricing came up in most interviews will often hear a percentage, whatever the readout actually said.

When you need to compare segments or track over time

Adaptive stimulus is not comparable across respondents, so it is not comparable across segments or waves either. A quarterly satisfaction metric, a brand tracking wave, or a comparison between enterprise and mid-market buyers needs a fixed instrument.

A difference between segments in an interview study often reflects who was asked which follow-up, rather than how the segments actually differ. Use a survey, hold the wording constant, and track it.

When the topic predicts who will finish

Some topics correlate with willingness to complete an AI interview. In the Verasight experiment, people who finished the AI interview were 7.4 points more AI-optimistic than everyone assigned to that condition.

A study about attitudes to AI, technology comfort, or privacy is where that self-selection distorts the finding.

Use a survey, or at minimum compare the completers against the full invited sample before reporting anything.

When the study needs deep, specialist discovery

Nielsen Norman Group advises against AI moderators for exploratory discovery and for domains where "an AI interviewer may miss important details." A study of how clinical coordinators schedule trial visits, or how procurement teams run a security review, often needs a moderator who recognizes the thread worth pulling.

Use a human-moderated interview. In foundational work, the cost of a missed thread is typically higher than the cost of running fewer conversations.

When you need observed behavior

Rather than observing behavior, interviews report it. What people say they did on a checkout page and what the session shows are often different.

Recall is typically weakest for routine actions, which are frequently the ones a product team most wants to understand.

Use session replay, product analytics, or a usability observation.

When you need to know what caused a change

Neither an interview nor a survey establishes cause on its own. If the question is whether a new onboarding flow raised activation, participants can describe what they experienced, but they cannot report the counterfactual.

Use a controlled experiment, such as an A/B test with a pre-registered success metric, and use interviews afterward to explain the result. Instead of asking customers whether a change helped, measure the difference between people who saw it and people who did not.

What the six limitations have in common

Every limitation above is the same limitation seen from a different side. An AI interview is optimized for depth from each participant, and each of these situations needs something depth cannot supply: a base, a comparison, a representative completer set, a trained ear, an observation, or a counterfactual.

AI-Moderated vs. Human-Moderated Interviews

Human moderators still win on adaptive judgment, and AI moderators win on running one discussion guide identically across many conversations. Human-moderated and AI-moderated interviews are different instruments, and treating them as one hides the most useful routing decision in qualitative research: deep discovery to a person, breadth of reasoning to an AI.

Where a human moderator still wins

The independent literature places AI moderation between a self-administered survey and a skilled human moderator.

One preprint study of 2,739 participants found that "the AI agent did not reach the human benchmark in probing for richer information, especially during open-ended questions." A smaller study found human interviewers elicited 38 percent longer responses, and its authors stated the result was "by no means indicating that it can replace expert human interviewers."

Nielsen Norman Group reaches the same conclusion from practice: AI interviewers "still struggle with the skills that make human-led discovery interviews effective: adapting in real time, making nuanced tradeoffs, and picking on subtle cues."

The practical consequence is about where the stakes sit. A foundational study that will shape a product strategy for two years generally justifies a trained moderator, even at a fraction of the interview count.

Where an AI moderator is the better trade

AI moderators run the same discussion guide the same way at 2 a.m. and across languages. That consistency is valuable for structured qualitative work: feature feedback, reaction to a known concept, or a study that needs 80 conversations rather than 8.

Nielsen Norman Group names this directly, noting that short AI interviews "could be a strong alternative to surveys for gathering structured product or feature feedback." The routing that follows is three-way rather than two-way: deep discovery to a human, breadth of reasoning to an AI moderator, and incidence to a survey.

Human judgment remains essential at every step, including deciding which of the three a question needs. Researchers remain responsible for checking transcripts against the guide before trusting that consistency.

The Three-Question Test for Choosing an Instrument

The Three-Question Test routes a research question to an AI-moderated interview only when three conditions hold at once. It is built from the situations where interviews substitute for surveys, and it adds the selection check that the 2026 randomized evidence made necessary.

Ask these three questions of every research question before choosing an instrument:

  • Is the population small or unknown?
  • Will no number leave the room?
  • Is the topic unrelated to who will finish an AI interview?

Yes to all three routes the question to an AI-moderated interview. The first question treats a population as unknown whenever nobody needs its size, so a large customer base asked only for reasons passes it.

No to any of the three routes the question to a survey, or to a survey after the interview. A fourth question then separates the two kinds of interview by asking whether the study needs a moderator who can recognize an unexpected thread in a specialist domain. Yes to the fourth routes the interview to a human moderator rather than an AI one.

| Answers | Instrument | What the readout can say | |:---------------------------------------:|:------------------------------:|:------------------------------------------:| | Yes, yes, yes | AI-moderated interview | Reasons, language, conditions, examples | | Yes, yes, yes, and specialist discovery | Human-moderated interview | The same, with deeper probing | | No to question 1 or 2 | Survey | A share, a ranking, a trend, a segment gap | | No to question 3 only | Interview, then a survey check | Reasons, but no claim about who holds them |

The next section applies the same rules to 20 real research questions.

How to read a mixed result

Mixed answers are common, and they frequently mean the question is two questions. "Why are trial users churning, and how many churn for pricing" splits cleanly. The first half is an interview question and the second is a survey question.

Rather than forcing one instrument onto a compound question, split it and run each half on the instrument it needs. A research lead who splits questions this way spends interview budget only where interviews are the stronger instrument.

Why the third question exists

The first two questions come from long-standing research practice, and the third is new in 2026.

The Verasight experiment showed that the people who finish an AI interview can differ on the very topic being studied, which means the caveat applies to qualitative findings as well as to estimates.

A theme heard in 40 AI interviews about AI adoption may describe AI enthusiasts rather than customers. That is worth knowing before the theme reaches a roadmap.

The Question Router: 20 Research Questions and the Instrument Each Needs

The question router applies the Three-Question Test to 20 common research questions. Each row names the instrument and the reason in one line, so a team can compare its own backlog against a worked set.

| Research question | Instrument | Reason | |:----------------------------------------------------------------:|:---------------------------------------:|:-----------------------------------------------------------:| | Why did trial users stop at step 3 of setup? | AI-moderated interview | Reasoning, and the population of drop-offs is often unknown | | What share of admins want single sign-on? | Survey | A population share | | How do CFOs describe this category in their own words? | AI-moderated interview | Language is the deliverable | | Which of eight features matters most? | Survey | A ranking across respondents | | How did satisfaction change since last quarter? | Survey | A tracked metric needs identical wording | | What were buyers trying to do when they found us? | AI-moderated interview | Jobs and context | | Do enterprise users rate onboarding lower than mid-market users? | Survey | A segment comparison | | How do our 25 largest accounts run vendor security reviews? | Human-moderated interview | Small population, specialist domain | | What confuses people about the new pricing page? | AI-moderated interview | Concept reaction where reasoning matters | | What price makes this plan feel expensive? | Survey | Price sensitivity needs a distribution | | How do people feel about AI features in our product? | Survey, then interviews | The topic predicts who finishes an AI interview | | Why do detractors score us low? | AI-moderated interview after the survey | The survey finds them, the interview explains them | | Which features do mobile users open before they churn? | Product analytics, then interviews | Observe first, ask second | | Where do users hesitate in checkout? | Session replay, then interviews | Observe first, ask second | | Which headline makes the value clearest? | Survey | A message test compares variants | | What would make a churned customer return? | AI-moderated interview | Reasoning from a hard-to-reach group | | How aware is the market of our brand? | Survey | Awareness is a population share | | How do clinicians schedule trial visits today? | Human-moderated interview | Specialist workflow discovery | | What words do users use for this feature? | AI-moderated interview | Vocabulary before instrument design | | How big is the problem the interviews surfaced? | Survey | Sizing a finding |

The pattern in the router

Nine of the 20 questions route to a survey, seven to an AI-moderated interview, two to a human moderator, and two to observation before asking. The survey rows all end in a number that someone will repeat in a meeting: a share, a ranking, a change, a gap.

The interview rows end in a reason, a word, or an example.

Questions that change instrument halfway

Four rows use two instruments in sequence, and they are the most instructive. The detractor question starts with a survey that finds low scorers and ends with interviews that explain them.

The AI-sentiment question runs the survey first, precisely because the topic would bias who finishes an interview.

Sequencing is not a compromise, and it is typically the standard design.

What the Randomized Evidence on AI Interviews Shows

Two randomized experiments published in August 2026 compared AI interviews against written survey questions. Both found large depth gains.

They split on coverage, and the difference in their designs is the most useful thing in either paper.

The embedded design: Austin and colleagues

Austin, Hohe, Kennedy, Litman, Minozzi and Moses, writing in Public Opinion Quarterly on August 17, 2026, randomized 2,243 respondents between fixed survey prompts and AI interviewing embedded in the survey flow. AI interviewing added 57 to 70 words and 0.63 to 0.75 additional reasons per issue.

Satisfaction fell by a negligible 0.06 points on a seven-point scale. Only 10 of 2,253 participants failed to complete, with no meaningful differential attrition across arms.

One co-author, Leib Litman, co-founded CloudResearch, which is worth knowing when reading the paper.

The same study found a cost the headline numbers hide. After justifying their answers to the AI, respondents gave more polarized answers to later items.

The handoff design: Verasight

Morris, Bakhshi, Leff, Rothschild and Marshall, publishing through Verasight on August 31, 2026, randomized 3,160 panelists between written open-ends and a separate AI-moderated interview.

Outset provided the interview platform, and fieldwork ran May 22 to 31, 2026.

Depth won clearly. The interview condition produced 128 words per assigned respondent against 27, and 1.8 times the information on a quality score that ignores length.

The authors attribute roughly 80 percent of the gap to the additional probing.

Coverage lost clearly. Completion was 99.4 percent for written open-ends and 40.5 percent for the interview, and 16 percent of AI completes were fraud-flagged.

One clean AI interview required assigning 2.9 respondents, against 1.01 for a written response.

The dropout was not random. Completers were 7.4 points more AI-optimistic than the full assigned group, and roughly three-quarters of that bias survived a full demographic adjustment.

The authors conclude that "no qualitative richness compensates for losing six in ten people non-randomly when the estimand is a population share."

What the two studies agree on

Both studies show that AI probing produces more words and more reasons than a static written prompt. Rather than resting on vendor case studies, that finding now rests on two separate randomized designs.

Among Verasight completers, 86 percent rated the experience 4 or 5 out of 5, and 45 percent preferred the AI interview to typing.

Teams can plan on a depth gain and spend their scrutiny on who the depth comes from.

What the two studies do not settle

The studies differ on panel, topic, and design, so neither isolates the handoff as the cause of dropout. The Verasight authors are explicit that "most of the loss occurs at the hand-off between methods," that respondents "were not told in advance," and that "browsers can fail when redirecting someone across websites."

Their experience measures cover completers only, and interviewing cost was inferred from assignment ratios rather than measured. One panel, one interview platform, and four topics were tested.

Those limits travel with every number above, and a single study is not a verdict on any platform. The depth results favor the platform tested.

How to use these numbers in your own planning

Use the two studies to set expectations, not targets. The depth gains are likely to generalize, because two different designs found them.

The completion results are less likely to generalize, because they diverged sharply with design.

The practical step is to measure your own completion from the invitation in a small pilot, and to compare completers with the invited list, before committing a program to either design.

Are AI-Moderated Interviews Representative?

AI-moderated interviews are generally not designed to be representative, and the 2026 evidence shows the people who complete them can differ from the people invited. Representativeness is a property of sampling and completion, and an interview study usually controls neither well enough to support a population claim.

The Verasight result makes the risk concrete. When six in ten assigned respondents do not finish, and the finishers lean in a direction that correlates with the topic, reweighting on demographics does not fix it.

Roughly three-quarters of the attitude gap remained after adjustment.

Two practical checks follow. First, compare completers against the full invited sample on every attribute you already hold, such as plan, tenure, region, and usage.

Second, write the readout so that it describes the participants rather than the market.

Rather than claiming interview findings represent customers, claim that they represent a mechanism worth sizing. Screening respondents carefully and recruiting from research panels improves who enters an interview study, but it does not create an estimate.

What a representative interview study would require

An interview study could, in principle, support a population claim. It would need a probability sample or a well-specified quota design, identical core questions for every participant, completion high enough that nonresponse cannot plausibly move the result, and a reported check of completers against the invited sample.

Few interview studies meet all four conditions. When one does, it has effectively become a survey with an interview attached, which is a legitimate design in its own right.

How Many AI-Moderated Interviews Do You Need?

No published benchmark says how many AI-moderated interviews a study needs. Saturation, not sample size, governs qualitative work, and a larger count of interviews is still not an estimate.

Saturation is the point at which new conversations stop producing new themes. It generally depends on how varied the population is, how narrow the question is, and how many segments need their own read.

A study of one onboarding step in one segment typically saturates sooner than a study of purchase decisions across four buyer roles.

Two rules of thumb are defensible without an invented threshold. Plan interviews per segment rather than in total, because a theme that appears only in one segment needs enough conversations in that segment to be recognized.

And track new themes per batch, stopping when a batch adds nothing the previous batches did not.

Count until the themes stop moving, not until a round number is reached.

Account for dropout when planning, because the Verasight experiment needed 2.9 assigned respondents per clean interview, so a team that wants 40 usable conversations from a panel may need to invite far more than 40.

Treat that ratio as one study's result rather than a planning constant, and measure your own in a pilot.

If the stakeholder question becomes "how many people feel this way," the study has changed instruments. That calls for a survey sized to a margin of error, which is a different calculation entirely.

Should You Use AI Follow-Up Questions Inside a Survey?

Use AI follow-up questions inside a survey when an open-ended answer needs one more layer of detail, and place them after every item you plan to report. A randomized experiment by Austin and colleagues in Public Opinion Quarterly found that embedded AI probing added depth with no meaningful dropout, and it also found that justifying an answer to the AI shifted later answers toward the extremes.

Until it is replicated, that second finding is a reasonable risk to assume for any product that adds AI follow-ups to a survey, including research-grade interview tools. Adaptive follow-up inside a survey is also not an interview.

It is a probe attached to a fixed instrument, which keeps the survey's comparability for the closed items and adds depth for the open ones.

Follow these placement rules when adding AI follow-ups to any survey:

  • Place AI follow-ups after the closed items you will report, never before them
  • Cap follow-ups per respondent so that fewer later answers follow a justification
  • Switch follow-ups off for tracking waves where wording must stay identical across periods
  • Report follow-up themes as explanation of the closed results, not as separate estimates
  • Compare completion with and without follow-ups before adopting them program-wide

The rules exist because the same mechanism that produces depth can contaminate measurement. Placement is what keeps the two apart.

When AI follow-ups are the wrong addition

AI follow-ups are generally the wrong addition to a short transactional survey, such as a two-question satisfaction pulse after a support ticket. The respondent has seconds to give, and a probe typically adds length without adding a decision.

They are also the wrong addition when the open-ended answer will never be read. Instead of adding probes to a question nobody analyzes, remove the question or assign an owner to it.

Sprig note. Sprig's Field Agent generates follow-up questions in real time based on responses, inside the survey. Depending on configuration, AI follow-ups can be turned off for a study. Limitation: Sprig has not published a randomized test of its follow-ups against static open-ends, so the Austin completion result should not be assumed to transfer. Field Agent follow-ups are adaptive probes on a survey, not AI-moderated interviews, and Sprig does not offer voice moderation or live observation.

How to Combine AI Interviews and Surveys: The Handoff Workflow

The handoff workflow runs interviews first to find reasons and a survey second to size them, and it names where each instrument's evidence stops. It follows directly from the Verasight result, which shows interview depth arriving without a population base. In this guide, the handoff workflow means the planned move from interview findings to a sizing survey, not the redirect between platforms that Verasight measured.

The sequence has five steps:

  1. Run AI-moderated interviews on the reasoning question until themes stop moving
  2. Turn each theme into a testable hypothesis with a named population
  3. Write a fixed survey that measures each hypothesis in the customers' own words
  4. Field the survey to a sample sized for the precision the decision needs
  5. Report the number from the survey and the explanation from the interviews

Each step has a boundary. Interview evidence stops at step 2. It can say a reason exists and describe it, but not how common it is.

Survey evidence starts at step 4. It can say how common the reason is, but its explanation comes from the interviews.

A worked example of the handoff

A subscription software team hears in 35 AI interviews that some trial users stall because they cannot invite teammates before paying.

That is a finding about a mechanism, but it is not yet a finding about the business.

The team writes a hypothesis: trial users who try to invite a teammate and fail are less likely to convert.

Then the team fields a survey to trial users with an answer option taken verbatim from the interviews, "I couldn't add my team," and cuts the result by conversion. The readout now carries a share, a base, and a quote that explains the share.

Rather than presenting the interview theme as the size of the problem, the team presents it as the reason behind a number. Coding open-ended responses into themes and running a cross-tab analysis against conversion are the two analysis steps that connect the halves.

Where teams lose the handoff

The handoff most often breaks at step 3. A survey written without the interview vocabulary typically measures the team's framing, and the survey result then disagrees with the interviews for reasons that have nothing to do with the customers.

Sprig note. Sprig's Design Agent generates a programmed study from an uploaded document, so interview findings or a discussion summary can seed the sizing survey, and research panels can reach people outside the customer base. The Synthesize Agent can theme open-text responses as they arrive. Limitation: Sprig does not run the interview half of this workflow, and it does not publish sample size guidance, a weighting approach, or an estimation method. Researchers remain responsible for sizing the sample and validating the study design before launch.

How to Route Research Questions with Claude or ChatGPT

A language model can apply the Three-Question Test to a backlog of research questions and return a routing table, provided the prompt forces it to show its reasoning and check its own counts.

The prompt below classifies questions, and it does not replace the researcher's judgment about which decision each question serves.

Paste your question list into Claude or ChatGPT first, then the prompt below it. Put the data above the instructions and split long lists into smaller batches, because models tend to lose items in the middle of long inputs, a failure documented by Liu and colleagues in 2024.

Preparing the question list

Write each question as the decision owner would ask it, one per line, with the owner's name beside it. "Why are trial users churning" and "Why are trial users churning, according to the growth team" often route differently, because the second tells you whether a number is expected in the readout.

Remove questions that are really requests for a dashboard. "How many weekly active users do we have" is a query against product analytics, not a research question, and routing it wastes a row.

What this prompt does: routes each research question to an instrument using the Three-Question Test.
What it returns: a table named question-routing-table.csv and a summary of counts by instrument.

You are helping a research team choose instruments. For each question in [QUESTION LIST, default: the list pasted above this prompt]:

1. Skip any row that is blank, a heading, or a duplicate. List skipped rows by row number.
2. Split compound questions into separate rows before routing.
3. Estimate the population the question is about. If the population is under [MIN BASE, default 100] people, including zero or unknown, mark it "small or unknown".
4. Answer three tests with yes or no:
   A. Is the population small or unknown?
   B. Will the answer be reported without a number, share, ranking, or trend?
   C. Is the topic unrelated to attitudes toward AI, technology comfort, or privacy?
5. Route: all three yes = AI-moderated interview. Any no = survey, or interview then survey if the question has a reasoning half. If the question needs specialist discovery in [DOMAIN, default none], route the interview to a human moderator.

Use code to calculate this, not estimation, for every count in the summary.

Then recompute the routing a second way: for each question, write the one sentence the final readout would contain. If that sentence contains a number, share, ranking, or trend, the question needs a survey. If it describes attitudes toward AI, technology comfort, or privacy, the question needs a survey first. Otherwise it needs an interview. Compare this result with step 5 row by row.

If the two methods disagree on any row, do not silently pick one. Report every mismatch with both answers and the reason for each, and mark the row "needs researcher review".

Output: question-routing-table.csv with columns row, question, owner, population estimate, test A, test B, test C, instrument, readout sentence, mismatch flag.

Reading the routing table

Read the mismatch rows first. They are commonly compound questions that the model did not split, or questions whose decision owner has not said whether a number will be reported.

Both are problems in the question, not in the model.

Then read the survey rows with the readout sentence beside them. A sentence such as "38 percent of admins cite onboarding as the main barrier" confirms the route.

A sentence such as "admins describe onboarding as confusing" suggests the question may be an interview question after all.

What the prompt cannot do

The model cannot know the population size of your customer segments unless you tell it. It frequently guesses, which is why the prompt asks it to show the estimate.

Treat every estimate as a placeholder and replace it with a figure from your own data. The default of 100 in the prompt is also a placeholder, not a standard. Set it to the smallest base your team already reports percentages from.

Language models also tend to agree with the framing they are given, a pattern Sharma and colleagues documented at ICLR 2024.

Instead of asking the model to confirm that a question is an interview question, give it the list without a preferred answer. Researchers remain responsible for the final routing.

The Critique of AI Interviews You Should Know

The strongest critique of AI-moderated interviews is about who they reach, not how well they probe. The strongest counter-critique is that embedded AI probing reached almost everyone in a large randomized test.

A fair reading needs both, plus the critique of surveys that makes AI interviews attractive in the first place.

The case against AI interviews

Three published findings carry the critique. Verasight found completion of 40.5 percent against 99.4 percent for written open-ends, with non-random dropout that survived demographic reweighting.

A preprint study of 2,739 participants found AI probing below the human benchmark on open-ended questions. And Austin and colleagues found that answering to an AI shifted later answers toward polarization.

Each finding attacks a different claim. The first attacks coverage, the second attacks depth relative to a human, and the third attacks the idea that an AI probe is a neutral addition to an instrument.

The critique from measurement practice

Survey methodologists generally treat variation between interviewers as a source of error in interviewer-administered surveys, which is one reason self-administered questionnaires are the standard instrument for measurement. An AI interviewer removes the variation between human interviewers and replaces it with designed variation between conversations.

That design choice is deliberate and generally correct for qualitative work. But it means an AI interview inherits the measurement problem that interviewer-administered surveys spent decades controlling, and it inherits it on purpose.

The counter-critique

The counter-critique is strong and should be stated at full strength. In the Austin experiment, only 10 participants failed to complete, satisfaction barely moved, and 44 percent of qualitative feedback was positive, with six positive comments for every negative one.

That is a large randomized sample in a peer-reviewed journal, and it found no coverage penalty for embedded probing.

Nielsen Norman Group adds a practitioner version: for structured product or feature feedback, short AI interviews "could be a strong alternative to surveys." Among Verasight completers, 45 percent preferred the AI interview to typing.

The reconciliation is design, not verdict. Embedded probing and a redirect to a separate interview are different instruments, and the evidence currently favors the first on coverage and the second on depth.

Neither study tested that variable directly, so the reconciliation is a hypothesis rather than a finding.

The case against surveys

Surveys deserve their own critique, and this guide is written by a survey vendor. A free-text question nobody probes averaged 27 words per assigned respondent in the Verasight experiment.

Closed items capture only the options the researcher thought to include, and a survey written before exploratory work frequently measures the researcher's assumptions precisely.

The honest position on surveys is narrower than universal superiority. Surveys are the instrument for population claims, and they are often a weak instrument for discovering what to claim.

A survey also carries its own coverage problem. A low-response survey is exposed to the same selection risk this guide raises about AI interviews. A survey does not escape the coverage question by being a survey, which is why completers should be checked against the invited sample on either instrument.

The evidence problem on both sides

Much of the published comparison between AI interviews and surveys comes from vendors.

In the six ranked 2026 roundups of AI-moderated research platforms reviewed for this guide, each was written by a vendor, and each ranked its author first. Survey vendors publish comparisons too, and this guide is one of them.

That is why the two randomized studies carry so much weight here. The sensible order of weight is randomized evidence first, independent practitioner guidance second, and vendor material last.

What would change the verdict

Three results would move the boundary this guide draws. A randomized study showing high, non-selective completion for a standalone AI interview would weaken the coverage critique.

A published sampling and weighting method for interview-derived counts would weaken the estimation critique. And a replication of the Austin polarization finding, or a failure to replicate it, would settle whether embedded probes need the placement rules above.

Common Mistakes When Choosing Between Interviews and Surveys

The most common mistakes come from asking one instrument to do the other's job. Each is avoidable with the Three-Question Test.

Reporting interview counts as percentages

"Twelve of 40 participants mentioned pricing" becomes "30 percent of customers are concerned about pricing" somewhere between the readout and the board deck.

The first sentence is a description of participants. The second is a population claim that no interview study supports.

Write interview readouts in counts of participants and never convert them into shares.

Reporting themes without the question that produced them

In an adaptive interview, two participants who "mentioned pricing" may have been asked about it in different ways, or one may have been asked and the other not. A theme count without the prompts behind it hides that difference.

Report each theme with the discussion-guide question that surfaced it, and note where the AI interviewer probed and where the participant raised the topic unprompted.

Running the survey before the interviews

A survey written before anyone has talked to customers often has answer options in the team's language. The results look precise and measure the wrong categories.

Rather than surveying first by default, run a small round of interviews when the answer options are unknown.

Mixing interview and survey evidence in one chart

A slide that shows survey percentages beside interview theme counts invites the reader to compare them as if they shared a base. They do not, and the chart erases the distinction the readout depends on.

Keep interview evidence in quotes and participant counts, keep survey evidence in percentages with the base stated, and put them on separate slides.

Putting AI probes before the items you report

The Austin polarization result means an AI follow-up placed early in a survey can change the answers that follow it. Put probes after the reported items, and keep them off tracking waves.

Choosing the instrument you already own

Owning one instrument pulls every question toward it. A team with an interview platform runs interviews on pricing share, and a team with a survey platform runs surveys on why customers churn.

Rather than starting from the tool, start from the readout sentence. If the sentence has a number in it, the question needs a survey, whatever the team has licensed.

Accepting a vendor's completion rate as coverage

A completion rate measured against people who reached the interviewer can look high while the rate measured against people invited is low. Ask which denominator a figure uses before comparing it with a survey response rate.

Skipping the pilot

A small pilot typically reveals whether the discussion guide produces the probing the team wants, and whether invited participants actually reach the interviewer. Both failures are cheap to find in a pilot and expensive to find after a full study.

Measure completion from the invitation, not from the start of the interview, and read several full transcripts before scaling. Summaries are frequently smoother than the conversations behind them, so check the summary layer against the raw transcripts once.

Asking the AI interviewer to do specialist discovery

An AI interviewer generally follows a guide well and improvises less well. For discovery in a specialist domain, a trained human moderator is better placed to notice a thread worth pulling, according to Nielsen Norman Group.

Frequently Asked Questions About AI-Moderated Interviews and Surveys

When should you use an AI-moderated interview instead of a survey?

Use an AI-moderated interview instead of a survey when the answer is a reason rather than a number, the population is small or unknown, and the topic does not predict who will finish. Use a survey whenever the result will be reported as a share, ranking, trend, or segment comparison.

What is an AI-moderated interview?

An AI-moderated interview is a one-to-one research conversation in which a language model asks questions from a researcher's discussion guide, probes vague answers, and adapts its follow-ups to each participant. The output is a set of transcripts and summaries rather than a distribution of answers.

Is an AI follow-up question the same as an AI-moderated interview?

An AI follow-up question is not the same as an AI-moderated interview. A follow-up is a probe attached to a fixed survey, so the closed items stay comparable across respondents.

An AI-moderated interview is an adaptive conversation throughout, which produces more depth and gives up comparability.

Can AI interviews replace surveys?

No, AI interviews cannot replace surveys for population claims, because adaptive questioning gives up the identical stimulus that makes responses comparable. AI interviews can replace an unread free-text question when depth matters more than coverage, and they can come before a survey to find what it should measure.

Are AI-moderated interviews as good as human moderators?

AI-moderated interviews are generally not as good as skilled human moderators at deep discovery, according to independent studies and Nielsen Norman Group. They are often better at consistency, reach, and speed for structured qualitative work, such as feature feedback across 80 conversations.

Can AI-moderated interviews produce quantitative data?

AI-moderated interviews can produce counts, theme frequencies, and structured question types, and some platforms document methods such as MaxDiff.

A question type is not a methodology, though. Without sampling, a stated base, and controlled completion, those numbers describe participants rather than a population.

Do AI interviews have lower completion rates than surveys?

AI interviews had much lower completion than written open-ends in one randomized study of 3,160 panelists, 40.5 percent against 99.4 percent, with most loss at the handoff between methods.

A second randomized study, with AI probing embedded in the survey, found no meaningful differential dropout. Neither study tested design directly, so whether embedding explains the difference is still a hypothesis.

Should I use AI follow-up questions in a survey?

Use AI follow-up questions in a survey when an open-ended answer needs more detail, and place them after the closed items you plan to report. A randomized study in Public Opinion Quarterly found that justifying answers to an AI made later answers more polarized, so an early probe can change the measurements that follow it.

Is an AI-moderated interview cheaper than a survey?

An AI-moderated interview is not necessarily cheaper per usable response, because completion drives the cost. In the Verasight experiment, one clean AI interview required assigning 2.9 respondents, against 1.01 for a written survey response, although the authors inferred cost from those ratios rather than measuring it. Compare cost per usable response, not cost per invitation.

How do you combine AI interviews and surveys?

Combine AI interviews and surveys by running interviews to find reasons, turning each reason into a hypothesis, and fielding a fixed survey that measures it in the customers' own words. The survey supplies the number, and the interviews supply the explanation behind it.

Do I need an AI interview tool or a survey platform?

Most research programs need both an AI interview tool and a survey platform, because they produce different evidence. If your open questions are mostly about reasons and language, an interview tool carries more of the load.

If stakeholders routinely ask for base sizes, trends, and segment gaps, a survey platform carries more of the load.

What does a buyer need to ask an AI interview vendor about completion?

Ask an AI interview vendor what completion is measured against: people invited, people who started, or people who reached the interviewer.

Then ask for the fraud-flag rate on completes and a comparison of completers against the invited sample. The Verasight results show a headline completion rate can hide all three.

Bottom Line

If your question asks why, run an interview, and choose a human moderator when the domain is specialist and an AI moderator when you need breadth and consistency. If your question asks how many, how often, or which one, run a survey, and size it to the decision.

Most teams need both instruments in sequence. Interviews find the reasons, surveys size them, and the readout carries a number with its explanation beside it.

Rather than choosing a winner between the two categories, choose per question. Apply the Three-Question Test to each question, because it is built to prevent the most expensive mistake in this area, which is a percentage built from conversations.

Sprig is built for the survey half of that sequence. To size what your interviews found, upload the findings, have Design Agent draft the sizing survey, and field it to your customers or a panel.

Researchers remain responsible for the final design and for deciding which findings are ready to be counted.

Back to top
Solutions
Experience measurementStrategic & foundational discoveryJourney & behavioral researchMarket & consumer insightsConcept & prototype testing
Agents
DesignFieldAnalyzeSynthesize
Deploy
EmailPanelsWeb apps and websitesMobile app
Pricing
Community
EventsBlogGuides
CustomersIntegrationsCompare
Company
About usCareersService agreementPrivacy policyData addendumSystem status
Socials
LinkedInX