Introduction
An AI-moderated interview is a one-to-one research conversation in which an AI agent asks the questions, listens to each answer, and decides what to ask next. It works best when the topic is known well enough to write the questions in advance, when no number needs to leave the room, and when the people most willing to talk to an AI are not the people whose views you are studying.
To run an AI-moderated interview, treat it as a structured interview at scale rather than as open discovery. Write five to eight core questions, give each one written probe rules and a follow-up budget, disclose the AI before the first question, and pilot the guide before fielding. Size the study from the qualitative saturation literature, where homogeneous groups typically stop producing new themes within 9 to 17 interviews, then gross up for dropout and fraud.
Expect a coverage cost, because in one randomized study of 3,160 panelists, completion fell from 99.4 percent for a written survey to 40.5 percent for an AI interview, and the people who finished were not a random subset.
This guide covers:
- Decide whether a question suits an AI moderator, using the four-question Fit Test
- Write a discussion guide with core questions, probe rules and follow-up budgets
- Plan the number of interviews from saturation evidence, then gross up for dropout and fraud
- Disclose the AI, collect consent and meet the EU AI Act transparency rule
- Check quality, analyze transcripts, and hand findings that need a number to a survey
Sprig, which publishes this guide, makes survey software and does not sell AI-moderated interviews. The evidence cited was checked on September 25, 2026, and no prices appear anywhere in this guide.
What an AI-moderated interview is
An AI-moderated interview replaces the human moderator with a language model that runs the conversation from a discussion guide. The participant answers by voice, video or text, and the AI generates follow-up probes in real time based on what was said.
The researcher still writes the guide, sets the rules for probing, and interprets the result. The AI handles the conversation itself, which is what makes hundreds of interviews in a few days possible.
Historically, interviews scaled only as far as a moderator's calendar. AI moderation changes this by running conversations in parallel, so fielding time depends on recruitment and response rather than on how many sessions one person can run in a day.
Three things separate an AI-moderated interview from its neighbors:
- Adaptive follow-up questions
- A conversational format
- A transcript as the output
A fixed open-ended survey question typically collects one written answer and stops, while an AI-moderated interview keeps asking.
That difference is measurable, and in the randomized Verasight study described throughout this guide, follow-up probes accounted for roughly 80 percent of the extra words the interview condition produced.
An unmoderated usability test records what someone does on screen. An AI-moderated interview records what someone says about a topic, and screen observation is often limited or absent.
The term covers several formats that are often grouped together. Some platforms run a voice conversation in the browser, some add a video avatar, and some run the whole exchange as a text chat, but in every case the defining feature is a model deciding the next question.
Rather than thinking of it as an automated survey or a cheaper focus group, think of it as a structured one-to-one interview whose moderator follows written rules, and whose rules need checking in a pilot before they run at scale.
When to run an AI-moderated interview
Run an AI-moderated interview when you need reasons rather than rates, from more people than a moderator could reach.
The clearest case is explaining a known pattern. A team that already knows trial conversion dropped, and needs to hear why from 40 former trial users across three segments, is the job this method commonly does well.
It fits when the questions can be written before the first conversation. Structured topics such as onboarding friction, a cancellation decision, or reactions to a named feature generally suit an AI moderator.
It fits when speed and reach matter more than rapport. Conversations run overnight and across time zones, so fielding length depends mostly on how fast the assigned participants respond.
It fits when written open-ends have stopped producing usable answers. Austin and colleagues, in a randomized study of 2,243 respondents published in Public Opinion Quarterly in August 2026, found AI-led interviews drew 57 to 70 more words and 0.63 to 0.75 more reasons per topic than fixed survey questions.
Satisfaction in that study barely moved, falling 0.06 points on a seven-point scale even though interviews took roughly twice as long. One of its co-authors, Leib Litman, co-founded the research platform CloudResearch.
And it fits when the findings will stay qualitative. Themes, reasons, language and examples are the output. The next section covers what happens when a stakeholder wants a percentage.
The Fit Test
Answer four questions in order. A yes to all four means an AI moderator fits the job. A no at any step sends you to the instrument named on that branch.
- Can you write the questions before the first interview? No means human-moderated discovery interviews come first.
- Will the findings stay as reasons and themes, with no number leaving the room? No means a survey sizes the finding before it is reported.
- Is the topic unlikely to correlate with who is willing to finish an AI interview? No means a survey or human-moderated interviews, because the completers will skew on the topic itself.
- Can the topic be discussed without emotional risk to the participant? No means a trained human moderator.
The first three questions are the same test used in the companion buyer's guide to AI-moderated research platforms. The fourth is specific to running the interview yourself, where participant care becomes your responsibility instead of a vendor's.
When not to run an AI-moderated interview
Do not run an AI-moderated interview when you need a percentage, when the topic correlates with willingness to talk to an AI, when you need to watch what people do on screen, when you do not yet know what to ask, or when the topic carries emotional risk. Each situation has a better instrument, named below, and the first two are the most likely to produce a confident wrong answer.
You need to know how many people think or do something. This is the method's largest limitation. In the Verasight randomized trial, completion fell from 99.4 percent to 40.5 percent, and the completers differed from the people assigned.
The better instrument is a survey sized to a stated margin of error, with follow-up probes inside the survey if you also need depth. Even a perfectly random set of 100 respondents carries a margin of error of roughly plus or minus 10 points on a 50 percent figure, by standard arithmetic, and an AI study typically starts from a far less random set than that.
The topic correlates with willingness to talk to an AI. In that same trial, completers were 7.4 points more optimistic about AI than the full group assigned to the interview. Roughly three-quarters of that bias survived a full demographic adjustment.
Attitudes to AI, comfort with technology and privacy concerns are the obvious risk topics. Less obvious ones include anything that tracks age or education, since the Verasight completers also skewed away from people 65 or older and from degree holders.
The better instrument is a survey or human-moderated interviews, both of which avoid this particular self-selection step.
You need to see what people do on screen. Screen sharing is limited or unsupported on some AI moderation setups, and the Nielsen Norman Group flagged screen-based work as a poor fit in its 2026 study.
The better instrument is a moderated or unmoderated usability test. A participant describing a confusing checkout from memory is a much weaker source than a recording of the same participant getting stuck on it, and the interview will typically miss the steps people do not think to mention.
You do not yet know what to ask. AI moderators generally follow the guide rather than the participant.
The better instrument is a human-moderated in-depth interview or contextual inquiry, where a skilled moderator can abandon the guide when something unexpected surfaces. For example, a team exploring why a new market segment behaves differently from its core customers generally learns more from human conversations that can follow the unexpected than from AI conversations that stay on script.
The topic carries emotional risk. Grief, health, money trouble and workplace conflict need a moderator who can recognize distress and stop. Several AI moderation vendors advise against sensitive topics themselves.
The better instrument is a trained human moderator working to a distress protocol. A protocol typically covers how to recognize distress, when to pause or end the session, and which support resources to offer, and none of those steps can be delegated to a model with confidence today.
Rather than treating those limits as reasons to avoid the method, treat them as the boundary of the job you hire it for.
Designing the study, and the fork
The central design decision is what you are asking the AI moderator to do. Three jobs are common, and the evidence supports only one of them as a standalone use.
Use an AI moderator to explain a known topic at scale. Pair it with a human for open discovery. Hand any counting job to a survey.
The three jobs
Explain. A known topic, a structured guide, and the question of why, asked of dozens of people. Randomized evidence supports real depth gains here, and Verasight found 1.8 times the information per assigned respondent on a quality rubric that scores concrete details, specific examples, personal context and causal explanations.
Discover. Open exploration where the next question depends on what was just said. The Nielsen Norman Group, in a 2026 study of AI interviewers with 10 participants, found the moderators did not chase unexpected threads, and only 3 of 10 participants said the conversation felt natural. The authors concluded that structured interviews are likely to remain the practical ceiling for AI moderation. Rather than asking an AI moderator to explore, run a handful of human-moderated interviews first, then use what they surface to write the structured guide the AI will run at scale.
Count. Sizing or comparing segments across hundreds of transcripts. Saturation typically arrives long before that point, so a very large run spends most of its interviews counting categories that emerged during the same run. The resulting counts can look precise on a slide, but they describe a self-selected group coded against a theme list that did not exist when the study began, which is why the counting job belongs to a survey.
The modality fork
Choose text, voice or video based on what the topic needs rather than on what the platform demos best.
Text is the lowest-friction starting point and produces a transcript with no speech-recognition errors. It also suits participants in shared offices or public places, who may prefer typing to speaking aloud about their work. Voice produces longer answers, although the published ratios come from vendor data, not independent studies. Wuttke and colleagues, in a 2026 preprint with 571 respondents who could choose voice or chat, found satisfaction at or above standard surveys in both modes, although the study had no human-interviewer comparison.
Video adds faces, and faces can backfire. Zhu and Broadbent, in a 2025 randomized trial of 160 participants in Computers in Human Behavior, found a human-like virtual interviewer increased socially desirable answers on sensitive items compared with a text chatbot.
The oversight fork
Rather than choosing between fully automated and fully human, build oversight into three points.
- A researcher reviews the discussion guide and the probe rules before launch.
- A researcher reads the first 10 to 20 transcripts before the rest of fielding continues.
- A few human-moderated sessions run alongside the AI sessions when the decision is high-stakes.
The review count in step two comes from published platform guidance, not research. Set it to whatever lets you catch a broken probe before it repeats across a hundred conversations.
The recruitment fork
Recruit from your own customers or from a panel. No published study compares the two for AI-moderated interviews.
Panels bring reach and a visible cost.
In the Verasight trial, 16 percent of completed AI interviews were flagged for fraud, and one clean interview required assigning 2.9 respondents.
Your own customers bring context you can attach to each transcript, such as plan, tenure and usage, but only people who agree to talk. That context is often more valuable than reach, because a theme that clusters among customers on one plan or in their first month is a much sharper finding than the same theme floating free.
Platform settings differ on every fork, so check your tool's support for modality, screen sharing, follow-up limits and probe instructions first.
Writing the discussion guide
To write a discussion guide for an AI moderator, draft five to eight core questions, order them from a concrete recent event to reflection, keep each question to one idea, and give every question a goal, the exact wording, conditional probe rules and a stop condition. Then open the guide with a plain disclosure that the participant is talking to an AI.
A discussion guide for an AI moderator is more explicit than one written for a person. A human moderator fills gaps with judgment, while an AI moderator fills them with whatever the model does by default, which is often a generic "can you tell me more?"
Write five to eight core questions, a narrower band than the published platform range, chosen to leave room for probes. Published platform guides from 2025 and 2026 generally recommend a range of five to ten core questions and a total length of 10 to 20 minutes for voice, with shorter sessions for text.
Those figures are vendor guidance rather than research findings, so treat them as a starting point and let the pilot set the final length.
Order the questions from concrete to reflective. Rather than opening with an opinion question such as whether the product is easy to use, open with a recent event the participant can describe, then move to decisions, alternatives and finally evaluation.
Keep each question to one idea. A double-barreled question such as "what did you like and dislike about onboarding" typically gets half an answer from a person and a confused follow-up from an AI moderator, which then probes whichever half it heard.
Give every core question four parts:
- A goal
- The question wording
- Probe rules
- A stop condition
The goal tells the moderator what the question is for. The stop condition tells it when the answer is good enough to move on, such as a substantive answer that includes at least one specific example.
The discussion guide template
Here is a lift-ready guide for a common structured topic: why trial users of a project-management tool did not convert. Replace the product and the topic, and keep the structure.
- Warm-up. Goal: context. "Tell me about the project you were working on when you started the trial." Probe once if the answer names no specific project.
- Expectation. Goal: the job the trial was hired for. "What were you hoping the tool would help you do?" Probe if the answer describes features instead of an outcome. Budget: up to two follow-ups.
- Last session. Goal: a real event, not an average. "Think about the last time you opened the tool during the trial. What were you trying to do?" Probe for what happened next. Budget: up to three follow-ups.
- Friction. Goal: the moment it stopped working. "Was there a point where it stopped being worth the effort?" If they mention price, ask what they were comparing it to. If they mention a teammate, ask what the teammate said. Budget: up to three follow-ups.
- Alternative. Goal: what won instead. "What are you using for that work now?" Probe once for why.
- Decision. Goal: who decided and how. "Who else had a say in whether to keep using it?" Probe once if the answer is "nobody."
- Close. Goal: anything missed. "Is there anything about the trial we did not ask about that mattered to you?" No probes.
Seven questions with these budgets typically fit inside the 10-to-20-minute range once the pilot confirms timing.
Notice what the template avoids. Rather than asking whether people liked the product, it asks about specific events, decisions and alternatives, which are the answers people can give accurately.
Disclosure and consent at the top of the guide
Open every interview with a plain disclosure before the first question. Here is a version to adapt:
"You'll be talking with an AI interviewer, not a person. It will ask follow-up questions based on your answers. The conversation is recorded and transcribed. [STATE WHETHER RESPONSES ARE USED TO TRAIN AI MODELS.] You can skip any question or stop at any time."
The consent section later in this guide covers why that line is a legal requirement in some markets and not only good practice.
Probe design and follow-up budgets
Probes are where AI-moderated interviews generally earn their depth, and where they most often go wrong.
Write probe rules as conditions rather than as general instructions. "If the participant mentions price, ask what they compared it to" produces a better follow-up than "probe for detail," because the moderator knows what detail is for.
Set a follow-up budget per question. Platforms commonly let you choose a range, from zero follow-ups to five or more, and some let you leave it automatic. The budget governs depth, not a guaranteed duration.
Spend the budget where the decision lives. In the template above, the friction and last-session questions get three follow-ups each, and the warm-up gets one.
Watch the total, since seven questions with an average of two follow-ups each is roughly 21 exchanges, and the pilot's median duration shows whether that total fits the length you set.
Put moderator-only guidance in the field meant for it. Many platforms separate the text participants see from instructions only the moderator reads, and mixing the two often leaks your hypotheses into the conversation.
Four probe rules that prevent the common failures
- Ban praise words such as "fascinating" or "great answer" in the moderator instructions
- Require one probe for a specific example before accepting a general answer
- Prohibit probes that introduce a reason the participant did not mention
- Cap total follow-ups so the last questions still get fresh attention
The first rule answers a documented problem. The Nielsen Norman Group observed AI interviewers praising answers, which signals to participants what the moderator wants to hear.
The third rule matters more than it looks. A probe such as "was it because of the price?" hands the participant a reason, and a sycophantic model will commonly accept the agreement as a finding.
Rather than trusting the rules to work, read every pilot transcript specifically for probe quality. Look for leading follow-ups, repeated generic prompts and questions skipped because an earlier answer seemed to cover them.
How many AI-moderated interviews you need
Plan the number of interviews from the saturation literature, then gross it up for dropout and fraud. Vendor claims of hundreds of interviews describe what a platform can run, not what a question needs.
What the saturation evidence says
Three peer-reviewed sources carry this section.
Guest, Bunce and Johnson, in Field Methods in 2006, coded 60 interviews and found 73 percent of all codes applied to the Ghana data had appeared within the first six transcripts, and 92 percent within twelve.
Hennink and Kaiser, in a 2022 systematic review of 23 studies in Social Science and Medicine, found interview studies reached saturation within 9 to 17 interviews for homogeneous populations with narrowly defined objectives.
Guest, Namey and Chen, in PLOS ONE in 2020, proposed measuring saturation directly as the share of new information in each run of interviews. At a threshold of 5 percent new information, their datasets reached saturation at a median of six or seven interviews.
None of these studies tested AI-moderated interviews. No published study has measured saturation for AI moderation specifically, so this guide applies the human-interview evidence and says so.
The derived planning table
The table below is derived, not copied. It multiplies a per-segment target by the number of segments, then divides by the share of assigned participants who produce a clean interview. Assigned participants are everyone offered the interview, whether or not they started it.
The low target of 12 per segment comes from Guest's twelve-interview mark. The high target of 17 is the top of the Hennink and Kaiser range. The 2.9 assignment ratio is Verasight's measured figure from one panel study.
| Segments | Clean interviews at 12 each | Clean interviews at 17 each | Assigned at 2.9 per clean interview, 12 each | Assigned at 2.9 per clean interview, 17 each |
|:--------:|:---------------------------:|:---------------------------:|:--------------------------------------------:|:--------------------------------------------:|
| 1 | 12 | 17 | 35 | 50 |
| 2 | 24 | 34 | 70 | 99 |
| 3 | 36 | 51 | 105 | 148 |
| 4 | 48 | 68 | 140 | 198 |
The formula is simple: clean interviews equal the per-segment target times the number of segments. Assigned participants equal clean interviews times the assignment ratio, rounded up.
The Verasight ratio is a pessimistic case, because it measured a panel sample that was not told in advance about the interview. Your own customers, told up front, may convert at a better rate, although no published study has tested it, so replace 2.9 with your pilot's own ratio as soon as you have one.
The two published figures agree with each other. A completion rate of 40.5 percent, with 16 percent of completes flagged, leaves about 0.34 clean interviews per assignment, which is about 2.9 assignments per clean interview.
When larger runs make sense
Run more interviews when you need to compare segments qualitatively, which often means one saturation target per segment, not when you need a percentage.
A theme that appears in 30 of 100 AI interviews is not a 30 percent finding. The completers are self-selected, and the themes emerged during the run.
If a stakeholder needs the percentage, the finding goes to a survey. The same rule applies to comparisons between segments: an AI study can show that one segment talks about price and another talks about permissions, but whether that difference holds in the population is generally a question for a sized sample.
Recruiting and screening participants
Screen on the behavior the study is about, not on willingness to be interviewed, which typically selects for the most talkative customers. For the trial-conversion example, the screener confirms that the person started a trial in the last 90 days and did not convert.
Keep screeners short. A long screener in front of a long interview typically compounds dropout, and the people who push through both are often the most motivated rather than the most representative. Rather than screening on attitudes, which invites participants to guess the answer that qualifies them, screen on facts such as dates, plans and actions that can often be checked against your own records.
Record who was assigned, who started and who finished. Rather than reporting only completed interviews, keep the full funnel, because the gap between assigned and finished is where self-selection hides.
Set quotas on the segments you plan to compare. An AI study without quotas risks over-representing whoever is most willing to finish, and in Verasight's trial those completers were more optimistic about AI than the people assigned.
Match the incentive to the burden. A 15-minute spoken interview asks more of a participant than a five-minute survey, and the Verasight authors named a higher incentive for the heavier task as a test worth running.
Tell participants what they are signing up for. The Verasight authors named the lack of advance notice as a possible cause of the dropout they measured, and most of the loss happened at the handoff from the survey to the interview platform.
Check technical readiness in the invitation. Voice interviews need a working microphone, a supported browser and a quiet place, and each is a common reason a willing participant never reaches the moderator.
Plan for mobile. Many participants may open an invitation on a phone, and some AI moderation setups do not support screen sharing on mobile, so test the full path on the devices your participants actually use.
Set a contact-frequency limit for your own customer list so the same people are not interviewed repeatedly, and record each participant's recent research contact as a screener field.
Consent, disclosure and fielding
To get consent for an AI-moderated interview, disclose that the participant is talking to an AI before the first question, collect consent to record and transcribe, and state whether responses train AI models and how long they are kept. In the European Union, disclosing the AI is a legal requirement, and voice or face data can trigger written-consent rules in some US states.
The EU AI Act transparency rule
Article 50 of the EU AI Act has applied since August 2, 2026. It requires AI systems that interact with people to inform them they are interacting with an AI, unless that is obvious to a reasonably well-informed, observant person.
Article 50 also requires deployers of emotion recognition systems to inform the people exposed to them. That matters if your platform labels emotion from voice or video.
Article 5 of the same act prohibits emotion recognition in workplaces and educational settings, with narrow medical and safety exceptions. Employee research that uses emotion labels needs a hard look before launch.
Voice and face data in the United States
Illinois's Biometric Information Privacy Act covers voiceprints and face geometry. It requires a written policy and written consent before collection, and an amendment in August 2024 made electronic signatures valid consent.
If a platform uses face matching or voiceprints for fraud checks, those checks may count as biometric collection.
Data use and retention
Ask the platform three questions before fielding: whether recordings and transcripts are used to train its models, how long they are retained, and whether you can delete them on request.
Put the answers in your consent language instead of a separate policy participants will not read. Some vendors publish that they do not train on study data and delete it within a fixed window, but terms vary, so confirm them for your own contract.
This section is general information, not legal advice. Confirm the rules for your participants' locations with your own counsel before fielding.
Fielding the study
Pilot first by running five to ten interviews, a common practitioner range, reading every transcript, and fixing the guide before launching the rest.
Then field in a single window. Rather than letting a study trickle along for weeks, open it, fill the quotas and close it, so every interview describes the same moment.
Read transcripts as they arrive during the first day. A broken probe or a misunderstood question typically repeats in every interview until someone stops it.
Timing and cadence
Run an AI-moderated interview study when a specific decision needs reasons within days rather than weeks.
A typical structured study moves from guide to readout in about one to two weeks: two or three days to write and pilot the guide, a few days of fielding, and several days of reading and synthesis. That timeline is a practitioner estimate, not a sourced figure.
No source publishes a repeat cadence for AI-moderated interviews, and this guide does not invent one.
Repeat a study when the thing being explained has changed, such as a new onboarding flow or a new pricing page. Keep the guide identical across repeats so differences come from participants, not wording.
Avoid fielding across a known disruption. A study that spans a pricing change or a major release commonly mixes two populations in one set of transcripts.
Time the study to the event it explains. Interviewing trial users within a few weeks of the trial ending generally produces sharper recall than interviewing them months later, when the decision has been reconstructed into a tidier story.
Quality control and fraud
Check three things in every AI-moderated study: fraud, participants using AI to answer, and moderator failures.
Fraud. The only published rate comes from the Verasight trial, where 16 percent of completed AI interviews were flagged. The report does not state how fraud was defined, so treat the figure as one study's measure, not a norm.
Participants using AI. Text interviews are exposed to participants pasting answers from a chatbot. Veselovsky and colleagues, in a 2023 preprint, estimated that 33 to 46 percent of crowd workers on one platform used a language model on a writing task. Westwood, in PNAS in 2025, showed an AI agent could pass survey attention checks while holding a consistent persona.
Voice and video raise the cost of this kind of fraud, but no published study measures by how much. Rather than assuming a conversational format is a fraud barrier, check for it.
Moderator failures. Read a sample of transcripts for skipped questions, leading probes and interviews that ended early.
The transcript review checklist
- Flag answers with uniform polish, perfect grammar and no personal detail across every question.
- Flag interviews where answers contradict the screener, such as a non-converter describing a renewal.
- Flag duplicate or near-duplicate transcripts across participants.
- Flag interviews far shorter than the pilot's median duration. Half the median is a working convention, not a sourced cutoff.
- Record exclusions with the rule used, the count removed and the final count by segment.
Report exclusions with the findings. A theme that holds only after removing a fifth of the interviews needs that context attached. Rather than deleting flagged interviews silently, keep them in a separate file with the reason attached, so a reviewer can check whether the exclusions changed the story.
The saturation check
Measure saturation during fielding rather than assuming it. The method from Guest, Namey and Chen is simple enough to run in a spreadsheet.
Code the first four interviews and count the distinct themes. That count is the base, and it typically grows quickly before it flattens.
Then code the next run of two or three interviews and count only the themes that did not appear before. Divide the new themes by the base count.
When that share falls to 5 percent or less, the run has reached saturation at that threshold.
Here is a worked example with illustrative numbers. The first four interviews produce 20 themes. Interviews five and six add 2 new themes, which is 10 percent of the base. Interviews seven and eight add 1, which is 5 percent, so this segment reached saturation at eight interviews on a 5 percent threshold.
Run the check per segment, not across the whole study. A theme set that saturates for one segment says nothing about another. A study with three segments therefore runs three saturation checks, and the segment that takes longest to saturate usually sets the size of the whole study.
Rather than stopping at the first run that meets the threshold, confirm it with one more run. Saturation claimed on a single quiet run is commonly a coincidence.
Benchmarks and what a good result looks like
There is no independent, cross-industry benchmark for AI-moderated interview completion, length or response depth, and the published figures come from single studies whose designs differ enough that their completion rates contradict each other.
Verasight measured 40.5 percent completion in a panel study where the interview sat on a separate platform and participants had no advance notice. Austin and colleagues reported minimal drop-off in a study where the AI interview ran inside the survey itself, with no platform redirect.
The likeliest explanation for that gap is the handoff. Verasight reported that most of its loss happened at the redirect between the survey and the interview platform, and its authors name failed browser redirects, disabled microphones and respondents in public spaces as possible causes.
Depth figures vary the same way, and Verasight reported 128 words per assigned respondent against 27 for written answers, and 291 words among people who answered at all. Geiecke and Jaravel, in the 2024 working-paper version of their study, reported 460 words for AI interviews against 190 for open text.
Those numbers describe different questions, audiences and interview lengths, so none of them is a target for your study.
Vendor figures for depth, length and candor are often cited without a method. Treat them as marketing until a method is published.
Build your own benchmark instead. Measure completion from everyone assigned, not from starts, and record median duration, words per interview and the fraud-flag rate for every study. Your second study then has something valid to compare against.
A good result is defined by the decision, not by a number. A structured study has done its job when the team can name the main reasons behind the pattern it set out to explain, quote participants on each one, and write the closed answer options for the survey that will size them.
Interpreting and acting on the result
Read the transcripts for reasons, language and mechanisms, and report them as such. An AI-moderated study licenses statements about why and how, not about how many.
Start with the funnel. State how many people were assigned, how many started and how many finished, because the reader needs to know how self-selected the evidence is before reading any theme.
Then the themes, each generally supported by two or more quotes. Report how many interviews raised a theme only as a rough indicator of prominence, and never as a percentage of a population.
Then the mechanisms. The most useful output of a structured study is usually a chain: a trigger, the moment of friction, and the alternative that won. For the trial-conversion example, that might read as a teammate refusing to adopt the tool, followed by a switch back to spreadsheets.
Then the contradictions. Themes that split along a segment line often point to the most useful follow-up question. Rather than smoothing those splits into a single story, report them as open questions, because a split between segments is the first thing the sizing survey should test.
What the result licenses
A structured AI-moderated study licenses hypotheses, language for messaging, and the answer options for a follow-up survey. It does not license a market share, a prevalence rate or a claim that one theme outweighs another in the population.
Rather than presenting a theme count as a finding, turn the strongest themes into a closed question and field it to a sample that represents the population.
In Sprig: the Design Agent can ingest a research brief and propose a structured survey with question flow and logic, and written interview findings can serve as that brief, so interview themes can become the answer options of a sizing study. Sprig does not run AI-moderated interviews, and the Design Agent drafts from your written findings rather than from transcripts as data. A researcher decides what goes live.
Analyzing transcripts with Claude or ChatGPT
Analyzing AI interview transcripts is open-text theming, and the mechanics of theming already have their own guide. Theme creation, the recode loop and checking quotes as a coding step belong to the advanced cross-tab guide, and comparing themes across segments belongs to the simple cross-tab guide.
This section covers only what is specific to interview transcripts: separating core answers from probe answers, verifying quotes against the source, and keeping counts out of the readout.
Three checks specific to transcripts
Quotes get invented. Wachinger and colleagues, in Qualitative Health Research in 2024, found ChatGPT produced themes that overlapped considerably with a human analyst's, but it misquoted or invented quotes and stayed descriptive where the human found latent themes. Verify every quote against the transcript with an exact text match.
Probe answers are not core answers. An answer given after three follow-ups was shaped by the moderator's questions. Tag each quote with whether it came from a core question or a probe.
Emergent themes do not get percentages. A theme discovered in the same run it is counted in has no stable definition to count against.
The transcript coding prompt
WHAT THIS PROMPT DOES: applies my theme list to AI-moderated interview
transcripts and extracts verified supporting quotes.
WHAT IT RETURNS: a quote file, a respondent-by-theme matrix and a method note.
Files: [PATH TO TRANSCRIPT FILES, no default, I must supply it]
Theme list: [PASTE YOUR THEME LIST WITH ONE-LINE DEFINITIONS, no default]
Segment field: [SEGMENT COLUMN, default segment]
Minimum interviews per theme: [THRESHOLD, default 3, a working convention]
Use code to calculate this, not estimation. Do not count anything by
reasoning about it in prose.
Rules:
1. Code only against my theme list. Put anything that fits no theme in
"Other or unclear". Do not create new themes.
2. Exclude any interview with a blank transcript or no answer to any core
question, and list the excluded interview IDs.
3. For every coded passage, save the exact quote, the interview ID, and
whether it answered a core question or a probe.
4. Verify every quote with an exact string match against its transcript.
Report any quote that does not match, and do not paraphrase it into a
match.
5. Count interviews per theme, not mentions. Any theme raised in fewer than
the minimum, including zero, is reported as below threshold.
6. Report counts as interviews out of the total. Do not report percentages.
Recompute every theme count a second way: in a fresh session, code each
transcript again from the theme list without reading the first pass, then
compare interview counts per theme. If any count differs by more than
[TOLERANCE, default 1 interview], say you cannot reconcile it and show both
passages. Do not choose the more plausible number.
Save as aimi_quotes.csv, aimi_theme_matrix.csv and aimi_method_note.md.
The general failure modes of language-model analysis apply here too. Hill and colleagues, in PLOS Digital Health in 2026, found models far better at applying a codebook than at discovering themes, and Ashwin, Chhabra and Rao, in Sociological Methods and Research in 2025, found rare codes systematically over-predicted.
Thinking Machines Lab reported in 2025 that 1,000 identical requests at temperature zero produced 80 different outputs, so produce any reportable count twice. Liu and colleagues showed in 2024 that material in the middle of long inputs typically gets missed, so split long transcript sets.
Sharma and colleagues documented at ICLR 2024 that models tend to agree with the user, so never ask a model to confirm a theme you already suspect. Transcripts often contain names, employers and places, so check the data terms of the tool before uploading.
Researchers remain responsible for the theme list, every quote in the readout, and what the findings mean.
The critique you should know
Four objections apply to AI-moderated interviews specifically. The counter-evidence is real, and it is narrower than vendor marketing suggests.
Selection: the people who finish are not the people you asked
This is the strongest objection, and it comes from a randomized design.
Verasight's Depth at a Cost, published August 31, 2026 by Morris, Bakhshi, Leff, Rothschild and Marshall, assigned 3,160 panelists to written open-ends or an AI-moderated video or audio interview. The experiment was conducted jointly with Outset, an AI interview platform, and the report is an industry publication, not a peer-reviewed paper.
Completion fell from 99.4 percent to 40.5 percent. Completers were more likely to be male and Black, less likely to be Hispanic, 65 or older, degree holders or from six-figure households, and leaned 6.2 points more Republican.
The authors concluded that "no qualitative richness compensates for losing six in ten people non-randomly when the estimand is a population share."
That finding runs against the partner platform's commercial interest, which adds to its credibility. The authors also list its limits: participants were not told in advance, incentives were not raised for the heavier task, and the study covered one panel, one platform and four topics.
Moderation quality: structured, not skilled
The Nielsen Norman Group's 2026 study found AI interviewers interrupted participants, praised answers, summarized without letting participants respond, and did not follow unexpected threads.
The study was small, with 10 participants, and it tested two tools. It remains one of the few independent observations of AI moderators at work.
Measurement: the interview changes the answer
Austin and colleagues found that following a closed question with an AI interview slightly polarized the closed answers, with reported effects of 0.064 and 0.045 on the two topics measured.
Barari and colleagues, in a 2025 preprint with 1,800 respondents, found live AI coding inflated false positives through acquiescence, meaning participants agreed with the moderator's framing.
Scale: counting before defining
Carl Pearson, a research practitioner, argued in a March 2026 blog post that large AI interview runs quantify qualitative data before the categories are defined, because researchers "can't effectively count 'it' before you know what 'it' is." The objection is not peer-reviewed, but it matches what the saturation literature predicts: most themes appear early, so the later interviews in a very large run mostly add to counts, not to understanding.
The counter-critique
Three findings push the other way.
Depth is real and repeatable. Verasight found 1.8 times the information per assigned respondent, and the gain survived the dropout. Austin and colleagues found more words and more reasons at almost no cost to satisfaction, and both studies are randomized.
Participants often prefer it. In the working-paper version of Geiecke and Jaravel's study, since listed as accepted at the Review of Economic Studies, 43 percent of respondents said they would prefer an AI interviewer next time, against 19 percent who preferred a human. Among Verasight's clean completers, 86 percent rated the experience 4 or 5 out of 5.
People may disclose more to a machine. Lucas, Gratch, King and Morency, in Computers in Human Behavior in 2014, found people who believed a virtual interviewer was automated feared self-disclosure less and were rated as disclosing more. In the Verasight study, 34 percent of clean completers said they were more honest with the AI, although that is a self-report from people who chose to finish.
Where this guide lands
Depth is the method's strength and coverage is its cost, and both findings come from randomized designs. Use AI moderation to explain, report what it finds as reasons, and send any number to a survey.
Common mistakes
Treating a theme count as a percentage. Thirty of 100 self-selected interviews is not 30 percent of anyone. Report counts as interviews, and size anything important with a survey.
Running discovery on an AI moderator. A moderator that follows the guide cannot find what the guide does not ask about. Run a few human-moderated interviews first.
Writing probes as "tell me more." Generic probes produce generic answers. Write conditional probe rules tied to the decision.
Skipping the pilot. A broken probe repeats in every interview. Read every pilot transcript before launching the rest.
Measuring completion from starts. Completion measured from people who reached the moderator hides the handoff loss. Measure it from everyone assigned.
Ignoring who dropped out. In the one randomized study, completers were more optimistic about AI. Compare completers with everyone assigned on every variable you hold.
Letting quotas fill first come, first served. Completers in the one randomized study skewed toward AI optimism. Set quotas on the segments you plan to compare.
Leaving disclosure until the end. Participants should know they are talking to an AI before the first question, and in the European Union it is a legal requirement.
Using emotion labels without a check. Emotion recognition triggers disclosure duties under the EU AI Act and is prohibited in workplace settings.
Trusting every quote in an AI summary. Language models misquote. Match every quote in the readout against its transcript.
Asking leading probes. A probe that names a reason teaches the participant the answer. Prohibit it in the moderator instructions.
Changing the guide mid-field. An edit halfway through splits the study into two instruments. Fix the guide at the pilot and freeze it.
Skipping the invitation details. Participants who do not know the session is a spoken AI interview, or who open it on a phone without a working microphone, often drop out at the handoff. State the format, the length and the device requirements in the invitation.
Mistakes in analysis and reporting
Reading only the summary. Platform summaries are useful for triage, but they commonly flatten the contradictions that make a finding useful. Read a full sample of transcripts from every segment before trusting any summary.
Mixing probe answers with core answers. An answer given after three follow-ups was shaped by the moderator's questions, and it typically sounds more articulate than the participant's first response. Tag each quote with where it came from.
Letting the model discover and count in one pass. A model asked to find themes and count them at once is counting against categories it invented a moment earlier. Build the theme list first, then apply it.
Reporting without the funnel. A readout that shows 60 interviews without saying 200 people were assigned hides the most important fact about the evidence. Put the funnel on the first page.
Synthetic respondents
Synthetic respondents are not a substitute for AI-moderated interviews.
An AI interviewing a real person collects that person's experience. An AI answering as a simulated person generates plausible text about an experience nobody had.
Some platforms now offer synthetic personas or digital twins beside their interview products. Rather than treating those outputs as evidence, use them to pressure-test a draft discussion guide for confusing wording or missing probes before a real participant sees it.
Nothing a synthetic respondent says belongs in a readout as a finding, however fluent it sounds, and a study that mixes synthetic and real transcripts cannot say which of its themes came from people. Bisbee and colleagues, in Political Analysis in 2024, found synthetic survey responses matched human averages closely while 48 percent of regression coefficients differed significantly from the human data, and an interview study is almost entirely about relationships between a trigger, a decision and an outcome.
Where the interview ends and the survey begins
An AI-moderated interview ends where a finding needs a number. The handoff is a survey built from what the interviews found.
The sequence has four steps:
- Interview to find the reasons, in the participants' own words.
- Turn the strongest themes into closed answer options.
- Field a survey to a sample that represents the population.
- Report the number with its margin of error, and quote the interviews as the explanation.
Each instrument does the job it is built for. The interviews supply the options a survey writer would otherwise guess at, and the survey supplies the coverage the interviews cannot.
Here is how the trial-conversion example travels through the sequence. Forty interviews surface five reasons for not converting, including a teammate refusing to adopt the tool and confusion about permissions, and the survey then asks every recent non-converter which of the five applied, with an "other" option for anything the interviews missed.
The interviews also improve the survey's wording. Participants' own phrases often make better answer options than a researcher's labels, because they describe the problem the way the people answering the survey will recognize it.
Sprig makes the survey half of this sequence. A team can field the sizing study to a panel, to its own customers by email, or inside its product, depending on who the interviews could not reach.
Sprig does not run AI-moderated interviews, voice moderation or emotion analysis, so the first step happens on an AI interview platform or with a human moderator.
In Sprig: upload the written findings from your interviews and the Design Agent drafts the follow-up survey, with logic and question flow proposed from the brief. The survey then fields to the people your interviews could not represent. Sprig does not run AI-moderated interviews. The agent drafts, and a researcher decides what goes live. Sprig publishes no sample-size guidance of its own, so size the study with standard margin-of-error arithmetic.
Alternatives and adjacent methods
Three families of methods sit next to AI-moderated interviews, and they belong to different categories.
Qualitative methods. Human-moderated in-depth interviews when you need discovery, rapport or care. Focus groups when you need to hear people react to each other. Diary studies when the behavior unfolds over days or weeks.
Usability methods. A moderated or unmoderated usability test when you need to see what people do on screen rather than hear what they say about it.
Survey methods. A closed-question survey when you need to size a finding across a population with a stated margin of error. A survey with open-ended questions when you need a written answer from everyone in a representative sample. A survey with AI follow-up questions when you want some depth inside the survey. Whether that avoids the dropout Verasight measured is untested, and a follow-up inside a survey is still a survey question, not an interview.
The three families answer different questions. An AI-moderated interview is a qualitative tool that reaches more people than a human moderator, and it does not become a survey by running at scale.
Choosing between them generally comes down to two questions. The first is whether you need to hear reasoning or measure prevalence, and the second is whether you need to watch behavior or hear about it. Reasoning without observation points to interviews, prevalence points to surveys, and observation points to usability methods, whichever moderator runs the session.
Frequently asked questions
What is an AI-moderated interview?
An AI-moderated interview is a one-to-one research conversation in which an AI agent asks questions from a discussion guide and generates follow-up probes in real time. Participants answer by voice, video or text, and the output is a transcript the researcher analyzes.
When should you use an AI-moderated interview?
Use an AI-moderated interview when you can write the questions in advance, when the findings will stay as reasons rather than numbers, and when the topic is unlikely to correlate with who is willing to talk to an AI. It commonly works best for explaining a known pattern, such as why trial users did not convert.
How many AI-moderated interviews do you need?
No study has measured saturation for AI-moderated interviews specifically. Human-interview evidence shows homogeneous groups typically stop producing new themes within 9 to 17 interviews per segment. Multiply by the number of segments, then gross up for dropout and fraud using your pilot's assignment ratio.
How long should an AI-moderated interview be?
Published platform guidance generally puts voice interviews at 10 to 20 minutes, text interviews shorter, and core questions at five to ten. Those are vendor recommendations, not research findings, so set the final length from your pilot's median duration.
Are AI-moderated interviews as good as human moderators?
AI-moderated interviews produce more depth than written survey answers in randomized studies, but a small independent study found AI moderators interrupted participants, praised answers and did not follow unexpected threads. They generally suit structured topics, and human moderators remain the better choice for discovery and sensitive subjects.
When should you not use an AI-moderated interview?
Do not use an AI-moderated interview when you need a percentage, when the topic correlates with willingness to talk to an AI, when you need to watch screen behavior, when you do not yet know what to ask, or when the topic carries emotional risk. A survey, a usability test, human-moderated discovery interviews and a trained human moderator are the better instruments across those five situations.
How do you get consent for an AI-moderated interview?
Get consent for an AI-moderated interview by telling participants before the first question that they are talking to an AI, asking permission to record and transcribe, and stating whether responses train AI models and how long they are kept. Where a platform collects voiceprints or face geometry, some US states such as Illinois require written consent.
Do people answer honestly when talking to an AI?
People may disclose more to an interviewer they believe is automated. A 2014 study in Computers in Human Behavior found lower fear of self-disclosure with an automated virtual interviewer. But a human-like virtual interviewer increased socially desirable answers compared with a text chatbot in a 2025 trial, so a text format may be the safer choice on sensitive items. Voice was not tested in that trial.
Do you have to tell participants they are talking to an AI?
Participants should always be told before the first question. In the European Union, Article 50 of the EU AI Act has required disclosure of AI interaction since August 2, 2026, unless it is obvious to a reasonably well-informed person. Emotion recognition carries a separate duty to inform.
Can AI-moderated interviews produce quantitative data?
AI-moderated interviews can include structured questions, but the completers are typically self-selected, so their answers do not represent a population. In one randomized study, completion fell to 40.5 percent and completers skewed on the study topic. Size any number with a survey.
How do you analyze AI interview transcripts with Claude?
Build a theme list from your first reading, then have Claude apply that list, extract quotes with exact-match verification, and count interviews per theme using code. Theme mechanics are covered in the advanced cross-tab guide, and every quote in the readout should be checked against its transcript.
How do you detect fraud in AI-moderated interviews?
Flag answers with uniform polish and no personal detail, answers that contradict the screener, duplicate transcripts and interviews far shorter than the pilot's median length. One randomized panel study flagged 16 percent of completed AI interviews for fraud.
What drives the cost of an AI-moderated interview study?
The cost of an AI-moderated interview study is driven by the platform plan, the number of participants you must assign to reach your clean-interview target, the incentive per participant, and researcher time for piloting and reading transcripts. The assignment ratio often matters most, because at 2.9 assigned participants per clean interview, recruitment scales with nearly three times the interviews you keep.
How do you choose an AI-moderated interview platform?
Choose an AI-moderated interview platform by checking six things against your study design: supported modalities, control over probe rules and follow-up budgets, fraud checks, data training and retention terms, built-in AI disclosure and consent, and whether completion can be reported from everyone assigned. A companion buyer's guide covers the vendor comparison.
How do you combine AI-moderated interviews with a survey?
Use the interviews to find reasons in participants' own words, turn the strongest themes into closed answer options, and field a survey to a representative sample. The survey supplies the number, and the interviews supply the explanation.
The bottom line
An AI-moderated interview is a structured interview that scales. Used for that job, it produces more depth than a written survey answer from more people than a moderator could reach.
Its limits are just as clear. It reaches a self-selected group, it follows the guide rather than the participant, and it cannot turn a theme into a percentage. Two of those limits have design answers in this guide, and the third, the percentage, is answered by handing the number to a survey.
Run the Fit Test first. Write a guide with conditional probes and a follow-up budget, pilot it, and size the study from saturation evidence. Disclose the AI, measure completion from everyone assigned, and verify every quote.
If the question needs reasons from more people than a moderator can reach, an AI moderator is often a practical way to get them. If the question needs a number, the interviews supply the answer options and a survey supplies the answer.