The first time I handed a real piece of thinking to AI, I expected to get an afternoon back. I didn't.
AI was acting useful as a thought partner, asking me things I hadn't asked myself. But I remember sitting there mid-session thinking that if I could just get a colleague on a call for fifteen minutes, we'd be at an answer faster than the time I was spending feeding context to a machine. The tool worked exactly as advertised; it just didn't do the thing I'd quietly assumed it would do.
That gap between what a tool does and what we assumed it would do, is where almost every AI walk-back in research lives. And there have been a lot of walk-backs this year. Nobody publishes a post about the workflow they rolled back, but if you talk to enough research leaders, the same story surfaces over and over.
The tools aren't failing, but they're drifting into work they were never supposed to own that nobody signed off on.
Most teams I talk to now are planning to cut their research stack, collapsing four or five point tools into one or two that actually talk to each other. Procurement is tired, and so is everyone who has to remember which tool holds which study.Consolidation is the right instinct and forces teams to address what decision-making responsibility each tool should actually be allowed.
Every research stack has three layers
The mechanical layer is the work that has always been slow: transcription, cleaning, deduping, repairing a broken SQL query, sanity-checking a translation across markets, and reformatting the same finding for the fourth stakeholder. . This you can give all away to AI.
The interpretive layer is first-pass pattern work: clustering open-ends, surfacing candidate themes, and flagging where a segment diverges. AI is genuinely good but incomplete here, which is a harder combination to manage than just being bad. AI will confidently surface one theme and quietly drop three others here, and the dangerous thing is that the output reads the same either way.
The judgment layer is what counts as evidence: whether a finding is strong enough to put in front of leadership and what the business should actually do about it. This layer is fully yours. This is the only layer where being wrong is expensive and invisible at the same time.
Walk-backs happen because a tool drifted up a layer without permission.
The most common walk-back patterns
Most commonly, AI drifts up a layer because of a lack of guardrails or sloppy instructions. The most tempting move in the stack and the one team's reverse fastest is pointing AI directly at raw research data and asking for the finding. The output isn't a bad answer, it's a fluent one. I've heard versions of the same story more than once: a researcher catches a fabricated quote, asks the model to correct it, and gets a second fabricated quote back. It only got caught because that researcher knew the dataset cold, which leaves a large blind spot on the studies where nobody does.
What one leader memorably called the “magic wand dump” involves uploading the whole survey with no instructions, no framing, and seeing what comes out. While it feels like using the tool, it's actually abdicating the interpretive layer and hoping for a clear and accurate result.
Another pattern isn't something teams do to themselves; it's stakeholders feeding a team’s deck into a model to generate their own executive summary. You lose control of the interpretation after delivery, which is a new problem that no amount of internal guardrails solves.
What makes these all so dangerous is that failure is undetectable by reading. A fabricated quote and a real one are formatted identically. A theme the model invented and a theme it found are the same number of words, in the same confident register, sitting in the same bullet. In every other part of research, bad work looks bad. You can hear a leading question, or you can see a thin sample. AI is the first tool we've had where output quality is invisible on the surface, and our entire professional instinct for catching problems is visual.
That's not an argument against the tool. It's an argument for knowing which layer it's working on.
The cost that never makes it into the business case
Three things worry me more than hallucinations, because all three are slow.
- The first is atrophy. A researcher told me they felt noticeably less sharp after two years of heavy AI use, and the specific thing they'd lost was synthesis intuition. That's a real cost and it doesn't show up in any efficiency metric. Synthesis is a muscle. It responds to disuse the way muscles do.
- The second is the junior pipeline. Senior researchers can spot a flawed synthesis because they spent years doing the grueling manual version. If someone entering the field never does that version, where does the instinct come from?
- Some organizations didn't just “add” AI to their research teams. Instead they cut the team and bought the tool. Here's the part that should sit uncomfortably with the people who made that trade: A team out of MIT's Project NANDA reported last year that 95% of enterprise GenAI deployments produced no measurable return, against more than $30 billion in spend. That 95% got picked apart, and fairly. Success was defined narrowly as P&L impact measured six months post-pilot, the sample was 153 leaders recruited at industry conferences, and plenty of real value never shows up that way. So I won't lean on the number. But the direction isn't seriously in dispute: Buying the tool and getting the value are different projects, and most companies have only done the first one. Cutting the judgment layer to afford the mechanical layer is how you end up in the majority.
So what's actually missing from the stack?
The same gaps come up again and again.There's no verification layer, as every research too right now is built to produce output. Almost nothing is built to check it. The entire burden of validation sits on a human reading carefully, which is exactly the thing that erodes when the work gets fast.
Provenance is missing, and trust in a synthesis should be verifiable in one click, not taken on faith. When a report tells me participants felt friction at checkout, I want to click that and land on the four responses it came from; not a citation, the actual raw data.Consistency across sessions is also unreliable. If you run the same analysis twice, you will get two different types of answers. Memory bleeds between sessions in ways that are hard to predict and harder to explain to a stakeholder who read the first version.
The last mile is barely tooled at all, as we automate study design and synthesis, but then everyone still opens Google Slides and hand-builds the deck. The gap between a finished analysis and a decision a stakeholder will actually make is where research impact is won or lost, and it's the least automated stage in the entire workflow.
And governance keeps arriving after adoption. Nearly every team I talk to bought the tools first and wrote the rules later, if at all. That's backwards, and it's the same mistake research teams made with democratization a decade ago.
The Sprig layer
I spent years assembling this research stack by hand: question banks in a wiki, templates in a folder, a review process held together with Slack and goodwill. Now I work at the company building it as software:
The Design Agent works the mechanical layer of study construction: logic, branching, neutral phrasing, bias checks. It's the part where a non-expert most reliably breaks a study, and it's the part a machine can hold a standard on better than a busy human can.
The Field Agent runs the study and adapts to each participant in real time instead of pushing everyone through the same static form. Mechanical work that used to be impossible rather than just slow.
The Synthesize Agent works the interpretive layer and the last mile at once, and it keeps the evidence attached. That's the design decision I care most about. A finding you can trace is a finding you can defend.
And Sprig MCP exists because your judgment layer doesn't live in our product. It lives in your head, in Claude, and in the doc where you're actually thinking. Pulling real study data into that environment beats asking you to do your thinking inside somebody's dashboard.
None of that removes you from the loop. The point of a good stack isn't that you stop checking. It's that you know exactly where checking matters, instead of spot-checking everything at random and calling it rigor.
Draw the line before you need it
The teams handling this well aren't the ones with the best tools, they're the ones who decided, out loud and in advance, which layer each tool is allowed to operate on and who owns the call when something drifts.
Write down your three layers. Name what's mechanical, name what's interpretive, and be ruthless about what stays judgment. Then go automate the chopping without apology.
The stack got fast this year, but what it didn't change is where you belong in it. That part is still your job.
Curious where Sprig’s research agents can sit in your stack? Book a demo and we'll map it with how your team works today.