AI user research: what it does well, and where it breaks

Language models are genuinely good at the part of qualitative research that used to eat a week. They are also very good at producing a confident theme out of three comments. Here is where the line sits.

Updated 9 min readBy Karla Cruz, UX researcher
In short
AI does the mechanical half of qualitative analysis — reading everything, clustering similar statements, tagging, drafting — at a speed no team can match by hand. It does not do the half that requires knowing how much weight a finding can carry. Keep the counts, the denominators and the quotes attached to every theme and the output is defensible; strip them out and you have a very fluent guess.

What "AI user research" actually refers to

The phrase covers two very different things, and conflating them is how teams end up disappointed.

The first is AI-assisted analysis: real evidence from real customers, analysed by a model. Transcripts, open-text survey answers, support tickets, sales call notes, churn reasons, app store reviews. The model reads them, groups them, and drafts findings. The evidence is human; the labour of reading it is not.

The second is synthetic users: asking a model to role-play a customer segment and answer your questions. This produces fluent, plausible, internally consistent text that contains no information about your customers whatsoever. It is a reasonable way to rehearse an interview guide. It is not research, and a finding sourced from it cannot be checked against anything.

Everything below is about the first kind.

What AI is genuinely good at

These are the tasks where a model does not merely save time but produces better results than a tired human doing the same job at 6pm on a Thursday.

  • Reading everything. Manual analysis silently samples. Nobody re-reads all forty transcripts before writing the summary; they re-read the memorable ones. A model reads the boring middle of every document with the same attention as the quotable opening, which is where the unglamorous, frequent problems hide.
  • Clustering paraphrases. "I couldn't work out which file format to use", "the importer kept rejecting my spreadsheet" and "I gave up at the upload step" are one theme wearing three costumes. Matching them is exactly what language models are built for, and it is the single most tedious task in thematic analysis.
  • Working across mixed sources. Interviews, tickets and reviews use different registers for the same complaint — a customer is polite in an interview, terse in a ticket and furious in a review. Normalising across registers by hand is slow and inconsistent; a model does it evenly.
  • Surfacing the thing nobody asked about. Human coders look for the themes the research question primed them for. A model asked for everything repeated will hand back the topic that was never on the discussion guide, which is frequently the most valuable output of a study.
  • Producing the first draft. Getting from a coded dataset to written findings is a multi-hour writing task. A draft that is 70% right and fully cited is a far better starting point than an empty document.

The five ways it goes wrong

Each of these produces output that reads well. That is what makes them dangerous: the failure is invisible in the prose and only shows up when someone checks the evidence.

1. Consensus manufactured from a tiny base

Ask for themes and you get themes, whether or not the data supports any. Three people mentioning pricing becomes "pricing is a significant concern for users". The model is not lying; it was asked for patterns and it found the best available ones. The fix is structural: never let a theme appear without the number of people it came from.

2. The lost denominator

This is the subtle one. If eight of forty participants used the export feature and six of those eight found it confusing, the finding is "75% of people who used export found it confusing" — not "15% of users". Both numbers are arithmetically true and they point at opposite decisions. A model summarising without being told which base each question applies to will quietly pick the bigger denominator, because it is the one in front of it.

3. Flattened disagreement

Summarisation pulls toward the mean. Two segments wanting opposite things — power users asking for more configuration, new users drowning in it — compress into a mushy statement about "balancing flexibility and simplicity" that no one can act on. The disagreement was the finding. Ask explicitly for contrasts by segment, not just for themes.

4. Stated preference read as behaviour

Customers are unreliable narrators of their own future behaviour, and a model has no way to know that "I would definitely pay for that" is the least predictive sentence in qualitative research. Analysis has to keep what people did separate from what they say they would do, and behavioural claims need behavioural evidence.

5. Lost provenance

A finding with no link back to a quote cannot be checked, cannot be defended in a roadmap review, and cannot be distinguished from a hallucination. This is the failure that makes all the others unfixable, because without provenance you have no way to detect that any of the first four happened.

Making AI analysis defensible

Six checks. If a tool or a prompt cannot produce these, treat the output as a hypothesis generator rather than as findings.

  1. Name the study type. Eight exploratory interviews and a 400-response survey support completely different claims. A report that does not say which it is invites the reader to assume the stronger one.
  2. Give every finding its own denominator. Count against the people who could have experienced the thing, not against everyone in the study.
  3. Cite the quotes. Every theme, linked to the verbatim evidence under it. This is non-negotiable and it is the check that catches the others.
  4. Separate observation from inference. "Six participants abandoned setup at the import step" is an observation. "Users don't understand our data model" is an inference. Label the second as one.
  5. Ask what is missing. A good analysis states what the evidence does not cover — which segments were not represented, which questions the data cannot answer. Silence about gaps reads as coverage.
  6. Spot-check by hand. Pick two findings and read their source documents yourself. If both hold, the rest probably do. If either does not, stop.

Where it fits, method by method

  • Customer interviews. Strong fit for analysis. No fit for collection. See how to analyse user interviews for the full pass.
  • Open-text survey responses. The best fit of all — high volume, short answers, one consistent question, a known denominator.
  • Support tickets. Strong fit, with a caveat: tickets over-represent problems severe enough to be worth complaining about, and say nothing about the people who left instead.
  • Usability testing. Partial fit. A model can code think-aloud transcripts but cannot see the hesitation, the cursor circling, or the participant's face.
  • Product analytics. Poor fit for a language model, and the wrong tool entirely — behavioural frequency questions belong in your analytics stack. Qualitative analysis explains the why behind what analytics measured.

What this looks like in practice

A team exports 340 support tickets from the last quarter and twelve onboarding interviews. The analysis returns a theme — setup abandonment at the data-import step — carrying 41 tickets and 7 of the 12 interviews, with the interview finding scoped to the nine participants who actually reached that step.

"I had the file ready, I just couldn't tell if it wanted the raw export or something I was supposed to reformat first. I closed it and meant to come back."

That is an observation with a base and a quote behind it. The inference — that the product's accepted formats are undiscoverable at the moment of need — is stated separately, and the opportunity it points at is a customer outcome rather than a feature: new customers need a way to know their file will be accepted before they commit to uploading it. Turning that chain into decisions is covered in from evidence to opportunity.

The wider practice of working across sources like these is covered in customer feedback analysis.

Common questions

What is AI user research?
AI user research is the use of language models to do the analytical part of qualitative research: reading interview transcripts, open-text survey answers, support tickets and reviews, grouping what they say into themes, and drafting findings. It does not mean AI-generated participants or synthetic answers — the evidence still has to come from real customers.
Can AI replace user interviews?
No. AI can analyse interviews far faster than a person can, but it cannot generate the interview. A model has no access to a customer's context, their workarounds or the thing they only mention at minute 34 because you asked a follow-up question. Synthetic-user tools produce plausible text, and plausible text is exactly what is dangerous when you cannot tell it apart from evidence.
Is AI analysis of qualitative data reliable?
It is reliable for the mechanical parts — clustering similar statements, tagging, finding repeated language across hundreds of documents — and unreliable for judgement calls it is not given the base rates to make. The failure mode is confident consensus built from a handful of comments. It becomes reliable when every theme carries its count, its denominator and its supporting quotes, so a reader can check the claim instead of trusting it.
How many interviews do you need before AI analysis is worth it?
Around five to eight interviews is where manual analysis starts to hurt and AI starts to pay for itself. Below that you can hold the whole dataset in your head and should. Above about twenty, manual thematic analysis stops being realistic for most teams under deadline, and that is where the time saving becomes large.
Does AI user research raise privacy problems?
It can. Interview transcripts and support tickets routinely contain names, employers, account details and things said in confidence. Before uploading anything, check what your tool retains, whether it trains on your data, and whether your participants consented to processing by a third party. Strip direct identifiers when the analysis does not need them.
What should an AI research report contain to be trustworthy?
The study type and what it can support, the sample size and how participants were selected, each finding's count against its own denominator, the verbatim quotes behind each finding, an explicit separation of what customers said from what the analyst inferred, and a statement of what the evidence does not cover.

Run this analysis on your own evidence

Upload interviews, tickets, survey exports or reviews and humsait reads all of it, groups the patterns, and writes the report with every finding linked to the quotes behind it.

Try it free

Keep reading