Customer discovery: AI writes the interview guide, real humans give the findings

Business Management Marketing Intermediate 240 min Free. Campus Zoom or Meet produce transcripts at no cost; source-grounded tools are free tiers.

The situation

My entrepreneurship students always want to skip the interviews and go straight to the pitch, and AI has made that temptation worse — you can generate a beautiful customer persona in ninety seconds and never talk to anyone. So AI is allowed to help build the guide and find patterns across transcripts they actually collected, and nothing else. What I assess is whether they can tell a hypothesis from a finding.

Steps

  1. Write the hypothesis sheet before you touch AI

    Paper, or any doc editor. Deliberately no AI.

    One page: who the customer is, what problem they have, how acutely. This is the control document — you cannot detect confirmation bias later unless you recorded your priors first.

    What you only learn by doing it: Make them number the hypotheses. In the final memo they have to write “H3: disconfirmed” next to each. Unnumbered hypotheses quietly mutate into whatever the data turned out to say, and students genuinely do not notice themselves doing it.

  2. Generate an interview guide, then cut it down hard

    Any free tier

    Prompt for open-ended questions probing past behaviour rather than future intent — “tell me about the last time you…” not “would you pay for…” Then require students to delete at least half. AI over-produces and drifts toward leading, solution-shaped questions.

    What you only learn by doing it: The tell for a bad generated question is that it contains the product. Any question a student cannot ask without describing their idea first is a pitch, not a discovery question. This single rule improves guides more than any amount of prompt engineering.

  3. Run the interviews yourself, in pairs, with consent

    Otter.ai

    Students interview in pairs — one talks, one notes tone, hesitation, and what was not said. Consent to record is asked on the record. No AI participates in this step, which is the point: the messy, emotionally charged material is exactly what the model has never seen.

    What you only learn by doing it: Require the note-taker to log every moment the interviewee corrected the premise of a question. Those corrections are where the actual pivot lives, and they are the first thing lost when you only keep the transcript.

  4. Synthesise in a source-grounded tool, not an open chat window

    Gemini Notebook (formerly NotebookLM)

    After five to ten interviews, ask for recurring themes with the supporting passage quoted for each. A source-grounded tool cites back to the document by default, which makes verification mechanical rather than aspirational.

    What you only learn by doing it: Run this in small batches of three or four transcripts, not all at once. Hallucination drops measurably with shorter, focused prompts. Students who paste everything get themes that sound sophisticated and cite nothing.

  5. Verify every theme against the transcript by hand

    The raw transcripts and Ctrl-F

    For each theme, locate the actual quote, paste it with speaker initials and date, and count how many distinct interviewees support it. Themes with fewer than two independent sources get demoted to signal, not finding.

    What you only learn by doing it: Expect roughly one theme in five to evaporate — either the quote does not say what the summary claimed, or three supporting quotes turn out to be the same person. Tell students the evaporation rate in advance so they treat it as expected yield rather than personal failure.

  6. Write the memo against the hypothesis sheet

    Any doc editor; AI for line editing only, disclosed

    Structured as H1 through Hn, each marked confirmed, disconfirmed or unresolved with evidence, closing with what the AI predicted that the humans contradicted. Students submit initial prompts, AI output and final memo together.

    What you only learn by doing it: Grade the disconfirmations hardest. A memo where every hypothesis was confirmed is almost always a memo where the student steered the interviews — an all-green memo is evidence of a problem, not of insight.

Where this breaks down

AI will generate a fluent, demographically specific, completely fictional customer persona on request, and students cannot tell it apart from a researched one because it is written in the same register. This is the most common failure in the assignment.

Market sizing is the other reliable fabrication zone. Ask for a total addressable market and you get a confident number with a plausible-looking source that does not contain it.

Models flatten affect, so sarcasm, hedging and politeness get coded as agreement, which systematically inflates how much customers claim to want the thing.

If transcripts contain identifiable personal information, strip names before uploading anywhere, and check whether your campus IRB treats the project as research.

Provenance: documented practice. NSF I-Corps Great Lakes Hub publishes this exact sequence — AI for question generation, human interviews for validation — with the explicit caution that AI output must be validated through real interviews. The transcript-analysis method is documented by Child Trends, which published its actual AI-assisted coding workflow and its observed failure modes.