New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Sector trends8 min read

If the AI Can''t Show You the Quote, Don''t Trust the Theme: Provenance in AI-Summarised Course Feedback

An AI-generated theme you cannot trace to verbatim student comments is a claim, not a finding. Abstractive summarisers hallucinate, and course evaluation is exactly the setting where an unverifiable summary is dangerous. The fix is provenance: every theme linked to the exact comments behind it.

Koji Education Team

Product · July 29, 2026

Bottom line up front: An AI-generated theme that you cannot trace back to the specific student comments behind it is a claim, not a finding. Language models that summarise open text are known to hallucinate — to assert things the source never said — and course evaluation, where those summaries flow into personnel files, programme reviews and league-table narratives, is precisely the setting where an unverifiable summary is most dangerous. The requirement that separates a trustworthy AI summary from a plausible fabrication is provenance: every theme must link to the verbatim comments that support it, so a human can check before acting.

The seduction of the tidy summary

A programme director facing 900 free-text comments and a Tuesday committee will take the clean list of "top five themes" every time. That is entirely rational — thematic triage is exactly what AI should do, and doing it by hand is infeasible at scale. The danger is not that AI summarises; it is that a summary looks equally authoritative whether it faithfully reflects the comments or quietly invents them. Fluency is not fidelity.

The faithfulness problem is measured, not hypothetical

This is not vendor fear-mongering. The foundational study is Maynez, Narayan, Bohnet and McDonald, On Faithfulness and Factuality in Abstractive Summarization (ACL 2020), which ran a large human evaluation of neural summarisers and found that hallucinated content appeared in a substantial majority of model-generated summaries — over 70% of single-sentence summaries in their analysis — and that the majority of these were extrinsic hallucinations: information not present in, and not inferable from, the source text at all. The paper also distinguishes intrinsic hallucination, where the model misassembles facts that were in the source into a false statement.

Modern large language models are meaningfully more faithful than the 2020 systems Maynez studied. But "more faithful" is not "faithful", and the failure has not been engineered away — it has become rarer and harder to spot, which for a decision that affects someone's career is arguably worse, not better.

What a hallucinated evaluation summary looks like

The abstract problem becomes concrete fast:

  • A theme nobody raised. The model generalises three complaints about one week's readings into "students found the course intellectually unchallenging" — a sentence with very different consequences.
  • A fabricated representative quote. Asked for an illustrative comment, the model produces a fluent, plausible quotation that no student actually wrote.
  • Invented frequency or severity. "Many students felt unsafe raising questions" when two did, or "widespread confusion about assessment" from a handful of remarks. Frequency and severity are where even faithful paraphrase distorts, because a summary flattens how many said what.
  • Mis-attributed sentiment, folding a sarcastic comment into the positive pile or vice versa.

None of these is exotic. All of them survive undetected if the reader never sees the underlying comments — and the whole point of the summary was that they would not.

Why this is worse in evaluation than in the newsroom

Three features make course evaluation an unusually unforgiving home for unverified summaries. The consequences are high: tenure and promotion cases, contract renewals, decisions to restructure or close a module. The base rate of serious problems is low, so a single fabricated serious theme — "a pattern of discriminatory remarks" — is both rare enough to alarm and damaging enough to act on before anyone checks. And the output is, to its reader, effectively unfalsifiable: a committee handed "top themes" without the comments behind them has no way to tell a grounded finding from a confident invention. A sentiment percentage has the same defect at a coarser grain — as we have argued, a sentiment score is not insight.

The fix is provenance, not abstinence

The answer is not to ban AI summarisation — the question of whether AI can reliably analyse open text has a qualified yes as its answer — nor merely to warn humans to stay vigilant, which is the automation-bias angle. Vigilance without verifiability is useless: you cannot check what you cannot see. The technical requirement is traceability:

  1. Every theme links to its evidence. A reviewer clicks a theme and sees the exact student comments assigned to it, with a count. No orphan claims.
  2. Verbatim quotes are marked as verbatim. Extractive, copy-exact quotations are distinguished from abstractive paraphrase, so a "representative comment" is a real one, not a generated one.
  3. Frequency is reported, not implied. "14 of 210 respondents" beats "many students", and it is checkable.
  4. A human spot-check is built into the workflow, not left to individual diligence — the reviewer confirms the grounding for any theme that will carry a consequence.

Provenance does not mean reading all 900 comments; that would defeat the purpose. It means that before you act on a theme, you read the six comments behind it. The AI does the triage; provenance makes the triage auditable.

"Modern models barely hallucinate — isn't this solved?"

The strongest objection is that faithfulness has improved so much that provenance is belt-and-braces. Two answers. First, improved is not solved, and the residual errors are now subtle enough to slip past a busy reader precisely because the surrounding text is fluent and correct. Second, and more fundamentally, the case for provenance does not rest on the hallucination rate at all. Even a perfectly faithful summariser compresses, and compression makes choices about frequency, emphasis and severity that a decision-maker must be able to inspect. Auditability is a property you want for consequential decisions regardless of how good the model is — the same reason we expect a citation in a systematic review even from a careful author. A second objection — "requiring quotes just re-creates the manual work" — misreads the workflow: you verify the handful of themes you are about to act on, not the entire corpus.

A test you can run in ten minutes

Before trusting any AI feedback tool, run one check. Take a single surfaced theme and ask the system to show you every comment assigned to it. Read them: does each comment actually express the theme, or has the model swept in tangential remarks to inflate a count? Now ask for the theme's "representative quote" and search the raw corpus for that exact string. If it is not there verbatim, the tool is generating quotations rather than retrieving them — a bright-line failure for evaluation use. Finally, check whether the stated frequency matches the number of comments you can actually see. A tool that passes all three is grounded; one that stumbles on any of them is narrating. This ten-minute exercise tells you more than any feature list, and it is exactly the scrutiny a defensible evaluation process should apply before a theme reaches a committee paper.

Where Koji fits

This is the design principle behind Koji's thematic analysis rather than an add-on. Every theme Koji surfaces is grounded in the specific interview passages that generated it: a reviewer opens a theme and sees the verbatim student quotes, with counts, that sit beneath it. Because Koji's feedback comes from AI-moderated conversational interviews, the provenance runs all the way down to a transcript you can read in context — not a detached sentence, but the exchange it came from. That is the difference between "the AI says students found assessment unfair" and "here are the eleven students who said so, in their own words, and what they were responding to." Koji does not ask you to trust the summary; it shows you the evidence and lets you trust the process. The same grounded, quote-linked engine powers koji.so for general user research, where acting on a fabricated theme is just as costly. If a tool gives you themes but cannot show you the quotes, it is not summarising your feedback — it is narrating over it.