Should AI Summarise Your Student Feedback? The Automation-Bias Problem Nobody Mentions
AI can summarise a thousand course-evaluation comments in seconds. The danger is not that the summary is wrong — it is that a committee will trust it more than it deserves. Meet automation bias.
Koji Education Team
Product · June 17, 2026
Short answer: When AI summarises open-text student feedback for a committee, the biggest risk is not a flawed summary but automation bias — the well-documented human tendency to over-trust output from an automated system and stop checking it against the underlying evidence. A summary that is 90% accurate can do more harm than a messy spreadsheet if it makes readers stop reading the actual comments. The fix is not to ban AI summarisation — that would throw away a genuine advance in handling qualitative feedback at scale — but to design the workflow so the human stays in the loop and every claim remains traceable to a real student voice.
The promise, and the quiet hazard
For decades, the open-text box was the part of the course evaluation everyone ignored. Quantitative scores got tabulated; the comments — often the richest signal — sat unread because nobody had time to read 800 of them. AI changes that economics overnight: large language models can cluster hundreds of comments into themes, summarise sentiment, and draft a paragraph a busy committee can absorb in a minute. We have argued elsewhere that AI can reliably help analyse open-text feedback at scale and that a sentiment score is not insight. This is a real gain.
But the moment a machine hands a person a tidy conclusion, a different problem starts. Automation bias is the human propensity to favour suggestions from an automated system over our own judgement — and, crucially, to stop looking for disconfirming evidence once the system has spoken. The term was crystallised by Skitka, Mosier and Burdick in 1999, who showed that people monitoring automated aids made new kinds of errors: omission errors (missing a problem the automation did not flag) and commission errors (following the automation even when other information contradicted it).
This is measured, not hypothetical
Automation bias is one of the better-replicated findings in human-factors research. In healthcare, a systematic review by Goddard, Roudsari and Wyatt (2012, Journal of the American Medical Informatics Association) documented clinicians changing correct decisions to incorrect ones after seeing wrong advice from clinical decision-support systems. In public administration, Alon-Barkat and Busuioc (2023, Journal of Public Administration Research and Theory) found officials exhibited both automation bias and "selective adherence" — over-relying on algorithmic advice, and especially when it matched their prior expectations. A broad theme across the literature is that people supervising automation lose situational awareness and fail to intervene to correct errors, with the propensity rooted partly in "cognitive laziness" — a reluctance to do the demanding mental work of checking.
Now place a teaching-and-learning committee in that frame. It is the end of term. There are forty modules to review. An AI summary says: "Students broadly valued the lectures; the main concern was assessment workload." That sentence will, for most readers under time pressure, become the feedback. The two students who wrote something serious about feeling unsafe asking questions, or about an inaccessible reading list, may never be read — an omission error at institutional scale. And if the summary confirms what the committee already believed, selective adherence makes it doubly sticky.
But isn't a human committee just as biased?
Yes — and this is the counterargument that keeps the discussion honest. Three points:
-
Humans were never neutral readers. Before AI, committees skimmed, over-weighted the most vividly negative comments, and ignored the long tail entirely. "AI introduces bias" is only half the story; the manual baseline had negativity bias, selective reading, and no consistency at all. The honest comparison is AI-with-good-process versus messy-human-process, not AI versus a mythical perfect reader.
-
The failure mode is different, and that matters. Manual bias is distributed and visible-ish; automation bias is concentrated and invisible — one authoritative-sounding paragraph that nobody second-guesses. A concentrated, confident error is more dangerous in a high-stakes decision than diffuse human sloppiness, which is precisely why it deserves explicit design attention rather than a shrug.
-
The fix is process, not abstinence. Refusing AI summarisation does not return you to careful reading; it returns you to unread comment boxes. The right move is to keep the speed and defend against over-trust. That is a solvable design problem.
This is also distinct from the worry we address in does AI-moderated evaluation introduce algorithmic bias? — that piece is about bias in collecting feedback; this one is about bias in reading it. Both have to be designed for.
How to keep the human in the loop
- Make every claim traceable to quotes. A theme should never appear without the ability to expand it into the actual comments behind it. Traceability is the single most effective antidote to automation bias, because it lowers the cost of checking — directly countering the "cognitive laziness" mechanism.
- Surface dissent, not just the majority. A good summary names the minority and outlier signals explicitly, so the two serious comments are not averaged away.
- Show uncertainty and volume. "3 of 210 respondents raised this" is decision-useful; an unqualified theme is not.
- Keep a human decision owner. AI drafts; a named person is accountable for the reading. Automation should inform the judgement, never be the judgement — the stakes-aware stance the literature recommends.
- Treat high-stakes uses with extra care. For promotion or programme-closure decisions, read the raw comments, full stop, and use the summary only as an index.
Where Koji stands
Koji is built on the premise that AI should expand what humans can attend to, not replace their attention. Its automatic thematic analysis clusters open-text feedback into themes, but every theme stays traceable to the real student quotes beneath it — so a committee can drill from "assessment workload" straight into the exact words students used, which is the practical defence against automation bias. Because the underlying data comes from AI-moderated conversational interviews rather than a one-shot box, the moderator has already probed why a student felt as they did, giving readers context rather than a decontextualised line. Quality scoring flags thin or low-information responses so they are not over-weighted, and programme- and institution-level reporting preserves the distribution and the dissent rather than collapsing everything into one confident sentence. The same human-in-the-loop philosophy runs through the main Koji research platform. Koji's claim is deliberately modest: it makes large volumes of feedback readable and checkable, so humans make better-informed decisions — it does not make the decision for you, and it keeps everything GDPR/AVG-compliant.
If your committees are starting to read AI summaries of student feedback, the question to ask a vendor is simple: can I click from the summary back to what students actually said? If the answer is no, you are buying automation bias. See how Koji for Education keeps every theme traceable to a real voice.
A procurement checklist for AI feedback tools
If you are evaluating any tool that summarises student feedback, four questions separate a genuine human-in-the-loop design from automation bias sold as efficiency. One: Can I click from any theme or summary line back to the verbatim comments behind it? If not, the tool is asking for blind trust. Two: Does it report how many respondents a theme represents, and does it surface minority and outlier signals rather than only the majority? A summary that hides the two serious comments is hiding the comments that matter most. Three: Is there a named human decision owner in the workflow, or does the tool position itself as the conclusion? Four: For high-stakes uses — promotion, programme review — does the process require reading the raw responses, with the summary used only as an index? A vendor who cannot answer these is selling speed at the price of judgement, and the human-factors literature says that trade goes badly under time pressure.
The bottom line
AI summarisation of student feedback is a genuine advance — and a genuine hazard, because humans over-trust confident machines. The risk is not the occasional wrong summary; it is the committee that stops reading. Design for traceability, surface the dissent, and keep a human accountable, and you get the speed without surrendering the judgement.