Can a Large Language Model Code Your Open-Text Course Feedback? What the Agreement Studies Show
Peer-reviewed evidence on how closely GPT-4-class models match human coders when categorising open-ended student comments — agreement statistics, where they fail, and how to use them responsibly in course-evaluation quality assurance.
Koji Education Team
Product
In brief
Recent peer-reviewed studies find that GPT-4-class large language models (LLMs) can categorise open-ended student comments at roughly the level of a second human coder for broad themes (often above 90% agreement on coarse categories), but with markedly lower reliability on fine-grained, interpretive constructs (Cohen's kappa frequently in the 0.4–0.65 range). The honest reading of the evidence is that an LLM is a credible first-pass coder and triage assistant for course-evaluation free text, not a replacement for human judgement on anything consequential. Used with human oversight and transparent reliability reporting, it makes large comment volumes analysable; used as an unchecked oracle, it manufactures false confidence.
What the research says
The most directly relevant peer-reviewed study comes from learning analytics, not marketing. Liu and colleagues (2025), writing in the Journal of Learning Analytics, tested GPT-4 against human-coded educational data across three datasets — algebra-tutoring transcripts, game-based-learning observations, and programming-debugging behaviours — and compared zero-shot, few-shot, and few-shot-with-context prompting against embedding methods. Their central conclusion is sobering and useful: no single method consistently wins, and crucially, "GPT-4 has the most difficulty with the same constructs that human coders find more difficult to reach inter-rater reliability on." In other words, the model is not uniformly weak; it is weak exactly where the underlying construct is ambiguous — which is also where human coders disagree.
Studies outside education converge on a similar pattern. A 2025 analysis in AI & Society applying LLMs to thematic analysis in the charity sector reported strong agreement on sentiment-type judgements (Cohen's kappa around 0.91–0.95) but near-zero reliability (kappa ≤ 0.01) on a more interpretive construct — evidence of impact — that demanded contextual reasoning the model could not perform. Across this emerging literature, multi-label thematic categorisation with GPT-4-class models tends to land in the "substantial" agreement band (kappa roughly 0.61–0.65), while inductive, open-coding tasks fall to "moderate" (kappa around 0.57). For comparison, the conventional human-coder benchmark for "acceptable" agreement is a kappa of about 0.60–0.70, with 0.80+ considered strong.
This work sits on top of the classic methodological foundation of qualitative analysis. Braun and Clarke (2006) established thematic analysis as a structured, reflexive process — familiarisation, coding, theme development, review — in which the coder's interpretation is part of the instrument, not noise to be removed. That framing matters for evaluating LLMs: the question is not "does the model find the one true coding?" (there is no such thing) but "does the model produce coding decisions a competent human analyst would accept and could audit?"
Why it matters for course evaluation in practice
Open-text comments are the richest part of any course evaluation and the least used, because reading and coding them at scale is expensive. A mid-sized European faculty running end-of-semester evaluations across 400 course instances can easily generate 15,000–40,000 free-text comments per cycle. Human thematic coding of that volume is rarely funded, so the text is either skimmed impressionistically (inviting negativity bias and cherry-picking) or ignored in favour of the Likert averages.
An LLM changes the economics. At broad-theme resolution — "this comment is about assessment workload," "this is about lecturer availability," "this is about lab equipment" — the agreement evidence suggests a model can do useful triage across the entire corpus in minutes, surfacing the distribution of concerns a programme director actually needs. That is a genuine capability gain for closing the feedback loop.
But the same evidence sets the boundary. The constructs that matter most for high-stakes decisions — whether a comment constitutes a safeguarding concern, whether criticism is substantive or abusive, whether feedback reflects a teaching deficiency versus a workload-design problem — are precisely the interpretive, context-heavy judgements where LLM reliability collapses. Using an automated theme count to inform a promotion or non-renewal decision would repeat, in a new technology, the oldest error in the student-evaluation literature: treating a noisy proxy as a clean measure. See our companion analysis on interpreting and reporting student ratings responsibly for why use-case, not instrument, governs validity.
Limitations and honest caveats
A critical reader should hold several objections in view.
Agreement is not accuracy. Cohen's kappa measures agreement between two fallible coders, not correctness against a ground truth. A model and a human can agree on a wrong label. High kappa on sentiment is reassuring only to the extent the human benchmark is itself valid.
Prompt and dataset dependence. Liu et al. (2025) show that the "best" method shifts by construct; reported agreement figures are not portable. A kappa obtained on charity-sector interviews or algebra transcripts does not transfer to your faculty's German-language engineering feedback. Any institution adopting LLM coding must re-validate on its own labelled sample, ideally with a held-out human-double-coded set per cycle.
Non-determinism and drift. The same model with the same prompt can return different codings on different runs, and vendor model updates silently change behaviour. This threatens the reproducibility that accreditation evidence requires. Versioning, fixed prompts, and temperature controls are necessary but not sufficient mitigations.
Construct ambiguity is shared, not solved. The finding that LLMs fail where humans fail is double-edged. It means the model is not pathologically bad — but it also means automation does not rescue you from a poorly specified construct. If "teaching quality" is vaguely defined, neither a human nor a model will code it reliably.
Language and demographic coverage. Most published agreement studies are English-language. Performance on the multilingual, code-switched feedback typical of European universities is under-evidenced, and bias in how models interpret comments about instructors from different gender, accent, or ethnic backgrounds is a live risk that mirrors known racial and ethnic bias in student evaluations.
Privacy and governance. Sending student free-text to a third-party model is a data-protection decision, not just a methods decision, and must be reconciled with GDPR and institutional ethics review before any pilot.
A minimum validation protocol
If you intend to use an LLM on real course-evaluation text, the agreement literature implies a concrete, repeatable check rather than a one-off vendor demo. First, draw a random sample of 100–200 comments from the current cycle and have two trained human coders categorise them independently against a fixed codebook; compute their inter-rater reliability so you know the human ceiling for these constructs. Second, run the model on the same sample with a frozen, version-pinned prompt and compute its agreement with the agreed human codes, reporting Cohen's kappa per construct rather than a single headline number — because, as Liu et al. (2025) show, performance varies sharply by construct. Third, accept automated coding only for constructs where the model reaches the human ceiling, and route the rest to human coders. Fourth, repeat the check every cycle, because model updates can silently change behaviour. This protocol turns "the AI coded our feedback" from a claim into auditable evidence, which is exactly what an accreditation panel will expect to see.
How Koji incorporates this
Koji is designed around the evidence above rather than against it. Three mechanisms are relevant.
First, Koji does not rely on a single Likert number plus an unstructured comment box. Its AI-moderated conversational interview probes behind a rating in the moment — when a student rates assessment clarity low, the system asks a targeted follow-up — so the open text it analyses is already more specific and less ambiguous than a typical end-of-form free-write. Reducing construct ambiguity at collection time is the most effective lever for raising downstream coding reliability, precisely because the literature shows reliability fails on ambiguous constructs.
Second, Koji's automatic thematic analysis of open text is positioned as first-pass triage with human oversight, not as a verdict. Themes are surfaced with the underlying verbatim quotes attached, so a quality-assurance officer can audit why a comment was grouped — the auditability Braun and Clarke's reflexive framing demands — and override the machine. The platform is built to support, not bypass, the human coder.
Third, Koji separates broad-theme reporting (where automated agreement is strong) from any individual-level or high-stakes inference (where it is weak), and frames machine output as "designed to assist interpretation," never as an objective score. This is a deliberate guard against the misclassification risk documented for ranking instructors by evaluation scores.
The same conversational-interview engine underlies Koji's core research platform at koji.so, where teams apply it to product and customer research; the education product brings that engine, plus accreditation-aware reporting, to course evaluation specifically.
Related resources
- Text Analytics for Open-Ended Student Comments: What NLP Can and Cannot Tell You
- How Reliable Is Your Coding of Open-Text Feedback? Inter-Rater Reliability and Thematic Analysis
- When Students Use ChatGPT to Write Their Course Feedback: AI-Generated Open-Text and Data Integrity
- What Can Open-Text Student Comments Tell You That Likert Scores Cannot?
- Do Students' Written Comments Match Their Ratings? What Concordance Tells You
- Why One Cruel Comment Outweighs Twenty Kind Ones: Negativity Bias in Reading Course Evaluations
References
- Liu, X., Zambrano, A. F., Baker, R. S., Barany, A., Ocumpaugh, J., Zhang, J., Pankiewicz, M., Nasiar, N., & Wei, Z. (2025). Qualitative Coding with GPT-4: Where it Works Better. Journal of Learning Analytics, 12(1), 169–185. https://doi.org/10.18608/jla.2025.8575
- Braun, V., & Clarke, V. (2006). Using thematic analysis in psychology. Qualitative Research in Psychology, 3(2), 77–101. https://doi.org/10.1191/1478088706qp063oa
- Wen, C., Clough, P., Paton, R., & Middleton, R. (2025). Leveraging large language models for thematic analysis: a case study in the charity sector. AI & Society. https://doi.org/10.1007/s00146-025-02487-4
Related articles
Text Analytics for Open-Ended Student Comments: What NLP Can and Cannot Tell You
Natural-language processing can turn thousands of free-text course-evaluation comments into themes and sentiment at scale — but the research (Cunningham-Nelson 2019; Sunar & Khalid 2023) shows where automated analysis is reliable and where human judgement is still required.
How Reliable Is Your Coding of Open-Text Feedback? Inter-Rater Reliability and Thematic Analysis
When you turn thousands of free-text comments into themes and counts, how do you know the coding is trustworthy? O Connor and Joffe (2020) on intercoder reliability, Braun and Clarke on thematic analysis, and what rigorous qualitative QA looks like.
Do Students'' Written Comments Match Their Ratings? What Concordance Tells You
Open-text comments and Likert scores usually agree — but the gaps are where the insight lives. What Alhija & Fresko (2009) and Brockx et al. (2012) found about the consistency between qualitative and quantitative course-evaluation data, and how to read it.
Why One Cruel Comment Outweighs Twenty Kind Ones: Negativity Bias in Reading Course Evaluations
Instructors and committees fixate on the harshest open-text comment and discount the praise. Baumeister''s "bad is stronger than good" and the negativity-bias literature explain why, and how to read qualitative course feedback fairly.