What Can Open-Text Student Comments Tell You That Likert Scores Cannot?
A research-grounded guide to open-ended student comments in course evaluation: what Alhija and Fresko found about who writes them and what they contain, how thematic analysis surfaces issues numbers miss (Stupans et al.), how rare abusive comments actually are (Tucker), and how to turn free text into reliable evidence.
Koji Education Team
Product
In brief: Open-text comments capture dimensions of the student experience that closed Likert items miss, but in traditional surveys only a minority of students write them, the comments skew general rather than specific, and they are laborious to analyse. The research (Alhija & Fresko, 2009; Stupans, McGuren & Babey, 2016) shows qualitative feedback adds genuine, otherwise-invisible signal — and that fears of widespread abuse are overstated (Tucker, 2014). The practical problem is not whether to collect open text but how to make it complete, specific, and systematically analysable.
The question, stated precisely
Every course-evaluation instrument that includes a free-text box is making an implicit bet: that what students write in their own words adds something the numbered scales do not. For a quality-assurance office, three empirical questions decide whether that bet pays off:
- Coverage — what proportion of students actually write anything, and is that group representative?
- Content — do comments repeat the closed items, or do they surface genuinely new dimensions?
- Tractability — can the free text be analysed reliably and at scale, or does it collapse into anecdote and cherry-picking?
The literature gives reasonably clear answers to all three, and they jointly explain why open text is simultaneously the most valued and the most under-used part of student evaluation.
What the research says
The anchor study is Alhija & Fresko (2009), "Student evaluation of instruction: What can be learned from students' written comments?" (Studies in Educational Evaluation, 35(1), 37–44). Examining written comments across 198 classes, they found that about 45% of students wrote any comment at all, that comments were more often positive than negative, and that they tended to be general rather than specific. Critically, the written comments addressed dimensions broadly similar to those covered by the closed-ended items — but also related to unique aspects of the courses that the structured items never asked about. In other words, open text both corroborates the numbers and extends beyond them, which is precisely the value proposition that justifies collecting it.
That extension is the centrepiece of Stupans, McGuren & Babey (2016), "Student Evaluation of Teaching: A Study Exploring Student Rating Instrument Free-form Text Comments" (Innovative Higher Education, 41, 33–42). Using the semantic-analysis tool Leximancer to mine cumulative free-text feedback for recurring concepts and themes, they identified issues that were not apparent from the quantitative data alone, producing a deeper, more actionable understanding of the curriculum. Their contribution is methodological: it demonstrates that systematic text mining can convert a pile of anecdotes into structured thematic evidence suitable for programme review — a substantial advance over reading a sample of comments and forming an impression.
The most common objection to open text — that anonymity invites abuse — is directly tested by Tucker (2014), "Student evaluation surveys: anonymous comments that offend or are unprofessional" (Higher Education, 68(3), 347–358). Across 30,684 comments from 17,855 surveys, Tucker found just 13 abusive comments (0.04%) and 46 unprofessional ones (0.15%). The clear conclusion: the overwhelming majority of students do not abuse anonymous feedback. Abusive comments are real and can cause disproportionate harm — and other work shows they fall hardest on women and marginalised academics — but they are statistically rare, and screening them is a tractable moderation problem rather than a reason to abandon qualitative feedback.
Read together, the three studies support a confident synthesis: open text adds unique, valid signal; it can be analysed systematically; and its most-feared failure mode is uncommon. The binding constraints are coverage and specificity, not abuse.
Why it matters for course evaluation in practice
For a QA process, the implications are concrete:
- Numbers tell you that; text tells you why. A 3.2 on "assessment was fair" is an alarm without a diagnosis. The comments are where the diagnosis lives — and Alhija and Fresko show they often point to dimensions the questionnaire never measured.
- Low coverage undermines the evidence. If only ~45% write comments, and writers differ systematically from non-writers (more motivated, more aggrieved, or more articulate), the qualitative record inherits a selection problem analogous to non-response bias in the numeric scores.
- Generality blunts action. "Great course" and "boring lectures" are not actionable. The value of open text rises sharply when comments are specific enough to change a syllabus, a rubric, or a teaching practice.
- Manual analysis does not scale and is not reproducible. Reading comments by hand invites confirmation bias and cherry-picking; systematic thematic analysis (as in Stupans et al.) is what makes free text defensible as accreditation evidence.
The goal, then, is an evaluation design that raises participation, pushes general reactions toward specific ones, screens the rare harmful comment, and analyses the rest reproducibly.
Limitations and honest caveats
A critical reader should hold several caveats:
- Coverage figures are context-bound. The ~45% writing rate from Alhija and Fresko reflects a particular instrument and setting; participation varies widely with survey length, mode, timing, and incentives, so it should not be treated as a universal constant.
- Thematic analysis embeds analyst choices. Tools like Leximancer surface co-occurring concepts, but theme definitions, granularity, and interpretation involve judgement. Inter-rater reliability and transparent coding protocols matter; an automated theme is not automatically an objective one.
- Sentiment is not validity. That comments skew positive (Alhija & Fresko) does not establish that the teaching was good — the same leniency and halo dynamics that affect numeric scores can colour written sentiment.
- Rare does not mean harmless. Tucker's 0.04% abuse rate is reassuring in aggregate but says nothing about the concentrated harm a single targeted comment can do to an individual instructor, which is why screening still matters.
These limits argue for treating open text as triangulating evidence to be analysed rigorously — not as unfiltered truth.
How Koji incorporates this
Koji for Education is designed around exactly the constraints this literature identifies — coverage, specificity, analysis, and screening:
- It turns the optional comment box into a guided conversation. Instead of a blank field that ~45% of students skip, Koji's AI-moderated interview invites every respondent into an adaptive dialogue, raising the share of students who contribute substantive qualitative feedback rather than leaving the text record to a self-selected minority.
- It converts general reactions into specific ones in real time. Alhija and Fresko's core limitation — comments that are too general to act on — is the exact problem conversational probing solves. When a student says "the lectures were confusing," Koji follows up ("which part, and what would have helped?"), eliciting the specificity a static form cannot.
- It performs automatic thematic analysis by design. What Stupans et al. achieved post hoc with Leximancer, Koji builds in: open-ended responses are clustered into themes and surfaced with representative quotes, so programme directors get structured, reproducible evidence instead of a folder of raw comments.
- It applies quality scoring and screening. Koji flags low-effort or potentially abusive responses, addressing the rare-but-real harm Tucker documented while preserving the 99.8% of feedback that is constructive.
- It triangulates text against structure. Because the same interview carries
scale,single_choice, andopen_endeditems, Koji can show where the qualitative themes confirm the numbers and where they reveal dimensions the closed items missed — operationalising Alhija and Fresko's central finding.
Koji is built to mitigate the weaknesses of open-text feedback — thin coverage, vagueness, unscalable analysis — not to claim qualitative data is bias-free. The same AI-moderated interview engine powers Koji's core research platform at koji.so, where turning open-ended customer responses into structured themes is the identical analytical task.
Related resources
- Does Grading Leniency Inflate Student Evaluations? — the numeric-bias counterpart to qualitative feedback
- Mid-Semester Feedback and the Power of Consultation
- Do Student Evaluations Measure Learning? The Uttl Meta-Analysis Revisited
- Gender Bias in Student Evaluations of Teaching: Evidence and Mitigation
- Response Styles and Likert Scales: Why Cross-Cultural Evaluation Needs More Than Numbers
Practical guidance for evaluation committees
To make open-text feedback genuinely useful, evaluation committees can act on each of the constraints the research identifies. To address coverage, move beyond an optional blank box: prompted, conversational collection raises the share of students who contribute substantive comments, so the qualitative record is not left to a self-selected minority of the most motivated or most aggrieved. To address specificity — Alhija and Fresko's finding that comments skew general — design follow-up prompts that push "the lectures were confusing" toward "which part, and what would have helped," because only specific comments change a syllabus or a rubric. To make analysis defensible, adopt systematic thematic methods rather than reading a convenient sample; as Stupans and colleagues showed, structured text mining surfaces issues invisible in the numbers and produces evidence robust enough for accreditation review. To handle the rare harmful comment Tucker documented, build a lightweight screening step that flags abuse and protects staff without discarding the 99.8% of constructive feedback. And always triangulate: report where the qualitative themes confirm the numeric scores and where they diverge, because the divergences are usually where the most actionable insight lives.
References
- Alhija, F. N., & Fresko, B. (2009). Student evaluation of instruction: What can be learned from students' written comments? Studies in Educational Evaluation, 35(1), 37–44. https://doi.org/10.1016/j.stueduc.2009.01.002
- Stupans, I., McGuren, T., & Babey, A. M. (2016). Student evaluation of teaching: A study exploring student rating instrument free-form text comments. Innovative Higher Education, 41, 33–42. https://doi.org/10.1007/s10755-015-9328-5
- Tucker, B. (2014). Student evaluation surveys: Anonymous comments that offend or are unprofessional. Higher Education, 68(3), 347–358. https://doi.org/10.1007/s10734-014-9716-2
Related articles
Response Styles and Likert Scales: Why Cross-Cultural Evaluation Needs More Than Numbers
Acquiescence and extreme response styles vary systematically by culture (Harzing, 2006; Baumgartner & Steenkamp, 2001), which means raw Likert averages are not directly comparable across nationalities in Europe's multinational classrooms. This article explains the evidence and how to evaluate fairly across diverse cohorts.
Gender Bias in Student Evaluations of Teaching: Evidence and Mitigation
A research-grounded synthesis of gender bias in student evaluations of teaching (SET), anchored to MacNell, Driscoll & Hunt (2015) and corroborated by Boring (2017) and Mengel, Sauermann & Zölitz (2019), with practical mitigations for European quality-assurance teams.
Do Student Evaluations Measure Learning? The Uttl Meta-Analysis Revisited
A research-grounded reading of the Uttl, White & Gonzalez (2017) meta-analysis on the SET–learning relationship, with implications for European course evaluation and quality-assurance policy.
Mid-Semester Feedback and the Power of Consultation: What the Meta-Analyses Show
Penny & Coe (2004) and Cohen (1980) on the formative effectiveness of mid-semester student feedback when combined with consultation, and the implications for European course-evaluation design.