Abusive Open-Text Comments in Course Evaluations: The Evidence and the Duty of Care
Lakeman et al. (2022) found that 91% of surveyed Australian academics received non-constructive anonymous comments — insults, threats, remarks on appearance. A research-grounded look at abusive student feedback, its effect on staff wellbeing, and how to moderate open text responsibly.
Koji Education Team
Product
In brief: Open-text comments are the most valuable part of a course evaluation — and the most dangerous. Lakeman and colleagues (2022) surveyed 791 Australian academics and found that more than 91% had received non-constructive anonymous comments, ranging from insults and remarks about appearance, attire and accent to allegations, blame and outright threats. These comments harm staff wellbeing and contribute to occupational stress, and the burden falls disproportionately on women and marginalised groups. The lesson is not to abandon open text — it is to moderate it. Institutions have a duty of care to filter, contextualise and act on qualitative feedback rather than passing raw, anonymous abuse straight to the people it targets.
The question, and why it matters
Quantitative scores tell you that students were dissatisfied; open-text comments tell you why. That diagnostic richness is exactly why good evaluation practice prizes qualitative feedback. But anonymity, low accountability and the emotional charge of end-of-term frustration combine to make the open-text box a channel for personal abuse as well as constructive critique. When an institution forwards unfiltered comments to academics, it can be — without intending to — delivering harassment under the banner of quality assurance. The question for any evaluation programme is therefore double-edged: how do we preserve the genuine value of open text while discharging a clear duty of care to staff?
What the research says
The anchor study is Richard Lakeman and colleagues' 2022 paper, "Appearance, insults, allegations, blame and threats: an analysis of anonymous non-constructive student evaluation of teaching in Australia" (Assessment & Evaluation in Higher Education, 47(8)). Surveying 791 Australian academics, the authors found that more than 91% reported receiving non-constructive comments about their teaching. A thematic analysis of the examples respondents shared identified five recurring themes:
- Allegations — unsubstantiated accusations;
- Insults — personal abuse unrelated to teaching;
- Comments about appearance, attire and accent — remarks on the body and identity of the teacher rather than the course;
- Projections and blame — attributing the student's own difficulties to the teacher;
- Threats and punishment — intimidating or retaliatory language.
The authors report that personally destructive, defamatory, abusive and hurtful comments were commonly described, and they argue these may have adverse consequences for the wellbeing of teaching staff and contribute to occupational stress. A 2023 follow-up by the same group, "Student evaluation of teaching: reactions of Australian academics to anonymous non-constructive student commentary" (Assessment & Evaluation in Higher Education), documents the emotional toll and the coping responses of staff who receive such feedback, reinforcing that the harm is not hypothetical.
The pattern is not limited to one country or one team. Troy Heffernan's 2022 synthesis, "Sexism, racism, prejudice, and bias" (Assessment & Evaluation in Higher Education, 47(1), 144–154), reviewing more than 130 publications, concludes that evaluations "include increasingly abusive comments which are mostly directed towards women and those from marginalised groups" — situating the Lakeman findings within a broader, well-documented equity problem. Earlier work by Tucker (2014), "Student evaluation surveys: anonymous comments that offend or are unprofessional" (Higher Education, 68), independently documented offensive and unprofessional comments in an Australian dataset, showing the phenomenon predates recent attention. The convergence across studies establishes that abusive open-text feedback is common, patterned and unequally distributed — not a handful of outliers.
Why it matters for course evaluation in practice
Two obligations sit in tension and must both be honoured. First, the methodological obligation: open text is indispensable for understanding why a course succeeded or failed, and stripping it out to avoid abuse would gut the instrument. Second, the duty of care: an employer that systematically exposes staff to anonymous abuse through an official process risks both real harm and, increasingly, legal and policy exposure under workplace health-and-safety frameworks.
Reconciling them requires moderation as a designed step, not an afterthought:
- Screen before disclosure. Comments that are abusive, defamatory, or focused on protected characteristics rather than the course should be filtered or flagged before raw text reaches the named individual.
- Separate the actionable from the abusive. A comment can be both critical and legitimate; the goal is to retain substantive criticism while removing personal attacks.
- Report patterns, not just raw quotes. Thematic summaries protect staff from the cumulative sting of reading dozens of individual barbs while still surfacing genuine issues.
- Monitor for equity. Because abuse disproportionately targets women and marginalised staff, evaluation systems should track whether qualitative tone differs systematically by instructor identity — a signal of bias, not teaching quality.
Limitations and honest caveats
A careful reader should note several constraints on the evidence.
- Self-selected, retrospective reporting. Lakeman et al. surveyed academics and asked them to recall and share examples. Those who experienced abuse may have been more motivated to respond, which can inflate prevalence estimates; the 91% figure describes the surveyed sample, not necessarily all academics.
- No denominator on comment volume. The studies establish that abusive comments are widely experienced, but not what fraction of all comments are abusive — most open text may still be constructive. Over-reading the prevalence risks unfairly stigmatising students as a group.
- Definitional latitude. "Non-constructive" spans a wide range, from genuinely threatening to merely blunt or negative. Robust criticism is not abuse, and moderation must not become a tool for suppressing legitimate dissatisfaction.
- National and disciplinary specificity. The anchor data are Australian; norms of expression, anonymity policies and grievance procedures differ across European systems, so prevalence and remedies should be validated locally.
These caveats argue for proportion — moderate the harmful tail without pathologising student feedback wholesale — rather than for inaction.
How Koji incorporates this
Koji treats responsible handling of open text as a first-class design problem, which maps directly to the duty-of-care gap this research exposes.
- Conversational framing that reduces abuse at source. Because Koji's AI-moderated interview engages students in a structured, purposeful dialogue rather than dropping them in front of an anonymous blank box, it nudges feedback toward the specific and constructive. A respondent asked targeted follow-ups about pacing, assessment or support is being steered away from the free-floating venting that the open-box format invites.
- Automatic thematic analysis and classification. Koji's text analysis is designed to cluster comments into themes and surface substantive issues, so a programme director can receive a pattern-level summary — the safer form of reporting the research recommends — rather than a raw, unmoderated dump of every individual remark.
- Flagging of non-constructive and identity-focused content. The platform can be configured to identify and separate comments that target appearance, identity or protected characteristics from those that address the course, supporting the screen-before-disclosure principle and the equity monitoring Heffernan's review implies.
- Quality scoring of responses helps distinguish substantive feedback from noise, so action planning rests on usable signal.
The framing is deliberately modest: Koji is designed to mitigate the exposure of staff to abusive feedback and to reduce its incidence through conversational structure — it does not eliminate abuse, and human judgement must remain in the loop for borderline cases and for any decision with employment consequences. Automated classification is an aid to responsible moderation, not a substitute for institutional policy and duty of care. Koji's core research platform at koji.so applies the same moderated, theme-aware approach to open-ended customer and product feedback, where unfiltered free text carries analogous risks.
What a defensible moderation policy looks like
Turning this evidence into governance requires a written policy rather than ad-hoc discretion. A defensible one names the threshold for intervention (comments that are defamatory, threatening, or focused on protected characteristics rather than the course), specifies who screens before disclosure, and guarantees that named staff receive a thematic summary rather than a raw feed when volumes are high or the tone is hostile. It should set out an appeal and support route for staff who receive abusive feedback, in line with the duty of care the Lakeman findings establish, and it should commit to monitoring whether qualitative tone differs systematically by instructor gender, ethnicity or background — the equity signal Heffernan's review flags. Crucially, the policy must protect legitimate criticism: moderation removes personal attacks, never substantive dissatisfaction. Publishing the policy to students also has a deterrent effect, signalling that the open-text box is a channel for course feedback, not anonymous abuse, and that comments are read by humans who act on them.
Related Resources
- What Can Open-Text Student Comments Tell You That Likert Scores Cannot?
- Text Analytics for Open-Ended Student Comments: What NLP Can and Cannot Tell You
- Gender Bias in Student Evaluations of Teaching
- Accent and Origin Bias in Course Evaluations
- Beauty Bias in Course Evaluations
- The Halo Effect in Course Evaluations
References
- Lakeman, R., Coutts, R., Hutchinson, M., Lee, M., Massey, D., Nasrawi, D., & Fielden, J. (2022). Appearance, insults, allegations, blame and threats: an analysis of anonymous non-constructive student evaluation of teaching in Australia. Assessment & Evaluation in Higher Education, 47(8), 1245–1258. https://doi.org/10.1080/02602938.2021.2012643
- Lakeman, R., et al. (2023). Student evaluation of teaching: reactions of Australian academics to anonymous non-constructive student commentary. Assessment & Evaluation in Higher Education. https://doi.org/10.1080/02602938.2023.2195598
- Heffernan, T. (2022). Sexism, racism, prejudice, and bias: a literature review and synthesis of research surrounding student evaluations of courses and teaching. Assessment & Evaluation in Higher Education, 47(1), 144–154. https://doi.org/10.1080/02602938.2021.1888075
- Tucker, B. (2014). Student evaluation surveys: anonymous comments that offend or are unprofessional. Higher Education, 68(3), 347–358. https://doi.org/10.1007/s10734-014-9716-2
Related articles
Text Analytics for Open-Ended Student Comments: What NLP Can and Cannot Tell You
Natural-language processing can turn thousands of free-text course-evaluation comments into themes and sentiment at scale — but the research (Cunningham-Nelson 2019; Sunar & Khalid 2023) shows where automated analysis is reliable and where human judgement is still required.
Gender Bias in Student Evaluations of Teaching: Evidence and Mitigation
A research-grounded synthesis of gender bias in student evaluations of teaching (SET), anchored to MacNell, Driscoll & Hunt (2015) and corroborated by Boring (2017) and Mengel, Sauermann & Zölitz (2019), with practical mitigations for European quality-assurance teams.
What Can Open-Text Student Comments Tell You That Likert Scores Cannot?
A research-grounded guide to open-ended student comments in course evaluation: what Alhija and Fresko found about who writes them and what they contain, how thematic analysis surfaces issues numbers miss (Stupans et al.), how rare abusive comments actually are (Tucker), and how to turn free text into reliable evidence.
Accent and Origin Bias in Course Evaluations: What Rubin (1992) Revealed
Donald Rubin's classic experiment showed students rated an identical recorded lecture as harder to understand when they believed the instructor was Asian — evidence that perceived accent and origin bias course evaluations. What it means for QA in multilingual European universities, and how to design evaluation that resists it.