New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Sector trends8 min read

When Course Feedback Turns Abusive: The Duty of Care Your Evaluation System Ignores

A documented minority of open-text course feedback is not criticism but personal, discriminatory, and abusive commentary — and it lands hardest on women, minority, and early-career staff. Forwarding it raw is a wellbeing and legal risk. The fix is not to abolish free text; it is to moderate it.

Koji Education Team

Product ·

Bottom line up front: A substantial, well-documented minority of open-text course-evaluation comments are not feedback at all — they are personal attacks, and they fall disproportionately on women, racialised, LGBTQ+, disabled, and early-career or casual staff. Most evaluation systems pass every comment through to the named academic unread and unfiltered. That is a duty-of-care failure hiding inside a quality-assurance process. The remedy is not to switch off the free-text box — qualitative feedback is the most useful thing students give you — but to put a moderation layer between the student and the staff member.

The problem quality offices do not measure

Course evaluation is meant to be an instrument of quality improvement. For a measurable share of academics, it is also a channel for abuse that no other workplace process would tolerate.

In the most direct study of the question, Lakeman and colleagues surveyed Australian academics and found that 59% reported having received comments containing abusive, offensive, or derogatory words or phrasing at some point in their careers (Lakeman et al., 2022, Higher Education). The comments they catalogued — later quoted in The Conversation — included instructions to "lose some weight" and epithets like "stupid old hag." These are not low ratings. They are comments about a person''s body, age, accent, gender, or race that happen to arrive through an official university system.

Crucially, the harm is not evenly distributed. Heffernan''s literature synthesis of sexism, racism, prejudice and bias in student evaluations concludes that derogatory and abusive commentary is disproportionately directed at women and marginalised academics, and that this makes evaluation a recurring source of stress and anxiety for exactly the staff a university most wants to retain (Heffernan, 2022, Assessment & Evaluation in Higher Education). The psychological effects are not transient: affected academics describe negative impacts lasting years, alongside concrete career consequences when hurtful comments sit permanently on a record used for promotion and contract renewal.

There is a name for this in the organisational-psychology literature: contrapower harassment — harassment directed at someone who nominally holds more institutional power (the lecturer) by someone with less (the student), enabled here by anonymity and an open text field.

This is a design failure, not a student-behaviour problem

It is tempting to file abusive comments under "some students behave badly." That framing lets the institution off the hook. The more useful diagnosis is that the system is designed in a way that predictably produces this output: take anonymity, add an unstructured free-text box, remove any moderation step, and deliver the result straight to a named individual whose promotion may depend on it. Anonymity is essential for honest feedback — but anonymity without moderation is also a disinhibition engine, a dynamic well established in the research on online commentary.

The institution owns that design. And in most European jurisdictions the institution also owns a legal obligation that the design cuts against. Employers carry a duty of care for the health, safety and welfare of staff, including psychological safety. A system the university itself operates, which routinely transmits discriminatory abuse to employees, is difficult to reconcile with that duty. Quality assurance and staff welfare are usually run by different offices; abusive feedback falls in the gap between them.

The counterargument: isn''t filtering just censorship?

The strongest objection deserves a direct answer. If you moderate comments, are you not suppressing criticism — sanitising the feedback academics need to hear, and hiding poor teaching behind a wellbeing rationale?

No — provided you are precise about the distinction. There is a clean line between harsh feedback about teaching and abuse aimed at a person. "The lectures were disorganised and the assessment brief was unclear" is critical, uncomfortable, and completely legitimate; it must reach the lecturer intact. "You are a fat, incompetent [slur]" contains zero pedagogical information; it teaches the recipient nothing except that the channel is unsafe. Moderation done well removes the second category while preserving — even amplifying — the first. The goal is not to soften criticism but to strip out the noise that carries no signal and does measurable harm.

The related worry is that moderation hides underlying bias patterns from the institution. The answer is that abusive comments should be filtered from the individual''s view but retained and analysed in aggregate. A pattern of appearance-based or racialised comments across a cohort is itself vital equity intelligence for the quality office — it just should not be delivered comment-by-comment to the target. Filtering for the individual and surfacing for the institution are not in tension; they are the same policy seen from two angles. (This mirrors the argument in our piece on disaggregated evaluation and the awarding gap.)

What a duty-of-care-aware evaluation system looks like

A responsible design has four properties:

  1. Moderation before the academic reads it. No member of staff should be the first human to see an abusive comment about themselves. A screening layer sits between submission and delivery.
  2. Separation of the actionable from the personal. The system distinguishes constructive criticism (delivered) from personal attack (withheld from the individual, flagged for the institution).
  3. Consistency. Screening applies the same standard to every comment for every academic, rather than depending on whether a sympathetic administrator happens to pre-read a particular cohort''s feedback.
  4. A route for serious disclosures. Some comments are neither feedback nor abuse but disclosures of harm — a safeguarding matter with its own duty, which we treat separately in our piece on safeguarding disclosures in free-text feedback.

Legacy survey tools — paper forms, or static online questionnaires from EvaSys, Qualtrics and similar — were never built for step 1. They collect free text and export it verbatim. The screening, if it happens at all, is a manual task nobody is funded to do.

Where Koji fits

This is a problem AI-moderated evaluation is unusually well placed to mitigate — and we choose that word deliberately; no system eliminates abuse, and any automated filter will miss edge cases and occasionally over-flag. Koji for Education helps in three concrete ways:

  • A consistent moderation layer. Because feedback is gathered through a standardised AI moderator rather than a raw text box, abusive and discriminatory content can be flagged consistently before it reaches a named academic — applying one standard across every cohort, without relying on a human administrator to pre-read thousands of comments (and absorb the abuse themselves in the process).
  • Separating actionable feedback from personal attack. Koji''s automatic thematic analysis is built to surface the substance of what students say. Constructive criticism is organised and delivered; content that is purely personal or discriminatory can be withheld from the individual view while still being retained for institutional equity monitoring.
  • A format that discourages drive-by abuse. A conversational interview that asks students to explain what would improve the course, and probes for specifics, changes the interaction from an anonymous one-line vent into a reflective exchange — which tends to raise the quality and lower the toxicity of what comes back.

Many of the same staff also run general user and customer research; the main Koji platform uses the same AI interview engine for that work, so the moderation and thematic-analysis capabilities carry across.

None of this makes the underlying bias disappear. It changes who has to absorb it, and whether the institution can see the pattern. For staff who currently brace themselves each semester, that is not a small change.

The bottom line

Open-text feedback is the richest, most useful part of course evaluation — and, unmoderated, the most dangerous. Abolishing it, as some of the affected researchers have proposed, throws away the signal to escape the noise. The better path is to keep the free text and finally build the moderation layer that a duty of care requires. Universities have accepted that responsibility everywhere else in the workplace. The evaluation system should be no exception.

See how Koji for Education handles open-text feedback with a bias-aware, moderated AI interview — explore Koji for Education.