Biased Course Evaluations Are Not Just Bad Data. For a Public University, They May Be a Legal Risk.
Most course-evaluation debates treat bias as a measurement problem. For a public university it is also a compliance one: the Public Sector Equality Duty requires due regard to eliminating discrimination before you act — and the randomised evidence that student evaluations are biased by gender is now strong enough that 'we never looked' is hard to defend.
Koji Education Team
Product · August 7, 2026
Bottom line: If your university is a public body, using teaching-evaluation scores you have reason to believe are biased to inform decisions about staff is not only a measurement problem — it engages a legal duty. In England, Scotland and Wales that duty is the Public Sector Equality Duty (PSED) under section 149 of the Equality Act 2010; across the EU, equivalent anti-discrimination obligations apply. The duty does not require you to prove your evaluations are perfect. It requires you to have due regard to eliminating discrimination and advancing equality of opportunity when you exercise a function — and the randomised evidence that student evaluations of teaching (SET) are biased by instructor gender is now strong enough that "we never looked at ours" is a weak position to argue from.
What the duty actually requires
Section 149 of the Equality Act 2010 requires public authorities — universities among them — to have due regard, when exercising their functions, to three aims: eliminating discrimination, harassment and victimisation; advancing equality of opportunity between people who share a protected characteristic and those who do not; and fostering good relations. Two features of the duty matter here.
First, it is a duty of process, not outcome. As guidance on the meaning of "due regard" makes clear, the law does not require you to achieve a particular result; it requires you to consider the equality implications properly, with rigour and an open mind. Second, the courts have consistently held that due regard must be exercised in substance and in advance — before and at the time a decision is taken, not reconstructed afterwards to justify a decision already made. You cannot have anticipatory regard to a risk you have deliberately refused to examine.
That combination is what makes course evaluation a live issue. The duty attaches to the university's functions — including its employment functions, where SET scores routinely feed probation, promotion, contract renewal and workload decisions.
The evidence the duty now has to reckon with
The reason this is not hypothetical is that the evidence base has moved from "students say they prefer" to causal, randomised findings.
- In a controlled experiment, MacNell, Driscoll and Hunt (2015) had online instructors present under a male or a female identity while teaching identical material; the same instructor was rated significantly higher when students believed the instructor was male.
- Boring (2017), using French data where students were effectively randomly assigned to instructors, found female instructors received lower ratings despite students performing no worse on anonymously graded final exams.
- Mengel, Sauermann and Zölitz (2019), in the Journal of the European Economic Association, analysed 19,952 evaluations in a setting with random allocation of students to instructors. Women received systematically lower evaluations than men even though neither students' grades nor their self-study hours differed by instructor gender. The bias was driven by male students and was larger in mathematical courses.
Random assignment is the crucial word. These are not studies where better teachers happened to be men; the instructor's gender was uncorrelated with everything else, and the ratings still diverged. That is a causal claim about the instrument, not a grumble about ratings.
Why "due regard" bites here
Put the duty and the evidence together. A public university that uses SET scores in personnel decisions is exercising an employment function using an instrument with published, randomised evidence of bias against a protected characteristic. Two distinct legal exposures follow. One is the risk of indirect discrimination under section 19 if a biased metric produces a disparate impact on women (or other groups) that the university cannot objectively justify. The other — and the point of this piece — is the anticipatory PSED itself: the duty required the university to have due regard to that risk when it designed and relied on the process, and having done nothing to examine its own data is exactly the failure the duty describes.
"We did not know" does not rescue the position, because the evidence is public and prominent. Neither does "the studies are American," because the honest response to uncertainty about transfer is to look at your own data, not to decline to look. The duty is satisfied by rigorous, documented consideration — including, where warranted, disaggregating your evaluation data by protected characteristic and assessing whether a pattern exists. It is breached by treating an unexamined instrument as neutral.
This is not only a UK question. The EU Charter of Fundamental Rights (Article 21) and the equality directives (2000/43/EC and 2000/78/EC) place analogous obligations on public bodies across member states; the specific statutory mechanism differs, but the principle — a public authority must have regard to equality when it acts — travels.
The counterargument, taken seriously
Critics will say, correctly, that the PSED is a procedural duty with a genuinely low bar: it does not ban SET, does not require any particular finding, and courts give public bodies latitude in how they weigh the aims. A university that considers the issue and decides to keep using SET, with reasons, has not obviously breached anything. That is fair. The claim here is not that using SET is unlawful — it is that using SET while refusing to examine it for bias is the specific posture the duty is designed to catch, and it is an avoidable one.
A second fair objection: bias findings do not transfer automatically from one system to another, and a raw gender gap in your own scores could reflect confounds rather than bias. True — which is why discharging the duty means a careful internal analysis, not a press release. But that concession cuts toward examining your data, not away from it. The instrument that is never audited is the one that cannot be defended.
Where Koji fits
Discharging the duty well requires three things a legacy survey rarely provides: the ability to disaggregate, the ability to separate uses, and an audit trail. Koji reports at a level that lets an institution examine patterns across cohorts and — where lawful and appropriate — protected characteristics, so that "having due regard" is backed by evidence rather than assertion. Its bias-aware, standardised AI moderation removes one well-documented source of inconsistency (the human moderator), and its design supports keeping formative improvement separate from summative personnel use, which is the single most important governance step for any institution worried about the role of SET in promotion and tenure. None of this eliminates bias — no instrument can — but it turns "we never looked" into a documented, defensible process. Teams that run wider staff and stakeholder research can use the same interview engine on the main Koji platform.
The duty does not ask you to have perfect data. It asks you to have looked. Most universities, on teaching evaluation, still have not.
Want to know whether your evaluation data would survive scrutiny? Koji for Education helps you audit, disaggregate, and act — so due regard is something you can evidence.
Frequently asked questions
Does the Public Sector Equality Duty actually apply to course evaluations? The duty applies to a public authority's functions. Where SET scores feed employment decisions — probation, promotion, renewal — the university is exercising a function to which section 149 attaches, so the duty applies to how it uses that data.
Is using student evaluations unlawful, then? No. The PSED is a duty of process, not outcome; it does not ban SET and does not require any particular finding. The exposure comes from using an instrument with public evidence of bias while refusing to examine your own data — that is the posture the duty is designed to catch.
How strong is the evidence that SET is biased by gender? Strong for gender specifically: multiple studies using experimental or random-assignment designs (MacNell 2015; Boring 2017; Mengel, Sauermann & Zölitz 2019) find women rated lower despite equal measured outcomes. Random assignment rules out the "better teachers were men" explanation.
We are not a UK institution — is this relevant? Yes. The specific statute is UK, but the EU Charter (Article 21) and the equality directives (2000/43/EC, 2000/78/EC) impose analogous obligations on public bodies across member states. The underlying principle — have regard to equality when you act — is common.
What is the minimum step to reduce the risk? Examine your own evaluation data for disparities, document that you did so, and separate formative use from summative personnel decisions. Auditing the instrument is what turns "we never looked" into a defensible, due-regard process.