Who Gives the Lower Scores? Student-Side Gender Bias in Course Evaluations
Most coverage of gender bias in student evaluations asks who gets penalised. The more actionable question is who does the penalising — and the evidence says the bias is concentrated, not uniform. That changes how you should read a department-wide average.
Koji for Education
Research & Editorial Team ·
Bottom line: The well-documented gender penalty in student evaluations of teaching (SET) is not produced evenly by the whole student body. The best quasi-experimental evidence shows it is driven disproportionately by male students, is larger in quantitative subjects, and falls hardest on junior women. This matters for interpretation: a raw departmental average silently weights every respondent equally, even though the bias lives in an identifiable slice of the respondent pool. If you cannot see the composition of who answered, you cannot know whether a low score reflects teaching or reflects who happened to show up.
The question most bias articles skip
There is now a large literature establishing that SET scores are biased against women. That part is not in dispute among methodologists. What gets far less attention — and what actually determines what a quality-assurance officer should do on a Monday morning — is the structure of that bias. Is it a small, uniform tax that every student applies to every woman? Or is it a large effect produced by a subset of respondents under specific conditions?
The distinction is not academic. If bias were uniform, you could in principle apply a flat correction and move on. If it is concentrated and conditional, then any single averaged number is unstable: it will swing with the gender composition of the cohort, the discipline, and the seniority of the instructor — none of which the teacher controls, and most of which your survey instrument never records.
What the strongest evidence actually shows
The cleanest study on this point is Mengel, Sauermann and Zölitz (2019), Journal of the European Economic Association. They exploited a setting at Maastricht University where students were assigned to instructors in a way that was effectively random, then analysed 19,952 evaluations. Female instructors received systematically lower teaching evaluations than male colleagues — even though the gender of the instructor had no effect on students' actual grades or on the hours they spent studying. In other words, the women were not teaching worse on any objective proxy; they were being rated worse.
The crucial detail is the structure. The bias was driven by male students' evaluations, was larger in mathematical courses, and was particularly pronounced for junior women (PhD students and early-career staff). The penalty was not a flat tax applied by everyone. It concentrated in a recognisable corner of the data.
This is consistent with Boring (2017), Journal of Public Economics, who analysed evaluations at a French university and found that male students in particular expressed a preference for male professors, rating them more highly on dimensions stereotypically coded as masculine — knowledge and class leadership — despite no evidence that students learned more from men. And it echoes the well-known experimental result from MacNell, Driscoll and Hunt (2015), Innovative Higher Education: in an online course where the same instructors were presented under either a male or a female name, the perceived-male identity received higher ratings — including on objective items such as promptness of grading, where the actual turnaround time was identical.
Three different designs — a natural experiment, an observational study, and a randomised identity manipulation — converge on the same uncomfortable picture. The penalty is real, and it is not evenly sourced.
Why this breaks the departmental average
Quality processes almost everywhere still run on a mean. A module gets a 4.6; another gets a 4.1; a threshold is drawn; a conversation is had. But a mean is only a fair summary if the thing it averages is measuring the same construct for every respondent. The rater-side evidence says it is not.
Consider two early-career women teaching equivalent quantitative modules. One cohort happens to be 70% male; the other happens to be 40% male. On the evidence above, the first instructor will tend to score lower for reasons that have nothing to do with her teaching and everything to do with who was sitting in the room. A panel comparing their averages is, in part, comparing the demographic accident of enrolment. The number looks objective. It is partly noise dressed as signal.
This is the deeper problem with treating an averaged Likert score as a measurement: it launders a structured, conditional bias into a single clean digit, and then strips away exactly the contextual information — respondent composition, subject, instructor seniority — that you would need to interpret it responsibly.
"But isn't this just an excuse to ignore bad scores?"
This is the strongest and most reasonable objection, and it deserves a direct answer. Sceptics — including many fair-minded academics — worry that "the evaluations are biased" becomes a universal escape hatch: any inconvenient result can be waved away as prejudice. That worry is legitimate, and the bias literature can absolutely be weaponised that way.
But the rater-structure evidence cuts the other way. Precisely because the bias is concentrated and conditional rather than uniform, it can be examined rather than assumed. You do not have to choose between "all low scores for women are bias" and "no low scores for women are bias." You can ask the specific, testable questions the evidence points to: Does the gap appear mainly among one group of raters? Does it shrink when you look at written justification rather than the global number? Does a woman's score on specific, behaviourally anchored items — "the instructor returned work within the stated time" — match the documented facts? When MacNell and colleagues asked exactly that kind of objective question, the bias showed up against the facts, which is how we know it is bias and not performance. The answer to weaponisation is not to ignore the evidence; it is to collect feedback granular enough to interrogate it.
A second fair objection: effect sizes vary across studies and not every replication finds a penalty. True. The honest framing is that bias in SET is well-evidenced but heterogeneous — it depends on field, level, and rater mix. That heterogeneity is the whole point. It is an argument for instruments that capture the conditions, not for a single context-free average that hides them.
What this implies for how you collect feedback
If the bias lives in the structure, then the remedy lives in the instrument. Three principles follow:
- Stop relying on a single global number. A one-item "overall, how good was this course" rating is the most bias-exposed quantity you can collect, because it invites exactly the stereotype-driven halo judgement the research describes.
- Ask about specific, observable things. Behaviourally anchored questions ("Was feedback returned within the stated timeframe?") can be checked against reality and are harder to answer from stereotype alone.
- Capture enough context to interpret. You cannot adjust for, or even notice, a rater-composition effect you never recorded.
This is where the conversational, AI-moderated approach Koji for Education takes is genuinely different from a static survey. Rather than collecting a lone Likert score, Koji runs a standardised AI-moderated interview that probes why a student holds a view, using structured question types — open-ended, scale, single- and multiple-choice, ranking, and yes/no — so a global impression can be unpacked into specific, checkable claims. Because the AI moderator applies the same calibrated, bias-aware prompting to every respondent, it removes the human-moderator inconsistency that contaminates focus groups, and its automatic thematic analysis surfaces what students are actually responding to rather than collapsing everything into an average. When a low rating is grounded in a concrete, repeated theme — "feedback came back late" — you can act on it; when it is grounded in nothing a student can articulate, that itself is signal. None of this eliminates bias — no instrument can — but it surfaces and helps mitigate it instead of averaging it into invisibility, and it does so with GDPR-compliant, EU-appropriate handling of the resulting data.
The same conversational interview engine underpins the main Koji platform for general user and customer research — the difference here is the education-specific framing, not a different method.
The reframe
"Student evaluations are biased against women" is true but incomplete. The version that changes practice is: the bias is produced disproportionately by some raters, in some subjects, against some instructors — and your averaged score is built to hide exactly that. Read your numbers with the composition of who answered in mind, collect feedback specific enough to interrogate, and you move from arguing about whether bias exists to actually managing it.
Ready to collect course feedback you can interrogate rather than just average? See how Koji for Education works.