New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Evaluation bias9 min read

Gender Bias in Student Evaluations of Teaching: What the Evidence Shows

Multiple experimental and quasi-experimental studies find that student evaluations of teaching rate women systematically lower than men for equivalent work — and the gap is unrelated to learning. Here is the evidence, the strongest counterargument, and what to do instead.

Koji for Education

Research & Editorial Team · June 1, 2026

The short answer: A well-replicated body of evidence shows that student evaluations of teaching (SET) systematically rate women lower than men for teaching of equal quality, and that these ratings are largely unrelated to how much students actually learn. The bias cannot be reliably corrected after the fact. The defensible response is not to silence student voice — it is to change how that voice is collected and interpreted, moving away from a single averaged score toward richer, bias-aware, evidence-led feedback.

Why this matters for European institutions

Across European higher education, SET data feeds into decisions that shape careers: contract renewal, promotion, teaching-award shortlists, and programme review evidence supplied to quality-assurance bodies. When the instrument carries a systematic gender skew, every downstream decision inherits it. For institutions committed to equality and to the Standards and Guidelines for Quality Assurance in the European Higher Education Area (ESG 2015), this is not a peripheral methodological quibble — it is a fairness and validity problem at the heart of how teaching is judged.

What the evidence actually shows

The strength of the gender-bias finding comes from research designs that hold teaching quality constant and vary only the perceived gender of the instructor.

The identity-swap experiment. In a frequently cited study, MacNell, Driscoll and Hunt (2015) ran an online course in which assistant instructors each operated under two gender identities. Students rated the instructor they believed was male significantly higher than the one they believed was female — regardless of the instructor's actual gender (MacNell et al., 2015, Innovative Higher Education). Because the teaching was effectively identical and only the perceived label changed, the rating gap can only be attributed to the label.

The random-assignment natural experiment. Mengel, Sauermann and Zölitz (2019) analysed 19,952 evaluations at Maastricht University, where students were randomly allocated to instructors. Women received systematically lower teaching evaluations than men even though students' grades and self-reported study hours were unaffected by instructor gender. The bias was driven primarily by male students, was larger in mathematical courses, and was especially pronounced for junior women (Mengel et al., 2019, Journal of the European Economic Association 17(2):535–566). The random allocation is what makes this powerful: there is no plausible selection story in which better-taught students happen to land with male instructors.

The "objective" items are not safe either. Boring, Ottoboni and Stark (2016) examined 23,001 evaluations of 379 instructors by 4,423 students across six mandatory first-year courses at a French university. They found the gender bias large and statistically significant, that it contaminated even ostensibly objective items such as how promptly assignments were returned, and — critically — that SET were more sensitive to students' gender bias and grade expectations than to teaching effectiveness (Boring, Ottoboni & Stark, 2016, ScienceOpen Research).

And the ratings barely track learning. The meta-analysis by Uttl, White and Gonzalez (2017) re-analysed decades of multisection studies and concluded that SET ratings explain at most around 1% of the variability in student learning — that is, they are essentially unrelated to how much students learn (Uttl et al., 2017, Studies in Educational Evaluation 54:22–42). If a score is biased and largely disconnected from learning, its use as a teaching-quality metric is hard to defend.

Beyond gender: the cluster of effects to watch

Gender is the best-documented bias, but it travels in company. The same literature points to accent and perceived ethnicity effects, attractiveness effects, and a robust grading-leniency / grade-expectation effect, whereby students who expect higher grades return higher ratings. Layered on top are the usual psychometric distortions — halo effects, central-tendency bias, and acquiescence. None of these are reasons to ignore students; they are reasons to stop treating a single end-of-term average as an objective measurement.

The strongest counterargument — taken seriously

"Doesn't this just prove SET are worthless, so we should scrap student feedback altogether?" No — and pretending otherwise would be its own error. Students are uniquely positioned to report on things they genuinely observe: whether sessions started on time, whether feedback arrived when promised, whether they felt able to ask questions, whether the workload was manageable. The problem is not that students have nothing valid to say. The problem is the instrument: a five-point Likert item averaged to one decimal place compresses rich, partly-biased, partly-valid signal into a single number that institutions then over-interpret.

A second fair objection: "Can't we statistically adjust for the bias?" Boring, Ottoboni and Stark argue you largely cannot, because the bias depends on too many interacting factors (course type, student composition, expectations) to model reliably. A correction factor that is itself uncertain does not restore validity; it launders the problem.

The honest conclusion is narrow and defensible: keep listening to students, stop reducing them to a contaminated average, and never use raw SET means as a primary input to high-stakes personnel decisions.

What a bias-aware approach looks like

If the average is the problem, the fix is to collect feedback that is harder to reduce to a biased number and richer in actionable, behaviour-specific signal:

  • Ask about specific, observable experiences, not global impressions of the person ("Was feedback returned within the stated time?" rather than "Rate the instructor overall").
  • Privilege qualitative depth. Open-text and conversational feedback surfaces the why behind a rating, making it possible to separate a legitimate concern from a halo or expectation effect.
  • Use consistent, neutral prompting. Human moderators introduce their own inconsistency; standardised, bias-aware prompting keeps the question framing identical for every instructor.
  • Triangulate. Combine student feedback with peer observation and learning evidence rather than letting one biased number stand alone.

Where Koji fits

Koji for Education is built for exactly this shift. Instead of a static Likert form averaged into a single score, Koji runs AI-moderated conversational interviews that probe beyond the number — asking students why, following up, and capturing the specifics that a 1–5 scale discards. Its six structured question types (open-ended, scale, single-choice, multiple-choice, ranking, yes/no) let you ask about observable behaviours rather than global impressions, and its automatic thematic analysis turns large volumes of open-text feedback into patterns you can act on.

Because the AI moderator applies the same bias-aware, standardised framing to every conversation, you remove the human-moderator inconsistency that creeps into interviews — and you gain programme- and institution-level reporting on a GDPR/AVG-compliant footing appropriate for European data handling. To be precise about the claim: this mitigates and surfaces bias and gives you far better evidence to interpret; it does not eliminate the underlying tendency of some students to judge instructors through a gendered lens. No tool can promise that. What it can do is stop that tendency from being silently baked into a single decisive number.

Teams that also run general user or customer research will recognise the engine: the same AI interview technology powers the main Koji platform.

What this means for promotion and review committees

The practical implications follow directly from the evidence. First, raw SET averages should never be the primary input to a high-stakes personnel decision. Use them, if at all, as one weak signal among several, and read them alongside peer observation, teaching portfolios, curriculum contributions and direct evidence of student learning. Second, calibrate your skepticism to the risk profile. Mengel and colleagues found the gender penalty largest for junior women and in mathematical subjects; a committee comparing a junior female mathematician's scores against a senior male colleague's is comparing two numbers contaminated to different degrees. Third, resist cross-instructor and cross-discipline league tables. A "below-faculty-average" flag can reflect who teaches what to whom far more than how well they teach. Finally, document the limitation in writing. Committees that record how they account for known SET bias protect both their staff and the integrity of their own decisions — and demonstrate exactly the reflective quality culture that European accreditation bodies look for. None of this requires discarding student feedback; it requires holding it to the standard of evidence its known properties can actually support.

The bottom line

The gender-bias finding in SET is not fragile or fringe; it survives identity-swap experiments, random-assignment natural experiments, and large multi-year datasets, and it sits alongside evidence that the ratings barely track learning. Treat raw SET averages as what they are — a noisy, biased signal — and rebuild your feedback around bias-aware, qualitative-rich, triangulated evidence. That is how you keep faith with students and with the staff being evaluated.