Gender Bias in Student Evaluations of Teaching: Evidence and Mitigation
A research-grounded synthesis of gender bias in student evaluations of teaching (SET), anchored to MacNell, Driscoll & Hunt (2015) and corroborated by Boring (2017) and Mengel, Sauermann & Zölitz (2019), with practical mitigations for European quality-assurance teams.
Koji Education Team
Product
Answer in brief
Gender bias in student evaluations of teaching is a robust, replicated finding: in controlled and quasi-experimental designs, instructors perceived as female receive systematically lower ratings than instructors perceived as male, even when teaching is held constant. The effect is small-to-moderate on average (typically 0.1–0.3 standard deviations of the rating) but large enough to affect tenure, promotion, and contract-renewal decisions when SET scores are used as primary evidence. Quality-assurance offices should treat raw end-of-term ratings as one weak signal among several, not as a comparable measure across instructors, and should triangulate with qualitative and peer evidence.
What the research says
The anchor study for this brief is MacNell, Driscoll, and Hunt (2015), "What's in a Name: Exposing Gender Bias in Student Ratings of Teaching," published in Innovative Higher Education (40(4), 291–303). Two assistant instructors co-taught an online course and each assumed a male and a female identity for different student groups, while delivering identical content. Students rated the instructor presented as male significantly higher than the instructor presented as female on overall effectiveness, professionalism, promptness, fairness, respectfulness, enthusiasm, giving praise, and being a good role model — even though the underlying instructor was the same person. The differences were on the order of 0.5 points on a 5-point scale for several items. Because the design held the instructor constant and varied only the perceived gender, the result is causal evidence that students attach gender-based expectations to their ratings.
The MacNell study is sometimes criticized for its small sample (n ≈ 43), so it is important that the finding has been corroborated in much larger natural and quasi-experimental designs.
Boring (2017), "Gender biases in student evaluations of teaching," in the Journal of Public Economics (145, 27–41), analysed six years of mandatory first-year evaluations at a French university (Sciences Po) where students were assigned to teaching sections quasi-randomly. Male students rated male instructors significantly higher than female instructors, and students valued different teaching dimensions for men and women in ways that tracked gender stereotypes (men were rated as more knowledgeable and authoritative, women as warmer). Critically, exam performance — an objective learning outcome — did not differ by instructor gender, so the rating gap was not justified by differential teaching quality.
Mengel, Sauermann, and Zölitz (2019), "Gender Bias in Teaching Evaluations," in the Journal of the European Economic Association (17(2), 535–566), exploited a setting at Maastricht University where students were randomly allocated to instructors. Across roughly 19,952 evaluations, female instructors received systematically lower ratings than men, with bias driven by male students, larger in mathematical/quantitative courses, and especially pronounced for junior women. As with Boring, neither student grades nor self-reported study hours differed by instructor gender. A complementary meta-analytic synthesis is provided in Spooren, Brockx, and Mortelmans (2013), "On the Validity of Student Evaluation of Teaching: The State of the Art," Review of Educational Research (83(4), 598–642), which catalogs gender among the demonstrated biases affecting SET validity.
In short: a controlled experiment, two large quasi-experiments using random allocation, and a state-of-the-art review converge on the same conclusion. The bias is real, modest in average magnitude, and unevenly distributed (worse for junior women, worse in quantitative disciplines, worse from male raters).
Why it matters for course evaluation in practice
For a Dutch, Belgian, German, or Nordic quality-assurance office that uses SET scores in tenure, promotion, contract renewal, or module-redesign triggers, three operational implications follow:
- Cross-instructor comparisons are not apples-to-apples. A 0.2-point gap between a male and female instructor teaching the same module is consistent with bias rather than with a real teaching-quality difference. Ranking instructors against each other on raw SET means is therefore unsafe — particularly for high-stakes decisions.
- Thresholds and traffic-light dashboards amplify small biases. Esarey and Valdes (2020) show that even when SET is reliable and unbiased at the population level, using a percentile cutoff (e.g. "investigate everyone below the 20th percentile") produces high error rates for individuals, because the noise floor and any residual bias both compound at the tails. If a bias of ~0.2 SD systematically pushes women below thresholds, the false-positive rate for "underperforming" women is materially higher.
- The most affected groups are the least senior. Mengel et al. (2019) find junior women bear the largest bias. Because junior staff are also those most exposed to contract-renewal decisions, gender bias in SET can have a compounding effect on academic career progression, which has obvious implications for institutional equality plans and Athena Swan–style charters.
Under the European Standards and Guidelines (ESG), standard 1.5 requires institutions to assure themselves that teaching staff are qualified and that processes for evaluating them are fair. Continuing to rely on unadjusted SET means for high-stakes decisions is increasingly hard to defend against the ESG fairness requirement.
Limitations and honest caveats
A careful reader should note several caveats before concluding "SET is broken."
MacNell et al. (2015) used a small online sample, limiting statistical power and generalisability beyond online instruction. The instructor-identity manipulation is also unusual — most students meet their instructor in person, and visible cues (name, voice, appearance, dress) interact with bias in ways the online setting does not capture. Boring (2017) and Mengel et al. (2019) have much larger samples and stronger external validity, but they are single-institution studies (Sciences Po and Maastricht respectively); generalisation to, say, a Swedish technical university or a British post-1992 institution should be tentative. Effect sizes vary across studies and disciplines, and a small minority of studies find no significant gender effect, particularly in humanities. There is no consensus that bias is uniform across contexts — what is consensus is that, on average, bias exists and is non-trivial.
There is also a methodological objection worth surfacing: SET measures perceived teaching, not teaching. If students respond differently to identical behaviour by men and women, the rating gap may be a faithful measure of student experience even though it is unfair as a measure of teaching quality. This distinction matters for how the data is used (formative dialogue: fine; summative ranking: unsafe).
Finally, the literature on bias in SET goes well beyond gender. Race, ethnicity, accent, perceived attractiveness, course difficulty, time of day, and grading leniency are also documented. Gender is the best-evidenced case, but it is not the only one.
How Koji incorporates this
Koji for Education (edu.koji.so) is designed on the premise that a single end-of-term Likert mean is the wrong primitive for high-stakes evaluation. We mitigate gender bias (and related biases) through five concrete mechanisms.
AI-moderated conversational interviews. Instead of asking students to rate "overall effectiveness" on a 1–5 scale, Koji conducts an adaptive conversation that probes what specifically worked or didn't in the course. Open conversational evidence is harder to compress to a stereotype-driven scalar, and it surfaces concrete behaviours (clarity of explanation, responsiveness to questions, fairness of assessment) that can be triangulated against peer observation.
Structured question types alongside narrative. Koji supports open_ended, scale, single_choice, multiple_choice, ranking, and yes_no items. We recommend QA teams use scales sparingly for global judgments (where bias is largest) and lean on open_ended and ranking items to capture multidimensional feedback. Koji's evaluation-design guidance maps each item type to the kind of inference it can support.
Automatic thematic analysis with bias-aware coding. Koji clusters open-text responses into themes and surfaces the underlying behaviours students cite. Reports separate evaluative language ("she was warm/cold/professional") from behavioural language ("she returned grading within a week"), so that committees can weight the behavioural signal more heavily. This is designed to mitigate — not eliminate — the stereotype-loaded evaluative layer that drives the gender gap in raw means.
Triangulation and segment reporting. Koji reports cohort-level results with attention to who is responding and how. Segment comparisons (e.g. by year of study, prior grade, declared major) make it visible when, for instance, a rating gap is concentrated in a subgroup whose ratings are known from the literature to be more biased.
Formative-first defaults. Following the evidence on formative feedback (see Mid-Semester Feedback and the Power of Consultation), Koji defaults to mid-cycle conversational check-ins whose results are returned to the instructor — not pooled into a year-end ranking. This shifts SET toward its empirically supported use (helping a teacher improve) and away from its empirically problematic use (ranking teachers for high-stakes decisions).
None of this eliminates gender bias in student perceptions. The point is to reduce reliance on the single number where bias is largest, expose the underlying behaviours where evidence is strongest, and give committees a defensible basis for fair decisions under ESG 1.5.
For teams that also run user, learner, or product research outside the course-evaluation context, Koji's core research platform at koji.so applies the same AI-moderated interview engine to general qualitative research.
Related resources
- Do Student Evaluations Measure Learning? The Uttl Meta-Analysis Revisited
- Selection Bias in Course Evaluations: What Goos and Salomons Found
- Mid-Semester Feedback and the Power of Consultation
- Closing the Feedback Loop in Programme Quality Assurance
References
- MacNell, L., Driscoll, A., & Hunt, A. N. (2015). What's in a name: Exposing gender bias in student ratings of teaching. Innovative Higher Education, 40(4), 291–303. https://doi.org/10.1007/s10755-014-9313-4
- Boring, A. (2017). Gender biases in student evaluations of teaching. Journal of Public Economics, 145, 27–41. https://doi.org/10.1016/j.jpubeco.2016.11.006
- Mengel, F., Sauermann, J., & Zölitz, U. (2019). Gender bias in teaching evaluations. Journal of the European Economic Association, 17(2), 535–566. https://doi.org/10.1093/jeea/jvx057
- Spooren, P., Brockx, B., & Mortelmans, D. (2013). On the validity of student evaluation of teaching: The state of the art. Review of Educational Research, 83(4), 598–642. https://doi.org/10.3102/0034654313496870
- Esarey, J., & Valdes, N. (2020). Unbiased, reliable, and valid student evaluations can still be unfair. Assessment & Evaluation in Higher Education, 45(8), 1106–1120. https://doi.org/10.1080/02602938.2020.1724875
Related articles
Do Student Evaluations Measure Learning? The Uttl Meta-Analysis Revisited
A research-grounded reading of the Uttl, White & Gonzalez (2017) meta-analysis on the SET–learning relationship, with implications for European course evaluation and quality-assurance policy.
Mid-Semester Feedback and the Power of Consultation: What the Meta-Analyses Show
Penny & Coe (2004) and Cohen (1980) on the formative effectiveness of mid-semester student feedback when combined with consultation, and the implications for European course-evaluation design.
Selection Bias in Course Evaluations: What Goos and Salomons Found
A research-grounded reading of Goos & Salomons (2017) on selection bias in online course evaluations, with practical implications for response-rate policy and reporting at European universities.
Racial and Ethnic Bias in Student Evaluations of Teaching: What the Evidence Shows
Peer-reviewed evidence shows students rate instructors of colour systematically lower than White peers on the same teaching. We review Reid (2010), Chávez & Mitchell (2020) and Bavishi et al. (2010), the limitations, and how to design evaluation so race-linked bias does not contaminate quality decisions.