How Do We Know It Is Bias? Natural-Experiment Evidence on Gender in Course Evaluations
Random-assignment and quasi-experimental studies — Maastricht, Sciences Po, and controlled online courses — provide causal evidence that gender bias in course evaluations is real and not explained by differences in teaching effectiveness or student learning.
Koji Education Team
Product
In brief
The strongest answer to "are lower evaluations for women instructors just measuring worse teaching?" comes from studies that hold teaching effectiveness constant by design. When students are randomly assigned to instructors, or when the same instructor teaches under a male and a female identity, women receive lower evaluations even though students' grades and learning do not differ. The bias is causal, comes disproportionately from male students, and is largest in quantitative subjects — which means raw evaluation scores cannot be compared across instructors of different genders without adjustment.
What the research says
Observational studies have long shown women instructors receive lower ratings, but they cannot rule out that women simply teach harder courses or weaker cohorts. The natural-experiment literature closes that gap.
The landmark European study is Mengel, Sauermann, and Zölitz (2019) in the Journal of the European Economic Association. At Maastricht University, students are as good as randomly assigned to instructor-led tutorial sections within each course, and the institution centrally grades coursework — giving the researchers both an exogenous source of variation in instructor gender and an objective measure of learning. Across 19,952 evaluations, female instructors received systematically lower teaching evaluations than their male colleagues, while neither students' grades nor their self-reported study hours differed by instructor gender. The bias was driven by male students, who rated female instructors about 21% of a standard deviation lower than male instructors; female students showed a smaller bias of roughly 8% of a standard deviation. The effect was larger in mathematical courses and most pronounced for junior women — precisely the group for whom early evaluations feed contract and tenure decisions.
A second quasi-experimental anchor is Boring (2017) in the Journal of Public Economics, using six years of data from a French university (Sciences Po) where first-year students are assigned to mandatory courses in a way that breaks the confound between instructor gender and student selection. Boring found students reward male instructors on dimensions associated with stereotypically male traits (leadership, knowledge) and women on stereotypically female ones (organization, clarity), even though women are at least as effective when effectiveness is measured by students' performance on anonymously-graded final exams. The ratings track gender stereotypes, not learning.
The cleanest causal isolation comes from the controlled online experiment of MacNell, Driscoll, and Hunt (2015) in Innovative Higher Education. Two assistant instructors each taught an online course under both a male and a female identity, so the only thing that changed was the perceived gender. Students rated the perceived-male identity significantly higher, regardless of the instructor's actual gender. Boring, Ottoboni, and Stark (2016) reanalyzed both the MacNell data and 23,001 French evaluations with nonparametric permutation tests and concluded that SET are biased against women by a statistically significant and sometimes large amount — biasing even ostensibly objective items such as the promptness of returning assignments — and are more strongly associated with students' gender bias and grade expectations than with teaching effectiveness.
Why it matters for course evaluation in practice
The methodological force of these studies is that they remove the usual rebuttals. Because assignment is random or identity is experimentally swapped, "the women taught harder courses" and "their students were weaker" are ruled out. What remains is bias — a difference in the rating that is not matched by any difference in learning.
For a quality-assurance office the implications are concrete. Raw scores are not comparable across instructors of different genders: a 0.2-point gap may be entirely a gender effect, so ranking or thresholding on uncorrected means systematically disadvantages women. The bias concentrates where stakes are high — quantitative disciplines and junior staff — so naive use of evaluations can compound into unequal tenure and renewal outcomes. Some items are not the neutral facts they appear to be: if students misjudge a woman's promptness because of her gender, an "objective" item offers no refuge. And because the bias is partly carried on stereotyped dimensions, dimensional reporting can make the pattern visible (e.g., women rated lower on "leadership" but equal on "clarity") rather than hiding it inside one global number.
Limitations and honest caveats
Several caveats keep the conclusion honest. The natural experiments are institution-specific: Maastricht's problem-based tutorial system and Sciences Po's structure may not generalize to large lectures elsewhere, and the magnitudes (21% / 8% of an SD) are particular to those settings. The MacNell experiment had a small sample (four sections), so its precision is limited even though its design is clean — which is exactly why the Boring-Ottoboni-Stark reanalysis matters. Gender effects are also context-dependent: Fan et al. (2019), in a large Australian observational study of more than half a million surveys, found significant bias against women mainly in science and weak or absent effects in fields with higher female representation, alongside a cultural/language penalty that sometimes exceeded the gender effect. That heterogeneity means the right correction is local, not a single universal constant, and that intersecting factors (discipline, accent, representation) must be modelled together. None of this weakens the core causal claim; it bounds its size and generality.
How Koji incorporates this
Koji for Education is designed to make bias visible and to reduce the weight placed on the biased global number — without pretending an instrument can erase a bias that lives in students' heads.
- Dimensional, low-inference questioning. Koji supports structured
scale,single_choice, andrankingitems that separate stereotype-laden dimensions (leadership, authority) from behaviourally specific ones, so a quality office can detect the gendered pattern Boring documented instead of collapsing it into one contaminated mean. - Bias-aware reporting and benchmarking. Koji's reporting is designed to discourage raw cross-instructor ranking and to contextualize scores within comparable cohorts, mitigating the unfair comparison that uncorrected means produce — particularly for the junior women and quantitative-course instructors the research flags as most affected.
- Qualitative evidence that resists stereotype shortcuts. Koji's AI-moderated conversational interviews and automatic thematic analysis surface specific, evidence-anchored feedback ("I could follow the proofs because she worked them step by step") that is harder to reduce to a gendered global impression than a single Likert number.
- Triangulation by design. Consistent with the finding that ratings track stereotypes rather than learning, Koji positions student feedback as one source among peer review and learning evidence, never the sole basis for high-stakes decisions.
These are mitigations, framed honestly: making bias legible and de-emphasizing the global score reduces harm, but does not eliminate a bias demonstrated to operate even on objective-seeming items. Koji's core research platform at koji.so applies the same AI-moderated interview engine to customer and product research, where rater stereotypes and halo effects distort feedback in structurally similar ways.
Why the design matters more than the dataset
The reason these studies are decisive where decades of observational work was merely suggestive is the logic of identification. In an ordinary dataset, instructor gender is tangled with everything else: who teaches the required quantitative core, who gets the large first-year cohorts, who is assigned the unpopular early slots. Any of those could produce a gender gap in scores that has nothing to do with bias. Random assignment severs that tangle — when a coin flip, not a timetable, decides which instructor a student gets, the only thing systematically different between male and female instructors is their gender, so a residual gap in ratings is attributable to gender itself. The identity-swap design goes further still, changing perceived gender while literally holding the teaching fixed.
For an institution, this reframes the practical question. The issue is not whether to "believe" that bias exists — the causal evidence settles that — but how to keep an instrument that is demonstrably biased from driving decisions it cannot fairly support. That means resisting the intuitive move of ranking a department's instructors on a single comparable number, and instead reading scores within gender- and discipline-comparable reference groups, foregrounding behaviourally specific and qualitative evidence, and reserving the global figure for the modest role it can sustain. It also means watching the high-stakes margins the research flags — the junior woman teaching first-year statistics is exactly the case where an uncorrected mean does the most damage.
Related resources
- Gender bias in student evaluations of teaching
- Racial and ethnic bias in student evaluations
- Gendered language in student comments (brilliant vs caring)
- Value-added learning versus student evaluations (Carrell & West)
- Is student evaluation of teaching valid? (Spooren state-of-the-art)
- Measurement invariance and differential item functioning
References
- Mengel, F., Sauermann, J., & Zölitz, U. (2019). Gender bias in teaching evaluations. Journal of the European Economic Association, 17(2), 535–566. https://doi.org/10.1093/jeea/jvx057
- Boring, A. (2017). Gender biases in student evaluations of teaching. Journal of Public Economics, 145, 27–41. https://doi.org/10.1016/j.jpubeco.2016.11.006
- MacNell, L., Driscoll, A., & Hunt, A. N. (2015). What's in a name: exposing gender bias in student ratings of teaching. Innovative Higher Education, 40(4), 291–303. https://doi.org/10.1007/s10755-014-9313-4
- Boring, A., Ottoboni, K., & Stark, P. B. (2016). Student evaluations of teaching (mostly) do not measure teaching effectiveness. ScienceOpen Research. https://doi.org/10.14293/S2199-1006.1.SOR-EDU.AETBZC.v1
- Fan, Y., Shepherd, L. J., Slavich, E., Waters, D., Stone, M., Abel, R., & Johnston, E. L. (2019). Gender and cultural bias in student evaluations: why representation matters. PLOS ONE, 14(2), e0209749. https://doi.org/10.1371/journal.pone.0209749
Related articles
Measurement Invariance: Can You Compare Course-Evaluation Scores Across Groups at All?
Comparing average ratings across departments, languages or online vs paper assumes the questionnaire means the same thing to every group. Measurement invariance is the test that assumption usually fails — and almost nobody runs it.
Gendered Language in Student Comments: Why Men Are ''Brilliant'' and Women Are ''Caring''
Bias in course evaluations is not only in the numbers — it is in the words. What Mitchell & Martin (2018) and Storage et al. (2016) found about systematically different language applied to men and women, and why open-text analysis must account for it.
Racial and Ethnic Bias in Student Evaluations of Teaching: What the Evidence Shows
Peer-reviewed evidence shows students rate instructors of colour systematically lower than White peers on the same teaching. We review Reid (2010), Chávez & Mitchell (2020) and Bavishi et al. (2010), the limitations, and how to design evaluation so race-linked bias does not contaminate quality decisions.
Gender Bias in Student Evaluations of Teaching: Evidence and Mitigation
A research-grounded synthesis of gender bias in student evaluations of teaching (SET), anchored to MacNell, Driscoll & Hunt (2015) and corroborated by Boring (2017) and Mengel, Sauermann & Zölitz (2019), with practical mitigations for European quality-assurance teams.