Do Anti-Bias Statements Reduce Gender Bias in Student Evaluations of Teaching?
A randomised experiment (Peterson et al., 2019) found that a short anti-bias statement placed on the evaluation form significantly raised ratings of women instructors without changing ratings of men. We review the evidence, the caveats, and how to deploy warning-label language responsibly in course evaluation.
Koji Education Team
Product
In brief
A short, explicit anti-bias statement placed at the top of a course-evaluation form can measurably narrow the gender gap in ratings that women instructors typically face. In a randomised field experiment across four large classes, Peterson, Biederman, Andersen, Ditonto and Roe (2019) found that students who read a brief statement warning them that bias can affect evaluations rated female instructors significantly higher than students who saw the standard form; ratings of male instructors were unaffected. The intervention is cheap, fast, and easy to deploy — but it is a mitigation, not a cure. Effect sizes are modest, replication remains thin, and a warning label cannot repair an instrument that is noisy and only weakly related to learning.
What the research says
The anchor study is Peterson et al. (2019), Mitigating gender bias in student evaluations of teaching, published in PLOS ONE. The authors ran a randomised experiment in four large-enrolment courses at a US public university — two taught by men and two taught by women. Within each course, students were randomly assigned either the standard institutional evaluation instrument or the same instrument prefaced with language explicitly asking students to consider how gender and other biases can influence their judgements of teaching. Because assignment was random within each classroom, the design holds the instructor, the course content, and the term constant, isolating the effect of the language itself.
The result was asymmetric and directionally clean. Students in the anti-bias condition gave significantly higher ratings to female instructors than students in the control condition, while ratings of male instructors did not differ between conditions. In the authors'' framing, a small linguistic nudge partially closed a gap that the same students would otherwise have reproduced. The effect appeared on the overall/summary items where personnel-relevant gender penalties tend to concentrate.
This finding does not stand alone; it sits on top of a substantial causal literature establishing that the gap it targets is real. MacNell, Driscoll and Hunt (2015), in a now-classic online experiment, held the actual instructor constant and manipulated only the perceived gender of the instructor in an online class. Students rated the instructor they believed to be male more highly than the identical instructor they believed to be female, across dimensions including promptness and fairness — a difference that cannot be attributed to teaching behaviour because the behaviour was identical. Boring (2017), analysing a large natural experiment in a French university where students were quasi-randomly assigned to teaching-group tutors, found that male instructors received higher scores on dimensions associated with male stereotypes even when female instructors'' students performed as well or better on anonymous final exams. Mengel, Sauermann and Zölitz (2019), using random assignment of instructors to sections at a Dutch business school, found women received systematically lower evaluations — with the penalty largest for junior women and from male students — despite no corresponding deficit in student learning. Peterson et al. is best read as the intervention response to this diagnosis: if the gap is a bias in the rater rather than a deficit in the teaching, then priming raters to notice the bias should attenuate it, and in their experiment it did.
Kreitzer and Sweet-Cushman (2022), in a review of measurement and equity bias in student evaluations of teaching published in the Journal of Academic Ethics, situate anti-bias statements within a broader menu of reforms — reweighting, contextualised reporting, and reducing the stakes attached to raw means — and treat warning-label language as one low-cost, evidence-informed component rather than a stand-alone fix.
Why it matters for course evaluation in practice
For a quality-assurance office, the appeal of an anti-bias statement is its cost-to-benefit ratio. It requires no change to the item bank, no new software, and no reduction in comparability across a department, because every respondent to a given form sees the same preface. It targets exactly the moment where bias is cheapest to counteract — the point of judgement — rather than trying to statistically undo bias after the fact. For institutions bound by the European Standards and Guidelines (ESG), which require that student feedback be collected and acted upon fairly, a documented, evidence-based step to reduce a known equity distortion is a defensible piece of a fairness case in front of an accreditation panel.
But the practical lesson cuts deeper than "add a sentence." Peterson et al. shows that evaluation results are constructed by the instrument, not merely recorded by it. The same students, on the same teaching, produce different numbers depending on framing. That should make any institution cautious about treating a raw mean as an objective measurement of teaching quality — and especially cautious about ranking instructors, or making promotion and tenure decisions, on small differences in those means.
Limitations and honest caveats
A methodologically careful reader should hold several reservations.
Single institution, few classrooms. The anchor experiment was conducted in four courses at one university. Randomisation gives it strong internal validity, but external validity — whether the effect generalises to other disciplines, cultures, languages, and to the specific wording an institution chooses — is not established by a single site. The size and even the existence of the effect may depend on the exact statement, the local student culture, and baseline levels of bias.
Replication is still thin, and null results exist. Interventions that work in one randomised trial do not always replicate. Anti-bias and "warning-label" manipulations in social psychology have a mixed track record, and some framings can backfire or fade. The honest position in 2026 is that anti-bias statements are promising and low-risk, not proven at scale. Institutions should treat deployment as an evaluable pilot, not a settled solution.
It narrows one gap, not all of them. The experiment addressed gender. Evaluations also carry biases linked to race and ethnicity, accent and non-native speech, instructor age, attractiveness, course difficulty, and grading leniency. A gender-focused statement may do nothing for these, and a generic "please avoid bias" line may be too diffuse to move any of them. Mechanistically, we should not assume a single sentence neutralises a family of distinct cognitive biases.
The outcome is the rating, not learning. Raising women''s scores toward men''s corrects an inequity in the measure. It does not make the measure a better proxy for actual learning — the multisection-validity literature (e.g. Cohen, 1981, and subsequent re-analyses) shows the ratings-learning correlation is modest at best. A fairer biased instrument is still a weak instrument.
Social-desirability and demand effects. A statement that foregrounds bias may induce students to consciously adjust — which is the point — but could also produce over-correction or a temporary Hawthorne-style shift that decays as the language becomes routine wallpaper. Longitudinal evidence on durability is limited.
How Koji incorporates this
Koji for Education treats the Peterson finding as a design principle, not a bolt-on. Three mechanisms are relevant, and we describe them as designed to mitigate bias rather than eliminate it.
Bias-aware framing at the point of judgement. Because Koji controls the respondent experience end to end, a configurable, evidence-informed bias-awareness preface can be presented consistently to every student before evaluation begins — the exact intervention Peterson et al. tested, deployed uniformly so comparability is preserved and the framing is logged as part of the instrument''s audit trail.
Moving beyond the single biased number. Peterson et al. shows how much a global rating can move on framing alone. Koji''s core method — AI-moderated conversational interviews that probe why a student holds a view — is designed to surface the reasoning behind a judgement rather than only its magnitude. When a student rates a course low, a follow-up probe can distinguish "the assessment brief was genuinely unclear" from an affective reaction, giving committees behaviour-anchored evidence that is harder to reduce to a stereotype-driven number. Structured question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) let institutions triangulate a scale item against elaborated text.
Bias-aware reporting. Koji''s reporting is built to show distributions, dispersion, and confidence intervals rather than a single decontextualised mean, and to support contextualised comparison rather than naïve league tables. This aligns with Kreitzer and Sweet-Cushman''s recommendation to reduce the stakes and increase the context around raw scores. Automatic thematic analysis of open-text responses lets a reviewer read what students actually said, which is a partial check on the halo and stereotype effects that inflate or deflate a global number.
Koji''s core research platform at koji.so applies the same AI-moderated interview engine to product and customer research, where framing and question-order effects distort self-report in exactly analogous ways.
Related resources
- /docs/gender-bias-student-evaluations-teaching
- /docs/gender-bias-course-evaluations-natural-experiment-causal-evidence
- /docs/gendered-language-student-evaluation-comments-brilliant-caring
- /docs/racial-ethnic-bias-student-evaluations-teaching
- /docs/attractiveness-bias-student-evaluations-teaching
- /docs/student-feedback-esg-accreditation-evidence
References
- Peterson, D. A. M., Biederman, L. A., Andersen, D., Ditonto, T. M., & Roe, K. (2019). Mitigating gender bias in student evaluations of teaching. PLOS ONE, 14(5), e0216241. https://doi.org/10.1371/journal.pone.0216241
- MacNell, L., Driscoll, A., & Hunt, A. N. (2015). What''s in a name: Exposing gender bias in student ratings of teaching. Innovative Higher Education, 40, 291–303. https://doi.org/10.1007/s10755-014-9313-4
- Boring, A. (2017). Gender biases in student evaluations of teaching. Journal of Public Economics, 145, 27–41. https://doi.org/10.1016/j.jpubeco.2016.11.006
- Mengel, F., Sauermann, J., & Zölitz, U. (2019). Gender bias in teaching evaluations. Journal of the European Economic Association, 17(2), 535–566. https://doi.org/10.1093/jeea/jvx057
- Kreitzer, R. J., & Sweet-Cushman, J. (2022). Evaluating student evaluations of teaching: A review of measurement and equity bias in SETs and recommendations for ethical reform. Journal of Academic Ethics, 20, 73–84. https://doi.org/10.1007/s10805-021-09400-w
- Cohen, P. A. (1981). Student ratings of instruction and student achievement: A meta-analysis of multisection validity studies. Review of Educational Research, 51(3), 281–309. https://doi.org/10.3102/00346543051003281
Related articles
Turning Student Feedback into ESG / ENQA Accreditation Evidence
A buyer's guide mapping the ESG 2015 internal quality assurance standards to concrete, accreditation-ready evidence you can generate from student feedback — and how AI-moderated evaluation closes the loop.
Gendered Language in Student Comments: Why Men Are ''Brilliant'' and Women Are ''Caring''
Bias in course evaluations is not only in the numbers — it is in the words. What Mitchell & Martin (2018) and Storage et al. (2016) found about systematically different language applied to men and women, and why open-text analysis must account for it.
Gender Bias in Student Evaluations of Teaching: Evidence and Mitigation
A research-grounded synthesis of gender bias in student evaluations of teaching (SET), anchored to MacNell, Driscoll & Hunt (2015) and corroborated by Boring (2017) and Mengel, Sauermann & Zölitz (2019), with practical mitigations for European quality-assurance teams.
Beauty Bias in Course Evaluations: Are Attractive Instructors Rated as Better Teachers?
What the research says about physical-attractiveness bias in student evaluations of teaching — from Hamermesh & Parker (2005) to German and laboratory replications — and how to design evaluation so appearance does not masquerade as teaching quality.