How Strong Is the Evidence That Student Evaluations Are Biased? Reading the Research Honestly
The literature on bias in student evaluations of teaching is large, contested, and easy to cherry-pick. Here is an honest reading of how strong the evidence actually is — and what it means for using SET responsibly.
Koji for Education
Research & Editorial Team · June 10, 2026
Bottom line up front: The evidence that student evaluations of teaching (SET) carry systematic bias is real and, in several domains, robust — but it is not uniform, and the strongest claims circulating in faculty lounges outrun what the data support. The defensible position is narrower and more useful than either "SET are worthless" or "SET are fine": averaged Likert ratings are a noisy, sometimes-biased signal that should never be used alone for high-stakes decisions, but well-designed feedback remains valuable for improving teaching. This post reads the research honestly — including the studies that complicate the anti-SET narrative.
Why the meta-question matters
Most writing on SET bias argues a single claim: gender bias exists, or accent bias exists, or grading leniency inflates scores. Each of those deserves its own deep-dive (and we have written several). But quality-assurance officers and promotion committees face a different question: taken together, how much should we trust this number? That requires stepping back from any one effect and asking how strong, replicated, and generalisable the body of evidence is. Treating a contested literature as if it were settled — in either direction — is itself a methodological error.
What the evidence supports well
SET ratings are weakly related, at best, to actual learning. The most-cited challenge is the meta-analysis by Uttl, White and Gonzalez (2017, Studies in Educational Evaluation), which re-analysed decades of multisection studies — the design that best isolates teaching from confounds. Once prior ability and sample size were accounted for, they found SET ratings explained essentially no variance in student learning, concluding ratings and learning "are not related." Earlier moderate correlations, they argue, reflected small samples and publication bias. A subsequent multisection conflict-of-interest meta-analysis (2019, PMC6611447) reached compatible conclusions. This is one of the better-replicated findings in the field.
Experimental designs show identity effects. Observational studies confound instructor identity with everything else about a course. The value of MacNell, Driscoll and Hunt (2015, Innovative Higher Education) is that it was a true experiment: in an online course where students could not see their instructors, the same instructor received lower ratings when assigned a female name than a male one. Because instruction was identical, the rating gap is attributable to perceived gender rather than teaching. Boring, Ottoboni and Stark (2016, ScienceOpen Research) reinforced this with French administrative data on 23,001 evaluations, finding SET "more sensitive to students' gender bias and grade expectations than they are to teaching effectiveness," and that bias touched even ostensibly objective items such as promptness of grading.
Grade expectations inflate ratings. That students who expect higher grades return higher ratings is among the most consistently reproduced associations in the literature, and it creates an obvious incentive problem when ratings drive personnel decisions.
Where the evidence is weaker or contested
Honesty requires naming the soft spots. First, effect sizes vary widely. Gender-bias effects are sometimes large in experiments but small or null in some large observational datasets, and they interact with discipline, country, and course level — so "SET are biased against women" is true on average but not a reliable predictor for any single instructor's scores. Second, publication and narrative bias cut both ways. The same small-sample and selective-reporting problems Uttl and colleagues diagnose in the pro-SET literature also affect attention-grabbing bias findings. Third, defenders make a serious case. Benton and Cashin (2012, IDEA Paper No. 50) argue that multidimensional, well-constructed instruments show reasonable reliability and modest validity. Catherine Linse (2017, Studies in Educational Evaluation) cautions administrators that the practical problem is less the instrument than how crudely the numbers are read — comparing means to two decimal places, ignoring confidence intervals, and treating a 4.1 as meaningfully worse than a 4.3.
The most useful synthesis comes from Esarey and Valdes (2020, Assessment & Evaluation in Higher Education): even if a SET instrument were unbiased, reliable, and valid, the noise alone means that using it to rank instructors misidentifies the better teacher a substantial fraction of the time — their simulations put the error rate around 37% under realistic conditions. The deepest problem, in other words, is not only bias but low signal-to-noise — and that problem survives even the most generous reading of the pro-SET evidence.
But doesn't this just mean we should scrap student feedback?
This is the strongest counterargument to the anti-SET case being weaponised, and it deserves a direct answer. No — and conflating two different claims is where the debate usually goes wrong. "Averaged Likert ratings are a poor instrument for high-stakes ranking" is well supported. "Students have nothing useful to say about their learning experience" does not follow and is false. Students are uniquely positioned to report whether a course was clearly organised, whether feedback arrived in time to act on, whether they felt able to ask questions. The failure is not in asking students — it is in compressing what they say into a single biased number and then over-interpreting it. The remedy is better instrumentation and more cautious use, not silence.
What responsible practice looks like
The evidence converges on a few defensible principles. Never use SET as the sole or dominant input to tenure, promotion, or renewal. Triangulate with peer observation, teaching portfolios, and learning evidence. Report distributions and uncertainty, not naked means. Separate formative feedback (for the instructor's improvement) from summative judgement (for the institution). And invest in capturing the qualitative substance of student experience rather than its numeric shadow.
How Koji fits
If the core problem is that a single averaged number throws away the information that actually helps, the response is to collect richer evidence and analyse it rigorously — which is what Koji for Education is built to do. Instead of a static Likert form, Koji runs AI-moderated conversational interviews that probe why a student felt a course was disorganised or unsupportive, surfacing actionable specifics a 1–5 scale cannot. Its automatic thematic analysis turns hundreds of open-text responses into structured, auditable themes, so qualitative feedback becomes usable at programme scale rather than skimmed and discarded. Standardised, bias-aware AI moderation applies the same probing consistently to every respondent, removing the human-moderator inconsistency that plagues focus groups. And because Koji captures distributions and themes rather than a lone average, it supports exactly the triangulated, uncertainty-aware reporting the evidence demands. Koji is careful about its claims: it mitigates and surfaces bias and noise — it does not eliminate them, and no honest tool should promise otherwise. (Teams running general user and customer research use the same AI interview engine on the main Koji platform.)
A reading checklist for committees and QA officers
When an evaluation report lands on a promotion committee's desk, a few disciplined habits turn this literature into practice. Check whether the response rate makes the sample representative before reading any mean at all. Look at the distribution and the confidence interval, not the point estimate — a 4.1 and a 4.3 from classes of thirty are statistically indistinguishable. Compare an instructor only to themselves over time, or to a genuinely like-for-like benchmark, never across disciplines, class sizes, or modalities, each of which independently moves scores. Treat any single number as a prompt for further enquiry, not a verdict. And read open comments for substance, weighting specific, verifiable observations above global praise or complaint.
It is worth being equally explicit about what the literature does not show. It does not show that good teaching and good ratings are unrelated for an individual — only that the relationship is too weak and too noisy to rank people fairly. It does not show that every instructor is biased against to the same degree. And it does not license an instructor in dismissing feedback they simply find inconvenient. The honest position is uncomfortable for everyone at once: the same student ratings are real evidence about the learning experience and a poor basis for a verdict on a career.
The evidence on SET bias is strong enough to change how you use student ratings — and nuanced enough that pretending it is settled, in either direction, is its own kind of dishonesty. The institutions that win trust are the ones that read the research as it is.
See how Koji for Education turns student voice into evidence you can defend — explore Koji for Education.