New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Evaluation bias9 min read

Racial and Ethnic Bias in Student Evaluations of Teaching: What the Evidence Shows

Instructors of colour tend to receive lower student-evaluation scores than white colleagues teaching identical content. Here is what the peer-reviewed evidence actually says, where the effects are strongest, and how to read SET data without penalising the faculty universities most want to retain.

Koji Education Team

Product ·

Bottom line up front: Across experimental and observational studies, instructors from racially and ethnically minoritised groups tend to receive systematically lower student-evaluation-of-teaching (SET) scores than white instructors delivering identical content. A 2022 review of more than 100 studies in the Journal of Academic Ethics concluded that "faculty of color, and other marginalized groups are subject to a disadvantage in SETs." Because those scores are averaged into a single number that feeds promotion, tenure and quality-assurance decisions, the bias is not abstract — it can quietly disadvantage exactly the staff European universities are working hardest to recruit and retain. The fix is not to stop listening to students. It is to change how feedback is collected and how it is read.

The evidence is broad, not anecdotal

The strongest single source is Kreitzer and Sweet-Cushman's 2022 review, Evaluating Student Evaluations of Teaching, published in the Journal of Academic Ethics. Drawing on a dataset of over 100 articles, the authors document both measurement bias (SET scores have low or no correlation with actual learning) and equity bias (women, faculty of colour and other marginalised groups are disadvantaged regardless of the methodology or data source used).

Experimental work isolates the effect from confounds. In Exploring Bias in Student Evaluations: Gender, Race, and Ethnicity (PS: Political Science & Politics, 2020), Chávez and Mitchell held the course constant — same content, same assignments, same schedule, same communications — and varied only the perceived identity of the instructor. Women and instructors of colour received lower ratings than white men despite delivering an identical course.

Large observational datasets point the same way. Reid's analysis of RateMyProfessors data in the Journal of Diversity in Higher Education (2010) found that Black and Asian instructors were rated lower than white peers on overall quality, helpfulness and clarity. More recent work outside the US — for example a 2023 study of an online teaching-evaluation platform published in Frontiers in Education — finds racial and gender patterns persist in other cultural contexts, which matters for European institutions that should not assume the literature is purely an American artefact.

Why it happens

Several well-documented cognitive mechanisms combine:

  • Stereotype expectations. Students arrive with implicit expectations about who "looks like" an authoritative expert. When an instructor does not match that template, ambiguous moments (a hard exam, a challenging reading) are more readily attributed to the instructor than to the material.
  • The halo effect. A single salient impression colours every other judgement. An instructor perceived warmly is rated higher on unrelated dimensions such as "knowledge of subject"; one perceived through a negative stereotype is marked down across the board.
  • Accent and name penalties. Instructors with non-native accents or names that signal a minoritised background face an additional, separable penalty — a finding that overlaps with the literature on accent bias in evaluations of non-native instructors.
  • Intersectionality. Effects compound. The review literature repeatedly finds that women of colour — and Black men in particular — fare worst, because race and gender biases stack rather than cancel.

The compounding problem: a biased reading of a weak signal

Here is the part that should worry any institutional-research office. Even setting bias aside, the numeric SET score is a weak proxy for teaching quality. Meta-analytic re-analysis by Uttl, White and Gonzalez (2017) found that once study size is accounted for, the correlation between SET ratings and actual student learning is essentially zero. So racial bias is not contaminating an otherwise pristine measure — it is adding a systematic distortion on top of a signal that already struggles to track what we care about. Averaging a contaminated, weak measure and then ranking instructors on the third decimal place is methodologically indefensible.

"But critics argue the effects are small and inconsistent"

This is the strongest counterargument, and it deserves an honest answer rather than a dismissal.

It is true that effect sizes vary across studies, that some well-designed studies find null results for particular comparisons, and that the magnitude of bias is conditional — it depends on discipline, class size, course difficulty and how prescriptive students' role expectations are. Kreitzer and Sweet-Cushman themselves note the effect of gender is conditional on other factors rather than uniform. A responsible reading is not "every minoritised instructor is always marked down by a fixed amount."

But three things follow even from the cautious reading. First, the direction of the bias is remarkably consistent across data types and countries — it is the magnitude, not the sign, that varies. Second, high-stakes personnel decisions are precisely where small, systematic biases do the most damage, because they tip marginal cases and accumulate over a career. Third, "the effect is sometimes small" is not a reason to keep using raw averages for ranking; it is a reason to stop pretending those averages are precise enough to rank people at all. The conservative interpretation still leads to the same operational conclusion: do not use a single biased number as if it were an objective measurement.

What to do about it

The evidence supports a small number of concrete changes:

  1. Stop ranking instructors on raw averages. Report distributions and uncertainty, not league tables. A 4.1 and a 4.3 in classes of 25 are statistically indistinguishable — see our piece on why averaging Likert scores misleads.
  2. Triangulate. Student feedback is one necessary source among several — peer observation, teaching portfolios, self-reflection. Never let it stand alone for promotion. See triangulation in teaching evaluation.
  3. Shift weight from global ratings to specific, behaviour-anchored evidence. "Was feedback returned in time to be useful?" invites far less stereotype contamination than "Rate the overall quality of this instructor."
  4. Read the qualitative comments — carefully. Open text carries the richest signal and the most overt bias (including abusive remarks). It needs structured analysis, not a quick scroll.

How Koji approaches it

Koji for Education does not claim to eliminate bias — no instrument can, because bias originates in the student, not the form. What an AI-native approach can do is reduce, surface and contextualise it:

  • Conversational, AI-moderated interviews probe beyond a global rating. When a student gives a low score, the moderator asks why and for what specifically — turning a vague, stereotype-prone number into concrete, actionable evidence that is harder to contaminate.
  • Standardised AI moderation removes the human-moderator inconsistency that itself introduces variance, and applies the same bias-aware prompting to every respondent and every instructor.
  • Automatic thematic analysis of open-text feedback makes it feasible to read every comment for substance, flag abusive or identity-focused remarks, and separate "the instructor was unclear" from "I did not expect someone like this to be my lecturer."
  • Distribution- and theme-level reporting at programme and institution level discourages the spurious decimal-point ranking that turns small biases into career outcomes.

Koji shares the same AI interview engine used for general user and customer research on the main Koji platform — the difference here is a method tuned for the equity and methodological demands of higher education.

If your institution wants student feedback that informs teaching development without quietly penalising your most diverse faculty, explore Koji for Education.

Why this matters specifically in Europe

Much of the primary literature is American, with a growing UK contribution, and it would be a mistake to import the effect sizes wholesale. But the structural conditions that produce the bias are, if anything, more pronounced across the European Higher Education Area. European faculties are increasingly international: mobility schemes, English-medium master's programmes and competitive recruitment mean a large share of lecturers teach in their second or third language, to multilingual, multinational cohorts. That is precisely the configuration in which accent and name penalties operate, and it intersects with the gender and ethnicity effects documented elsewhere. The 2023 study cited above, conducted in the United Arab Emirates, is a useful reminder that these patterns are not a peculiarly American phenomenon.

European institutions also carry specific legal and ethical obligations — equal-treatment directives and institutional equality duties — that make the uncritical use of biased SET averages in personnel decisions not merely poor measurement but a potential compliance exposure. The practical implication is not to assume American magnitudes apply, but to audit your own data: disaggregate evaluation outcomes by instructor characteristics, look for systematic gaps, and treat any you find as a signal about the instrument rather than the instructor. An institution that has never checked whether its own SET data shows these patterns cannot claim its promotion process is free of them.

The takeaway

The question is not whether students sometimes evaluate instructors of colour unfairly — the peer-reviewed evidence says they do, with a consistent direction if a variable magnitude. The question is whether universities will keep feeding those judgements, unexamined, into the decisions that shape academic careers. Better instruments and better reading habits are available now.