New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods10 min read

Racial and Ethnic Bias in Student Evaluations of Teaching: What the Evidence Shows

Peer-reviewed evidence shows students rate instructors of colour systematically lower than White peers on the same teaching. We review Reid (2010), Chávez & Mitchell (2020) and Bavishi et al. (2010), the limitations, and how to design evaluation so race-linked bias does not contaminate quality decisions.

Koji Education Team

Product

Answer box

Multiple peer-reviewed studies find that instructors perceived as Black, Asian, Latino or otherwise racially minoritised receive systematically lower student ratings than White instructors teaching equivalent material — with effects appearing even before students have any teaching to judge. The bias is real but moderate in size, entangled with gender and discipline, and large enough that raw evaluation means should never be compared across instructors of different backgrounds as if they measured the same thing. The defensible response is not to discard student voice but to change how it is collected and interpreted: probe the reasons behind ratings, separate course-design feedback from person-directed judgement, and benchmark within comparable contexts rather than against a single institutional average.

What the research says

The most-cited large-scale evidence is Reid (2010), who analysed 3,717 instructors' ratings scraped from RateMyProfessors.com across the 25 highest-ranked US liberal-arts colleges (3,079 White; 142 Black; 238 Asian; 130 Latino; 128 other). Controlling for the volume of ratings, racially minoritised faculty — Black and Asian instructors especially — were rated significantly lower on overall quality, helpfulness and clarity, while being rated higher on "easiness." A two-stage cluster analysis showed the very best-rated instructors were disproportionately White, and the worst-rated disproportionately Black or Asian. Notably, Reid found no reliable gender main effect in this dataset, underscoring that race operates as a distinct axis of bias, not a proxy for gender.

Chávez & Mitchell (2020) provide a cleaner causal design. In a quasi-experiment, instructors teaching otherwise identical online courses recorded short welcome videos — the only cue to their perceived gender and race/ethnicity, with all course content, assessment and workload held constant. Even with the teaching itself controlled, women and minoritised instructors attracted more negative and fewer positive evaluative comments than White male instructors. Because the course was held fixed, the difference is hard to attribute to anything other than the instructor's perceived identity.

Bavishi, Madera & Hebl (2010) isolate the "judged before met" component. Students evaluated professors purely from CVs that randomly varied race (White, Black, Asian), gender and discipline. Both Black and Asian professors were rated as having significantly less interpersonal skill than White professors before a single class was experienced — a pure expectancy effect. (Professors in science were also judged more competent than those in the humanities, a reminder that discipline confounds sit alongside race.)

Together these three designs triangulate the same conclusion from different angles: an observational scrape (Reid), a controlled field quasi-experiment (Chávez & Mitchell), and a randomised vignette study (Bavishi et al.). The convergence across methods is what makes the finding credible — each design's weaknesses are covered by another's strengths.

Why it matters for course evaluation in practice

Student Evaluation of Teaching (SET) scores feed reappointment, tenure, promotion, teaching-award and workload decisions at most universities. If race-linked measurement bias is baked into the raw number, then comparing a minoritised lecturer's 4.1 against a departmental "benchmark" of 4.3 is not comparing teaching quality — it is partly comparing skin colour, accent and name. For European quality-assurance officers this is not only an equity problem but a validity problem: a measure that varies with an irrelevant instructor attribute fails the basic measurement requirement of construct validity, and any accreditation narrative built on naive cross-instructor ranking is evidentially weak.

The practical consequence is that the interpretation layer matters as much as the instrument. Three habits do most of the damage: (1) treating small mean differences as meaningful; (2) ranking individuals against a single institution-wide average; and (3) reading numeric scores without the qualitative context that explains them. Each amplifies whatever bias is present.

Limitations and honest caveats

A critical reader should hold these findings with appropriate care:

  • Generalisability. Reid's data come from a voluntary, self-selected US website (RateMyProfessors) at elite liberal-arts colleges — not representative of official European evaluation systems, and subject to its own selection bias. Vignette studies like Bavishi et al. measure stereotype activation, not necessarily behaviour in a real classroom over a semester.
  • Effect sizes are moderate, not deterministic. These are average tendencies. Many minoritised instructors score highly; race explains a modest share of variance. The claim is "systematic bias exists," not "every rating is wrong."
  • Confounding. Race correlates with discipline, seniority, native language, class size and course type. Disentangling the unique racial component is hard, and estimates vary across studies.
  • Replication and context. Most large studies are US-based; European cohorts differ in composition, norms and survey culture. Some replication attempts of bias-mitigation interventions yield variable results across departments. The direction of the effect is robust; its precise magnitude in any given institution is an empirical question best answered with local data.

Stating these caveats is not hedging — it is the same methodological honesty a QA committee should demand before acting on any evaluation number.

How Koji incorporates this

Koji is an AI-native course-evaluation platform built around AI-moderated conversational interviews rather than a static Likert form. Several design choices are intended to mitigate (not eliminate) race-linked bias:

  • Probing the reason, not just the rating. When a student gives a low score, Koji's AI moderator follows up with open_ended questions that surface why. A theme of "the lecturer explained concepts unclearly" is actionable evidence; a pattern of vague negativity that evaporates under probing is a signal that the rating may reflect expectancy rather than teaching. Making the reasoning explicit gives evaluators a way to distinguish substance from bias.
  • Automatic thematic analysis with bias-aware reporting. Koji clusters open-text responses into themes and can flag when person-directed, identity-adjacent language (e.g., comments about accent or "fit") dominates over course-directed feedback — prompting human reviewers to read those comments critically rather than aggregate them blindly.
  • Structured separation of course vs instructor. Using distinct question types (scale, single_choice, open_ended, ranking) Koji can separate feedback on course design, materials and assessment from feedback on the individual — limiting the halo by which one identity impression colours every item.
  • Within-context benchmarking, not single-average ranking. Koji's reporting is designed to compare like with like (level, discipline, cohort, format) and to present distributions rather than a lone mean, discouraging the cross-instructor rankings that most amplify bias.
  • Triangulation across cohorts and cycles. Because Koji supports mid-cycle/formative collection and longitudinal tracking, an instructor's evidence base is broader than a single end-of-term snapshot, reducing the leverage any one biased cohort holds.

None of this removes bias from students' minds. Framed honestly, the aim is to reduce the chance that bias silently becomes a personnel decision — by adding reasoning, context and fair comparison around the number. Koji's core research platform at koji.so applies the same AI-moderated interview engine to product and customer research where probing beyond a rating is equally valuable.

Intersectionality and the European context

Race rarely acts alone. Women of colour, non-native-speaking instructors of colour, and minoritised faculty in quantitative disciplines can sit at the intersection of several documented bias axes at once — and the literature on gender, accent and discipline bias should be read alongside the race evidence rather than in isolation. Reid's (2010) finding of a race effect with no gender main effect does not mean gender is irrelevant; it means the two can operate on different students, in different courses, and sometimes interact in ways a single main-effect model misses. Committees that treat "bias" as one box to tick will under-count the compounded disadvantage facing instructors at these intersections.

The European picture also differs in important ways from the US studies that dominate this literature. Student cohorts are frequently multilingual and internationally mobile; "minoritised" is defined differently across national systems; and formal SET instruments — rather than voluntary websites like RateMyProfessors — are the norm, which changes both who responds and how. This is a reason to localise the evidence, not to dismiss it: a Dutch, German or French institution should test for race- and origin-linked patterns in its own anonymised data rather than assume US effect sizes transfer unchanged. The robust, cross-method direction of the finding — minoritised instructors rated lower for equivalent teaching — is the part that travels; the magnitude is an empirical question every quality-assurance office can and should investigate before any score feeds a personnel file. Doing so is itself good practice under European quality frameworks, which expect institutions to assure the fairness and validity of the evidence they act on, not merely to collect it.

Related Resources

References

  • Reid, L. D. (2010). The role of perceived race and gender in the evaluation of college teaching on RateMyProfessors.Com. Journal of Diversity in Higher Education, 3(3), 137–152. https://doi.org/10.1037/a0019865
  • Chávez, K., & Mitchell, K. M. W. (2020). Exploring bias in student evaluations: Gender, race, and ethnicity. PS: Political Science & Politics, 53(2), 270–274. https://doi.org/10.1017/S1049096519001744
  • Bavishi, A., Madera, J. M., & Hebl, M. R. (2010). The effect of professor ethnicity and gender on student evaluations: Judged before met. Journal of Diversity in Higher Education, 3(4), 245–256. https://doi.org/10.1037/a0020763