Do Students Rate Instructors "Like Them" Higher? Similarity Bias in Course Evaluations
Similarity (homophily) bias in student evaluations is real but smaller, more conditional, and more confounded than headline demographic-bias claims. The honest reading of the evidence — and what it means for how you should and should not use the scores.
Koji Education Team
Product ·
Bottom line up front: There is credible evidence that students sometimes rate instructors who resemble them — by gender, and to a lesser and more confounded degree by ethnicity — slightly more favourably. But the effect is smaller, more conditional, and more entangled with who happens to be in the room than the "students are biased against people unlike them" headline suggests. The defensible conclusion is not "correct the scores for demography." It is more unsettling: a single Likert average quietly bundles a student''s reaction to who is teaching with their reaction to how well they were taught, and you usually cannot tell the two apart. The fix is not a statistical adjustment. It is a better instrument.
What "similarity bias" actually means
Social psychologists have known since Donn Byrne''s similarity-attraction work in the 1960s that people tend to like, trust, and rate more highly those they perceive as similar to themselves. Applied to course evaluation, the worry is homophily: that a student gives a marginally warmer rating to an instructor who shares their gender, ethnicity, first language, or background — not because the teaching was better, but because similarity feels like rapport.
This is a distinct claim from the more familiar "demographic bias" finding. Gender bias research typically asks whether women (or minority instructors) receive lower scores on average. Similarity bias asks something subtler and interactional: does the match between rater and instructor move the score, regardless of the instructor''s group? The two questions have different answers, and conflating them is one of the most common mistakes in this literature.
What the evidence shows
The strongest similarity signal is for gender, and it is concentrated on the student side. In a widely cited study using data from a French university, economist Anne Boring (2017, Journal of Public Economics) found that male students gave significantly higher evaluations to male instructors, and that the teaching dimensions students rewarded tracked gender stereotypes — men were perceived as more knowledgeable and as stronger class leaders — even though students appeared to learn just as much from women. The asymmetry matters: the effect was driven largely by male raters, not by a symmetric "everyone prefers their own kind."
Experimental work points the same way. In MacNell, Driscoll and Hunt''s much-discussed 2015 online-teaching experiment (Innovative Higher Education), instructors were rated more highly when students believed they were male — regardless of the instructor''s actual gender — across most rating dimensions. Because the only thing that changed was the perceived identity, the design isolates bias from teaching quality.
Race and ethnicity are murkier. A 2023 study in Frontiers in Education used virtual avatars delivering identical lectures in the United Arab Emirates (n = 318) and found that students rated South Asian instructors as more approachable, more enthusiastic, and more worth learning from. But the authors were admirably candid about the confound: 65.7% of their participants were themselves South Asian, so the apparent "pro–South Asian" pattern is at least partly in-group favouritism produced by the composition of the rater pool — exactly the homophily mechanism, but a reminder that the same data can read as "bias for" or "bias against" depending on who is answering.
That candour is the whole point. Earlier and larger reviews — for example Centra and Gaubatz (2000) — found same-gender effects that were small and inconsistent across disciplines, with women rating women somewhat more favourably in some fields and no clean mirror image among men. Similarity bias is not a myth, but neither is it the dominant force in a typical evaluation score. It is one more small, hard-to-see contaminant sitting inside the number.
But doesn''t this just mean we should statistically correct for it?
This is the strongest counterargument, and it deserves a straight answer: no, and trying is usually worse than doing nothing.
First, the effect is interactional and context-dependent. To "correct" a score you would need to model the rater''s demographics, the instructor''s demographics, the discipline, the class composition, and their interactions — and you almost never have anonymous raters'' demographics, nor should you, because collecting them undermines the anonymity that makes feedback honest in the first place. Second, applying a blanket adjustment ("add 0.1 to female instructors in male-dominated cohorts") hard-codes a population-average estimate onto an individual whose real effect could be zero. You would be replacing an unknown bias with a known, indefensible one. Third, the UAE study shows the sign of the effect can flip with the rater pool, so any fixed correction is wrong somewhere by construction.
The honest position is the uncomfortable one: similarity bias is a reason to distrust the precision of a single global rating, not a quantity you can subtract out. It belongs in the same bucket as the halo effect and negativity bias — small distortions of a global impression that you mitigate by not relying on a global impression.
Why the instrument is the problem
Similarity bias does its damage when the evaluation asks for a diffuse, affect-laden judgement: Overall, how would you rate this instructor? That question invites the student to consult their feeling of rapport, and rapport is precisely where "this person is like me" leaks in. The more a question pulls for a global gut reaction, the more room there is for who-the-instructor-is to colour how-they-taught.
The mitigation is well established in survey methodology: ask about specific, observable teaching behaviours rather than overall impressions, and make students explain their judgements rather than just register them. "Which explanation or example helped you understand a difficult concept?" is far harder to answer on the basis of demographic rapport than "Rate this instructor 1–5." This is the same logic behind the case against averaging Likert scores and the broader question of whether evaluations are even valid.
How Koji reduces the surface area for similarity bias
Koji for Education cannot eliminate bias — no instrument can, and we will not claim otherwise. What it can do is shrink the space in which similarity bias operates and make what remains visible.
- Conversational probing instead of a single number. Koji''s AI-moderated interviews ask students to describe concrete moments — what helped them learn, where they got stuck, what they would change. Anchoring feedback to specific behaviours rather than a global impression is the best-evidenced way to dilute halo and affinity effects.
- Standardized, bias-aware moderation. Every student is interviewed by the same AI moderator following the same protocol, removing the human-moderator inconsistency that adds its own noise — and the moderation is designed not to flatter or lead.
- Six structured question types (open-ended, scale, single-choice, multiple-choice, ranking, yes/no) so quantitative items are paired with the reasons behind them, rather than left to stand alone.
- Automatic thematic analysis of open-text responses surfaces what students actually said about teaching, separating substantive feedback from diffuse sentiment.
- Programme- and institution-level reporting lets quality teams look at patterns across many instructors and cohorts, where idiosyncratic rater-instructor matches wash out, rather than over-reading one teacher''s mean.
Koji runs on the same AI interview engine as the main koji.so platform that teams use for general user and customer research — the difference is the education-specific question design, GDPR/AVG-appropriate data handling, and reporting built for quality assurance.
The takeaway for quality officers
Treat similarity bias as evidence about the limits of the number, not as a number to adjust. It is one more reason that a 0.2-point gap between two instructors should never decide anything on its own — see Is a 0.3-point difference real? — and one more reason to triangulate student ratings with peer observation and learning evidence. The most useful thing an evaluation can do is tell you what helped students learn and what got in the way. A score can''t do that. A conversation can.
Curious what behaviour-anchored, bias-aware course evaluation looks like in practice? See how Koji for Education works.