Beauty in the Classroom: Does Physical Attractiveness Bias Student Evaluations?
Better-looking instructors get better teaching ratings — an effect first documented two decades ago and replicated since. There is no defensible teaching-quality reason for it, which makes it one of the cleanest demonstrations that a Likert mean is not a pure measure of teaching.
Koji Education Team
Product · June 21, 2026
Short answer: Yes. Multiple studies, beginning with Daniel Hamermesh and Amy Parker's 2005 "Beauty in the Classroom," find that instructors rated as more physically attractive receive measurably higher student evaluations of teaching — an effect that holds within the same department and even the same course, and that tends to be larger for male instructors. Because physical appearance has no plausible connection to teaching quality, attractiveness bias is one of the clearest pieces of evidence that an evaluation score reflects more than how well someone teaches.
The finding that should unsettle every promotion committee
Of all the biases documented in student evaluations of teaching (SET), attractiveness bias is the one with the least possible justification. You can at least construct a story in which an "easy" course earns higher ratings because students genuinely had a better experience. You cannot construct any honest story in which an instructor's bone structure improves their pedagogy. And yet appearance moves the scores.
The foundational study is Hamermesh and Parker's "Beauty in the Classroom: Instructors' Pulchritude and Putative Pedagogical Productivity," published in Economics of Education Review in 2005. The design was straightforward: the authors took end-of-semester student ratings for 463 courses taught by 94 instructors at the University of Texas at Austin, and had six independent observers rate each instructor's physical appearance from a photograph on a beauty scale. The result: instructors viewed as better looking received higher instructional ratings, and the authors described the impact of moving from the 10th to the 90th percentile of the beauty distribution as substantial. The effect persisted within departments and within specific courses, and was larger for male than for female instructors.
This was not a fluke of one American campus. Replications and extensions in other countries — including European and Italian samples — have found broadly similar patterns, and the popular site RateMyProfessors, where students could once award a "hotness" rating, gave researchers a vivid (if informal) corpus showing attractiveness travelling alongside other ratings. The direction of the effect is consistent enough to take seriously.
Why this is the canary in the coal mine
Attractiveness bias matters out of all proportion to its size, because of what it proves. If a variable with zero teaching relevance can systematically shift evaluation scores, then the score is not a clean measurement of teaching. It is a measurement of teaching plus an unknown amount of noise and bias — and you usually cannot tell, from the number alone, how much of each you are looking at.
That single fact reframes how the number should be used. Course evaluations are frequently fed into hiring, contract renewal, pay, and promotion decisions. As Hamermesh and Parker themselves noted, allowing non-teaching factors like physical attractiveness to influence ratings that drive those decisions is exactly what makes the bias problematic. A difference of a couple of tenths of a point — the kind of gap that committees routinely treat as meaningful — can be entirely manufactured by appearance, gender, accent, or age before any teaching has been assessed at all.
Attractiveness rarely travels alone, either. It correlates with, and compounds, the other documented biases in SET — gender, age, and race and ethnicity. An instructor who is older, female, racialised, and judged less conventionally attractive does not face these as separate, additive penalties; they interact in ways a raw average will never reveal.
"But critics argue the effect is small and confounded"
The strongest objection deserves a fair hearing. Skeptics make two points. First, that the measured effect sizes are modest, and that "attractive" instructors might genuinely be more engaging — perhaps appearance correlates with confidence, energy, or warmth that does help students. Second, that beauty ratings made from photographs are not the same as the in-the-room perception that drives real evaluations, so the studies may not capture the actual mechanism.
Both points have merit, and an intellectually honest reading concedes them. The effect is not enormous; appearance does not swamp teaching. And the causal pathway is genuinely tangled — "attractiveness," "charisma," and "approachability" are hard to separate cleanly.
But neither objection rescues the Likert average. Even if part of the beauty effect runs through charisma rather than face shape, that does not make it a measure of teaching effectiveness — it makes it a measure of likeability, which is not the same thing and which we should not be quietly rewarding in tenure files. And "the effect is small" is cold comfort when high-stakes decisions are routinely made on differences smaller than the bias itself. A bias does not have to be large to be unfair when the decisions it feeds are fine-grained. This is the same trap we describe in Is a 0.3-Point Difference Real? — small, biased differences treated as if they were precise signals.
There is also a structural reason the bias is hard to police: it enters at the exact moment a student converts a complex, multi-week experience into a single global score. That conversion is where snap impressions — appearance among them — do their work, because a person reaching for one number leans on the most available overall feeling rather than reconstructing what actually happened in each session. This is why appearance bias and the halo effect are mechanically linked: both exploit the global rating. It also explains why the bias is nearly invisible in aggregate reporting. A department comparing instructor means sees only the final numbers, with no trace of the appearance, gender, or age judgements folded into them — which is precisely why a criterion-referenced reading against a fixed standard is safer than ranking instructors against each other on figures that quietly encode bias.
What evaluation should do instead
You cannot make students stop perceiving appearance. Bias of this kind is baked into snap human judgement, and no survey instrument will switch it off. What you can do is design evaluation so that a global, appearance-tinged impression has less room to dominate the result:
- Anchor questions to specific, observable events, not global impressions. "Was the structure of each session clear?" and "Did you get feedback in time to act on it?" pull responses toward what happened rather than how the instructor came across.
- Capture reasons, not just scores. A number is where halo and appearance bias hide. The explanation a student gives — what specifically helped them learn — is far harder to generate from a vague positive (or negative) gut feeling.
- Read context, never the bare mean. Differences of a few tenths are within the range that appearance alone can produce. Treat them as noise unless corroborated by other evidence — the triangulation principle again.
How Koji helps
Koji for Education is designed around the assumption that a single global rating is contaminated — by appearance, and by everything else students bring into the room. Rather than collecting one impression-driven number, Koji runs an AI-moderated conversational interview that pushes students from "I liked it" toward what specifically happened and why it helped. Its six structured question types let you replace global "overall, how good was this instructor" items with concrete, behaviourally anchored questions that give halo and appearance bias far less purchase.
Because the moderation is standardized and bias-aware, the same neutral probing reaches every student — there is no charismatic human interviewer to charm and no inconsistency between sessions. The automatic thematic analysis then organises feedback around evidence of teaching and learning — clarity, feedback, support, workload — rather than around an aggregate likeability score, and reports at programme and institution level so a committee can see what is driving a number before acting on it. Koji cannot make appearance bias disappear from human perception; what it can do is stop a single appearance-tinged number from being the whole evaluation, and surface the substance underneath.
The same conversational engine underlies the main Koji platform for user and customer research, where first impressions distort feedback in exactly the same way.
Beauty in the classroom is not a quirky footnote. It is the cleanest available proof that a teaching-evaluation average measures more than teaching — and a standing argument for collecting evidence rich enough to separate the substance from the surface.
Want evaluations that measure teaching, not first impressions? Explore Koji for Education.