New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Evaluation bias10 min read

Warmth vs Competence: The Two-Dimensional Bias Hiding in Your Course Evaluation Comments

Students judge instructors on two axes — warmth and competence — and the gap between them is gendered. Understanding the Stereotype Content Model explains why a "likeable" score and an "effective" score are not the same thing, and why women and minoritised faculty are caught in a double bind.

Koji for Education

Research & Editorial Team · June 22, 2026

Bottom line up front: Decades of social psychology show that people evaluate others along two near-universal dimensions — warmth (is this person friendly, trustworthy, well-intentioned?) and competence (is this person capable, intelligent, effective?). Student evaluations of teaching tap both at once, usually without distinguishing them. Because the two dimensions are stereotyped by gender, the same teaching behaviour can earn a man competence and a woman a warmth-or-competence trade-off she cannot win. If your evaluation instrument collapses warmth and competence into one "overall" number, you are not measuring teaching effectiveness — you are measuring a social judgement with a known, gendered structure.

The model: two dimensions, not one

The Stereotype Content Model (SCM), developed by Susan Fiske, Amy Cuddy, Peter Glick and colleagues, holds that social perception is organised around warmth and competence, and that nearly all everyday person- and group-judgements can be located on those two axes (Cuddy, Fiske & Glick, 2008). The two are partly independent: you can be seen as warm but incompetent (the patronised), cold but competent (the envied), or — the goal everyone is chasing — both. Crucially, the dimensions are stereotyped: in Western academic settings, women are stereotypically coded for warmth, nurturance and emotional sensitivity, men for competence, authority and dominance.

This matters for course evaluation because most instruments quietly assume teaching quality is one thing. It is not. A student rating an instructor "excellent" may be reacting to warmth (approachability, encouragement, an easy rapport) or to competence (clarity, rigour, command of the material) — and the two can point in opposite directions for the same teacher.

The double bind for women

The SCM predicts a specific trap, and the evaluation literature documents it. When a woman signals strong competence — runs a demanding course, holds a firm standard, projects authority — she risks being seen as insufficiently warm, "at variance with the female stereotype," and is penalised for the very behaviour that earns men competence credit. The same assertiveness read as "authoritative" in a man is read as "strident" or "cold" in a woman. Research on gender stereotypes in evaluation notes that it can be unusually hard for a woman to be rated highly on both liking and competence at once — the dimensions that students fuse into a single score work against her (Gender-biased evaluation or actual differences?, 2021).

The causal evidence is not subtle. In a five-year natural experiment covering 23,001 evaluations at a French university, Boring, Ottoboni and Stark (2016) found that student evaluations are biased against women by an amount that is "large and statistically significant," that the bias contaminates even ostensibly objective items (such as how promptly work was graded), and that the bias can be large enough to make a more effective instructor receive a lower score than a less effective one. The mechanism is exactly what the SCM predicts: students apply gendered warmth/competence expectations, and the instrument has no way to separate the judgement from the teaching.

This is why our discussions of gender bias in student evaluations and who gives lower scores keep returning to the same structural point: the problem is not only that bias exists, but that the standard instrument is designed to launder it into a number that looks objective.

Warmth is not noise — but it is not effectiveness either

A careful reader will object that warmth is not irrelevant. A warm, approachable instructor may genuinely create better conditions for learning; rapport is part of good teaching, not a contaminant to be scrubbed out. This is correct, and it is the strongest counterargument to "just measure competence."

But it cuts the other way too. If warmth matters, then warmth and competence should be measured separately and explicitly, so that a department can see whether a course is being praised for its rigour or merely for its congeniality — and so that a warm-but-thin course is not mistaken for an excellent one, nor a rigorous-but-demanding course punished for failing a popularity contest. The fault is not in caring about warmth; it is in fusing it with competence into a single global rating that no one can decompose after the fact. A 4.1 "overall" tells you nothing about which axis produced it.

"But aren't these just real differences in teaching?"

The most serious objection is that warmth/competence gaps might reflect genuine differences in how people teach, not bias in how they are perceived. The natural-experiment designs answer this directly. Boring, Ottoboni and Stark's data come from a setting where students were effectively randomly assigned across sections and where some measured outcomes (like final-exam performance in subsequent courses) were available as an objective check. The bias appeared on items where it could not reflect real differences — including the grading-promptness item, identical across the instructors being compared. When the perception of an objectively identical behaviour shifts with the instructor's gender, that is bias, not a difference in teaching. The same logic underlies the famous online-course experiments in which students rated an identical instructor more highly when assigned a male name. Real teaching differences cannot explain a gap produced by a name.

What to do about it

The remedy is not to abolish student voice — students see things no peer observer or self-report can — but to stop asking a single number to do work it cannot do honestly:

  1. Measure warmth and competence as distinct constructs. Ask about specific, low-inference behaviours ("explanations were clear," "feedback helped me improve") rather than global impressions ("overall, an excellent instructor"), so the two axes do not collapse.
  2. Read the comments, structurally. Gendered language — "caring," "nice," "rude" for women; "brilliant," "knowledgeable," "tough" for men — is a signal of which axis a student is rating on. At scale this requires systematic thematic analysis, not a manager skimming a sample.
  3. Triangulate. As we argue in our note on triangulating teaching evaluation, student perception should be one input among several (peer observation, learning evidence, self-reflection), never the sole basis for a high-stakes judgement.
  4. Brief your decision-makers on the SCM. A dean who knows the warmth/competence trap reads "students found her cold" very differently from one who takes it at face value.

How Koji addresses the two-dimensional problem

Koji for Education is built to surface the distinction the traditional Likert form erases. Instead of a single global rating, its AI-moderated conversational interview can probe why a student felt a course worked or did not — separating "I always felt comfortable asking questions" (warmth) from "I could apply the method on the exam" (competence). Its automatic thematic analysis of open-text feedback can flag the gendered descriptive language the SCM warns about, turning a vague pile of comments into a structured pattern a quality officer can actually examine. Because the AI moderator applies the same standardised, bias-aware prompting to every student and every instructor — with none of the inconsistency of a human moderator — it removes one source of variance even as it makes the underlying student judgement legible. To be precise about the claim: Koji cannot eliminate the gendered perceptions students bring to the interview; no instrument can. What it does is surface and disaggregate them, so that warmth and competence stop hiding inside a single deceptively objective number.

Researchers running general user studies face the identical warmth/competence confound when they reduce qualitative interviews to one satisfaction score; the main Koji platform runs on the same interview engine for that work.

The takeaway

There is no such thing as an "overall teaching score" that is innocent of the warmth/competence structure of human social judgement. The dimension is real, it is gendered, and the natural-experiment evidence shows it can reverse the ranking of genuinely effective teachers. The honest response is to measure the two axes separately, read the language students actually use, treat student voice as one evidence source among several, and refuse to let a single global rating quietly encode a stereotype.


Koji for Education replaces global Likert ratings with structured, AI-moderated interviews and thematic analysis that surface the warmth/competence distinction instead of burying it. See how Koji approaches bias-aware evaluation.