New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Evaluation bias8 min read

Shifting Standards: Why Subjective Rating Scales Hide the Bias Your Data Looks Clean Of

A 1-to-5 rating scale can show no gender gap and still be biased. Biernat's shifting-standards model explains why subjective scales mask stereotyping that objective measures reveal — and what it means for how you read course evaluations.

Koji Education Team

Product ·

The short answer: A course-evaluation item like "Rate this instructor's competence from 1 to 5" can return identical averages for men and women and still be contaminated by stereotyping. Monica Biernat's shifting standards model shows that subjective scales let raters apply a different yardstick to different groups, so bias hides inside a score that looks even-handed. To see it, you have to anchor judgments to common, objective referents — or ask people what they actually observed, not how they rate it.

Why a "no difference" result is not the same as "no bias"

Most institutions test their evaluation data for bias the obvious way: they compare mean ratings for male and female instructors, or across ethnic groups, and check whether the gap is statistically significant. When the gap is small or absent, the conclusion is reassuring — our process is fair.

The shifting standards literature is a warning that this test is underpowered by design. Biernat, Manis and Nelson showed across a programme of studies that subjective response scales (good–bad, weak–strong, 1–7) invite raters to use group-specific standards, while objective or common-rule measures (estimate the height in centimetres, count the behaviours, assign a grade in points) force a single yardstick and expose the stereotype (Biernat & Manis, 1994, Journal of Personality and Social Psychology; Biernat, Manis & Nelson, 1991).

The classic demonstration: when people judged a woman described as a "very good" parent against a man described the same way, the subjective labels matched — both "very good." But when asked to estimate objective behaviours, the woman was judged to perform significantly more parenting acts than the identically described man. The word "good" silently meant something different depending on whom it was applied to. The standard shifted.

How this plays out in a course evaluation

Translate that into a teaching-evaluation item. "This instructor is an effective teacher: strongly disagree … strongly agree." That is a subjective scale. The phrase "effective" has no fixed anchor; each student calibrates it against what they expect from this kind of instructor.

Shifting standards predicts two things that should unsettle anyone who reads evaluation dashboards:

  1. Null or even reversed gaps on subjective items. Because raters lower the minimum standard for a stereotyped-as-less-competent group ("she's good for a woman in engineering"), a woman can clear the bar for "agree" more easily — producing a mean that looks equal or higher, while the underlying judgment is still stereotyped. Biernat calls these the masking and contrast effects of subjective scales.
  2. Bias that surfaces only on consequential, objective decisions. The same raters who give equal subjective ratings apply higher confirmatory standards to the stereotyped group when something concrete is at stake — a nomination, a shortlist, a "would you recommend." Biernat and Vescio's memorable title captures it: "She Swings, She Hits, She's Great, She's Benched" — praised in the abstract, passed over in the decision (Biernat & Vescio, 2002, Personality and Social Psychology Bulletin).

This is not a fringe idea applied speculatively to teaching. Researchers have taken the model directly into higher education. A 2022 study in Higher Education Research & Development applied shifting-standards theory to student nominations of teaching excellence and found that student conceptions of excellence conformed to gender biases, with male students disproportionately less likely to nominate a female teacher (Heffernan, 2022). Experimental work in school settings shows the same signature: stereotype-based shifting standards appear in grading and written feedback depending on whether the scale is subjective or objective (Springer, Social Psychology of Education, 2021).

The famous experiment, re-read through this lens

The MacNell, Driscoll and Hunt experiment is usually cited as blunt proof of gender bias: in an online course where the same instructors taught under swapped identities, the instructor students believed was male scored 4.35/5 while the one believed female scored 3.55, despite identical conduct (Driscoll, MacNell & Hunt, 2015, Innovative Higher Education).

Read through shifting standards, the lesson is sharper. That study found a gap because it held everything else constant — same person, same work, same timing — so the only thing left to vary was the standard students applied. In real institutional data, instructors are not identical twins under swapped names. The "construct-irrelevant variance" of who is teaching what, to whom, gets tangled with genuine differences in teaching, and the shifting standard hides inside an aggregate mean that no significance test on raw scores will flag. The experiment removes the camouflage; your dashboard restores it.

But doesn't this prove your evaluations are useless?

No — and overclaiming here would be its own error. Three honest qualifications:

  • Shifting standards is a moderator, not a universal verdict. It predicts when bias is masked versus revealed; it does not say every evaluation is hopelessly biased. The size of bias in real student-evaluation data is genuinely contested, with strong experimental evidence on one side and large observational studies finding small or context-dependent effects on the other. We treat that debate honestly in How Strong Is the Evidence That Student Evaluations Are Biased?
  • The fix is not to ban ratings. It is to stop relying on a single subjective number to carry a personnel or quality judgment. Triangulate with peer observation, learning evidence, and structured qualitative data — see Triangulation in Teaching Evaluation.
  • Objective anchoring has costs. Asking "how many times did the instructor give feedback within 48 hours?" is harder to design and answer than "rate responsiveness 1–5." There is a trade-off between the bias resistance of behaviour-anchored items and the burden they place on respondents.

What shifting standards tells you to actually do

The model is unusually actionable because it names the mechanism. To reduce masked bias, move judgments from subjective scales toward common, observable referents:

  • Replace global trait ratings with low-inference, behaviour-anchored questions. Not "Was the instructor approachable?" but "When you asked a question, did you get a clear answer in the session?" Concrete referents shrink the room for a standard to shift. This is the same logic behind better evaluation-question design.
  • Probe the why behind a rating instead of accepting the number. A standalone 4/5 is a black box; a follow-up that asks what specifically made the course effective converts a subjective label back into observable evidence you can audit for stereotype content.
  • Read the qualitative layer for shifting language, not just sentiment. Gendered descriptors ("brilliant" vs "caring," "tough" vs "abrasive") are exactly where shifted standards leave fingerprints — and a sentiment score will not surface them.

Where Koji fits

Koji for Education is built around the part of this problem that a static Likert form cannot touch: it replaces the single subjective rating with an AI-moderated conversational interview that asks students to ground their judgments in what actually happened. When a student says a course was "well taught," the interview probes for the specific, observable behaviours behind the label — pulling the judgment back toward the objective end of the scale where shifted standards are far harder to hide.

Because the moderation is standardized across every interview, the calibration that drifts from one human interviewer (or one student's internal yardstick) to the next is held constant. Koji's six structured question types let you pair scale items with open-ended probes, and its automatic thematic analysis surfaces the gendered or stereotyped descriptors in open text at scale — the place shifting standards predicts bias will actually appear. It cannot eliminate bias; no instrument can. It is designed to surface and reduce the masking effect that makes biased data look clean. (Teams running general user research hit the same masking problem in customer feedback — the main Koji platform uses the same AI interview engine for that.)

If your evaluation process has never found bias, that may be good news — or it may be the shifting-standards problem doing exactly what the model predicts. The only way to know is to stop trusting the average and start interrogating the standard behind it.

See how Koji for Education replaces the average-a-Likert-score form with bias-aware conversational interviews — explore Koji for Education.