Central Tendency Bias: Why Low Variance in Course Evaluations Is Not Consensus
When every course scores between 3.8 and 4.2, administrators read agreement. The rater-error literature reads something else: central tendency, range restriction, and an instrument that has stopped discriminating. Here is how to tell the difference.
Koji Education Team
Product ·
Bottom line up front: When course evaluation scores cluster tightly around the middle-to-upper end of the scale — almost every course landing between 3.8 and 4.2 out of 5 — it is tempting to conclude that teaching is uniformly good and students broadly agree. The rating-error literature offers a less flattering explanation. Central tendency (raters avoiding the extremes) and range restriction (the scale failing to spread genuinely different teaching across its full width) can manufacture the appearance of consensus out of an instrument that has simply stopped discriminating. Low variance is not evidence of agreement. It is, just as often, evidence that your scale is no longer measuring anything.
This distinction matters because almost every consequential use of student evaluation of teaching (SET) data — ranking instructors, flagging "underperformers", comparing departments, informing tenure — depends on the differences between scores being real. If the spread is artefactual, every decision built on it inherits the artefact.
The rating-error tradition is forty years old, and SET ignored most of it
The systematic study of how raters distort ratings predates student evaluations as a mass instrument. The canonical reference is Saal, Downey, and Lahey's 1980 Psychological Bulletin review, Rating the ratings: Assessing the psychometric quality of rating data, which catalogued the recurring ways a rater's scores diverge from the thing being rated. Three of those errors are directly relevant to any Likert-style course evaluation:
- Halo — the rater's global impression of the instructor bleeds into every specific item, so "clarity of assessment", "quality of feedback", and "approachability" all move together regardless of their true, separate quality.
- Leniency (or severity) — the rater systematically shifts scores up (or down) relative to what the performance warranted.
- Central tendency — the rater avoids the ends of the scale and clusters responses around the midpoint, compressing real differences into a narrow band.
Saal and colleagues made an argument that the SET field has been slow to absorb: these are properties of the measurement, not of the teaching. A distribution that is tightly bunched can reflect genuinely uniform quality — or a central-tendency response style, a ceiling, or an instrument too blunt to separate a transformative course from a merely adequate one. You cannot tell which from the mean alone.
Central tendency and range restriction are different problems that look identical
It is worth separating two mechanisms that both produce the same flat, narrow distribution.
Central tendency is a rater behaviour. Faced with a 1–5 scale, many respondents are reluctant to award a 1 or a 5. A 5 feels like a claim of perfection; a 1 feels like an accusation. Cautious raters retreat to 3 and 4, and the available range collapses to two usable points. This is amplified by the social context of course evaluation: students often like their instructor personally, are aware that low scores can have consequences, and have no incentive to use the bottom of the scale.
Range restriction is a distributional property of the resulting data. When the observed scores occupy only a fraction of the theoretical scale, the variance shrinks — and shrunken variance has a brutal statistical consequence. Correlations and reliability coefficients are bounded by the variance available to them. An instrument whose scores never leave the 3.5–4.5 band cannot, by construction, demonstrate that it distinguishes good teaching from bad, because there is almost no variation for any criterion to align with. The classic ceiling-effect pattern — where most courses pile up near the top of the scale — is range restriction's most visible form, and we have written about it separately in why almost every course scores four out of five.
The two combine viciously. Central-tendency response styles and leniency push scores into a narrow upper band; range restriction then guarantees that whatever real differences in teaching exist are invisible inside it. The committee sees 4.1 versus 3.9 and treats it as a finding. It is noise inside a compressed range.
Why "everyone scored about the same" is the wrong inference
Administrators routinely read low dispersion as a quality signal: our teaching is consistently strong; look how little variation there is. Three reasons to distrust that reading:
-
A flat distribution is consistent with both uniform quality and a dead instrument. The data look identical in both cases. Without independent evidence — peer observation, learning outcomes, qualitative depth — you cannot adjudicate between them. We make the broader version of this argument in triangulating teaching evaluation across multiple evidence sources.
-
Bimodality hides inside a moderate mean. A 3.9 can be a genuine consensus around "good", or it can be two camps — a group who found the course excellent and a group who found it alienating — averaging to a centrist fiction nobody actually reported. The mean erases the disagreement. We dissect this failure mode in how the mean hides dispersion and bimodality.
-
Compressed ranges break every downstream comparison. If your usable range is one scale point wide, then ranking, benchmarking, and flagging are operating on differences smaller than the instrument's own measurement error. This is the same problem we raise in is a 0.3 difference a real effect: the spread is too small for the comparisons people insist on making.
But doesn't low variance sometimes just mean the teaching really is consistent?
This is the strongest objection, and it is sometimes correct. Teaching quality within a well-run department genuinely can be uniformly high, and in that case a narrow distribution is an honest summary, not an artefact. We should not pretend every flat distribution is a measurement failure.
The point is not that low variance is always bias — it is that you cannot diagnose which it is from the numbers alone, and most institutions never try. The differential diagnosis is straightforward when you look for it:
- If scores are compressed and the qualitative comments are rich, specific, and varied — describing different strengths and weaknesses across courses — you likely have genuine quality with a scale too coarse to register the texture. The teaching varies; the Likert items just cannot see it.
- If scores are compressed and the open text is thin, generic, or absent, you more likely have a response-style artefact: students disengaging from a ritual, defaulting to a safe 4, and telling you nothing.
The variance in the numbers cannot distinguish these. The variance in the language can. An instrument that collects only scale points throws away exactly the evidence needed to interpret the scale points. That is the design flaw at the heart of legacy SET.
What actually helps
The rater-error literature points to concrete mitigations, most of which legacy survey tools handle poorly:
- Stop relying on the global mean as the unit of analysis. Report distributions, not just averages; show how much of the scale is actually being used. A score with no visible spread should be treated as low-information by default.
- Use behaviourally specific items, not global impressions. Low-inference questions ("How often did you receive feedback you could act on?") resist halo and central tendency better than "Rate the overall quality of teaching", because they ask about observable events rather than inviting a single gut number. This is the practical upshot of writing better course evaluation questions.
- Treat open-text feedback as primary evidence, not decoration. The qualitative channel is where range restriction is broken, because language is not bounded to five points. The barrier has always been that nobody can read and code thousands of comments consistently — which is precisely where modern tooling changes the economics.
Where Koji fits
Central tendency and range restriction are, at root, failures of a static five-point scale to make people commit to a real signal. Koji for Education attacks the problem from a different direction: instead of asking a student to compress a complex experience into one number, its AI-moderated conversational interview probes beyond the rating — when a student gives a noncommittal answer, the moderator follows up, asks for a concrete example, and surfaces the why behind the number. That follow-up is exactly the information a flat Likert distribution destroys.
Because the moderation is standardized and bias-aware, every student gets the same calibrated probing — removing the human-moderator inconsistency that plagues focus groups, while avoiding the dead-flat response styles that plague self-administered surveys. Koji's automatic thematic analysis then reads the full open-text corpus consistently, so the qualitative channel that breaks range restriction is finally usable at scale rather than skimmed by an overloaded committee. Its quality scoring flags low-information responses instead of averaging them in as if they carried signal. And because Koji supports six structured question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no), scales remain available where they are genuinely informative — but they are no longer the only thing you collect.
The same AI interview engine powers the main Koji platform for general user and customer research, where the identical problem — people defaulting to a safe middle answer on a survey — costs teams real insight.
Koji does not claim to eliminate central tendency; no instrument can stop a cautious rater from being cautious. What it does is reduce the field's dependence on the one number that central tendency most easily corrupts, and surface the richer evidence that tells you whether a flat distribution means consensus or a dead instrument.
The takeaway
Low variance in course evaluation scores is a question, not an answer. It is equally consistent with uniformly excellent teaching and with a scale that has stopped discriminating — and the mean alone can never tell you which. Before you rank, flag, or reward on the strength of differences inside a compressed range, ask whether the range is compressed because your teaching is uniform or because your instrument is blunt. The qualitative evidence, read properly, is what answers that question.
Evaluating teaching at a European institution and tired of a flat 4.1 that tells you nothing? See how Koji for Education turns a single ambiguous number into evidence you can act on.