Why Almost Every Course Scores 4 out of 5: Ceiling Effects and Restricted Range in Course Evaluation
When nearly every instructor lands between 4.0 and 4.5 on a 5-point scale, the number has almost stopped measuring anything. Here is why ceiling effects and restricted range quietly break the comparisons universities make most, and what to do instead.
Koji for Education
Research & Editorial Team · June 13, 2026
Answer first: Course evaluation scores cluster near the top of the scale because the rating distribution is negatively skewed and the usable range is compressed — a ceiling effect. When most courses score between 4.0 and 4.5 out of 5, the differences between them are statistically tiny and substantively meaningless, yet committees routinely treat a 4.1 as worse than a 4.4. The fix is not a better number; it is reporting distributions instead of means, attaching uncertainty to every figure, and adding evidence that actually discriminates between courses.
The pattern every evaluation office has seen
Pull the institutional means for a department and the histogram is almost always the same shape: a tight cluster sitting high on the scale, a long thin tail trailing down toward the middle, and almost nothing at the bottom. On a five-point "overall satisfaction" item, the bulk of courses land somewhere between 3.9 and 4.5. A course at 4.5 is celebrated; a course at 3.9 triggers a conversation. The gap feels meaningful. It usually is not.
This is a ceiling effect: when a measure cannot register values above a limit, scores pile up against that limit and the instrument loses its ability to distinguish the genuinely excellent from the merely good. Closely related is restricted range — because the realistic spread of scores is so narrow, the variance that statistics need in order to separate cases has been squeezed out. A scale nominally running 1 to 5 is, in practice, operating across a one-point band near the top.
Why the scores pile up at the top
Several well-documented mechanisms push course evaluation distributions toward the ceiling:
- Genuine positivity. Most teaching is competent and most students are reasonably content. A negative skew partly reflects reality, not just artefact.
- Acquiescence and leniency. Respondents drift toward agreement and toward the kinder end of a scale, especially on quick end-of-term forms completed with little deliberation.
- Halo effects. A general positive impression bleeds across every item, so a well-liked instructor scores uniformly high on dimensions students never actually considered separately. In a controlled study across three universities, Jared Keeley and colleagues found that both halo and ceiling/floor effects in student evaluations were robust and persisted even when the rating scale was expanded from 5 points to 7 or 9 points (Keeley, English, Irons & Henslee, 2013, Educational and Psychological Measurement). Simply lengthening the scale did not buy back the lost discrimination.
The practical consequence is that the "4.2" on a report card carries far less information than its two significant figures imply. As Philip Stark and Richard Freishtat argued in their widely cited An Evaluation of Course Evaluations (2014), averaging responses to an ordinal scale and then comparing those averages to two decimal places gives "a false sense of precision" — the differences institutions act on are smaller than the noise in the measurement.
The real damage: over-interpreting tiny differences
Ceiling effects would be harmless if everyone treated a high cluster as "this department teaches well." The damage comes when the compressed range is mined for rankings. When the entire faculty sits between 4.0 and 4.5, a 0.3 difference looks like a meaningful gap — but it can be produced by two or three students, a slightly different response rate, or the random composition of who happened to fill in the form.
The evidence that humans over-read these differences is direct. Guy Boysen has shown that faculty and administrators interpret small mean differences in student ratings as practically significant even when explicitly warned not to (Boysen, 2015, Scholarship of Teaching and Learning in Psychology). The instrument compresses the signal; human judgement then amplifies the residual noise. That is the worst possible combination for a number used in promotion, tenure, and probation decisions — a use the broader evidence base already cautions against, as we discuss in whether student evaluations should decide tenure and promotion.
Restricted range also quietly corrupts the statistics built on top of these scores. Correlations between student ratings and any other variable — learning gains, peer review, future performance — are mechanically attenuated when one variable has little variance. A near-zero correlation in a ceiling-bound dataset can mask a real relationship; a modest one can be inflated by a handful of low outliers. Either way, the validity question (do these scores measure teaching effectiveness?) cannot be answered cleanly with a variable that barely moves. This is the same reliability-versus-validity distinction we unpack in are course evaluations valid?.
But isn't a high score just evidence of good teaching?
This is the strongest objection, and it deserves a fair hearing. If most teaching genuinely is good, a left-skewed distribution is the honest result, and demanding more spread would mean manufacturing dissatisfaction that does not exist. Critics of the "ceiling effect" framing argue that statisticians are pathologising good news.
Two things are true at once. First, yes — the skew is partly real, and no one should respond by re-engineering scales to force a normal distribution or by treating a 4.3 as a problem. Second, and decisively: even if the level is meaningful, the differences within the cluster are not. The objection defends the headline ("teaching here is well-regarded") while doing nothing to rescue the comparison ("Course A's 4.4 beats Course B's 4.1"). It is precisely the comparison, not the level, that evaluation offices are pressed to deliver and that committees misuse. Accepting that teaching is broadly good is an argument for abandoning fine-grained ranking, not against the ceiling-effect critique.
A second objection: "then widen the scale." Keeley's experiment already answered this — moving to 7 or 9 points did not dissolve the effect. The problem is not the number of boxes; it is that a single satisfaction rating asks students to compress a rich experience into one summary judgement, and summary judgements gravitate to the top when the underlying impression is positive.
What to do instead
You cannot average your way out of a ceiling. You can change what you collect and how you report it.
- Report distributions, not means. Show the full spread — proportion in each category, or at minimum the median and interquartile range. A course where 70% say "excellent" and 30% say "poor" and a course where everyone says "good" can share a mean and mean completely different things. Averaging hides exactly the bimodality that matters, a point we make at length in why averaging Likert scores misleads.
- Attach uncertainty to every figure. With typical class sizes and response rates, the confidence interval around a course mean often spans most of the compressed range. Once intervals are shown, the "ranking" between 4.1 and 4.4 visibly collapses into overlap.
- Stop deriving league tables from a saturated metric. Norm-referencing courses against each other on a ceiling-bound scale guarantees that random noise decides who sits below the line.
- Add a source that actually discriminates. The way to recover lost signal is not a finer scale but richer evidence — open-ended, probed, qualitative feedback that varies meaningfully across courses even when the numbers do not. This is the logic of triangulation: student ratings are necessary but not sufficient.
Where Koji fits
Koji for Education is built on the premise that the single number is the weakest part of course evaluation, not the centrepiece. Instead of asking students to collapse a term into one Likert response, Koji runs AI-moderated conversational interviews that probe beyond the rating — asking why a course worked, for whom, and where it broke down. The result is feedback that varies across courses even when satisfaction scores are uniformly high, restoring the discrimination a ceiling-bound mean has lost.
Because Koji performs automatic thematic analysis of that open-text and conversational data, evaluation teams get a structured, comparable picture of content — the recurring strengths and concrete failure points — rather than a spuriously precise decimal. Its reporting is designed to surface distributions and themes rather than encourage decimal-place ranking, and its quality scoring flags thin or low-information responses instead of averaging them into the noise. None of this eliminates ceiling effects in any scale items a university chooses to keep — Koji mitigates the problem by adding signal where the numbers have gone flat, and by reporting honestly on the numbers that remain.
The same conversational interview engine powers the main Koji platform for general customer and user research, where the "everything scores 8/10" problem is just as corrosive to product decisions as it is to teaching ones.
If your evaluation reports are full of 4-point-somethings that no longer tell anyone apart, the answer is not a new scale. It is better evidence. See how Koji for Education surfaces what the average hides.