New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
analysis-reporting9 min read

When Almost Everyone Scores 4.5: Ceiling Effects, Skew, and What They Do to Course-Evaluation Statistics

Course-evaluation ratings pile up at the top of the scale, producing a strong ceiling effect and negative skew that breaks the statistics most universities still report. Here is what the evidence shows and how to report ratings honestly.

Koji Education Team

Product

Bottom line: Student-evaluation ratings cluster near the top of the scale, producing a ceiling effect and a strongly negatively skewed distribution. That skew breaks the assumptions behind the mean, the standard deviation, and norm-referenced cut-offs that most universities still report — so two instructors with the "same" 4.3 average can have completely different distributions, and a fixed threshold can fail a large share of perfectly competent teachers. Report the full distribution, use rank- or criterion-based summaries, and never read a tenth of a point as a difference.

What a ceiling effect is, and why SET has one

A ceiling effect occurs when a measurement instrument cannot register values above a limit, so scores bunch against the top. Ho and Yu (2015), in their methodological treatment of modern test-score distributions, show that ceiling effects mechanically produce negative skew and distort every moment of the distribution — the mean is dragged toward the ceiling, the standard deviation is compressed, and skewness and kurtosis depart sharply from the normal-curve values that most statistical defaults assume. Their core warning is that summarising a ceilinged, discrete, skewed variable with statistics designed for a smooth normal distribution invites systematic misinterpretation.

Course evaluations are a textbook case. Most students, most of the time, are reasonably satisfied, and the scale (typically 1-5) has nowhere above "excellent" to put them. The result is that the modal response is the top or second-from-top point, the bulk of the mass sits in the top two categories, and the long thin tail points downward. The "average course" is not in the middle of the scale; it is near the ceiling.

What the research says

The most directly relevant evidence comes from Uttl and Smibert (2017), who analysed 14,872 publicly posted class evaluations reflecting input from more than 325,000 students. They document that SET rating distributions are "negatively skewed for most of the selected subjects due to ceiling effects" — a large proportion of students award the highest possible rating. This is not a quirk of one institution; it is the normal shape of SET data.

That shape has teeth when scores are used for decisions. Uttl and Smibert show that norm-referenced standards — labelling instructors below the average, or below a percentile, as "unsatisfactory" — behave very differently from criterion-referenced standards on a ceilinged distribution. Because the mass is squeezed near the top, small absolute differences translate into large rank differences, and an "Overall Mean" cut-off ended up labelling 43.0% of classes as failing to meet the standard. A distribution that is overwhelmingly positive can still, under a norm-referenced rule, brand nearly half of teaching as below par. Their broader finding — that instructors of quantitative courses are at substantially higher risk of being labelled unsatisfactory — is a direct downstream consequence of comparing ceilinged distributions with different shapes.

Two further sources corroborate the reporting problem. Stark and Freishtat (2014), in their widely cited An Evaluation of Course Evaluations, argue that because SET distributions are skewed and "lumpy" (discrete, with mass at a few points), the average is a poor summary and reporting a mean to one decimal place conveys false precision; they recommend distributions, medians, and explicit uncertainty instead. And the cross-cultural replication by Uttl and colleagues (2022) — pointedly titled Spain is not different — reproduces the skew and the quantitative-course penalty in a different national system, showing the ceiling effect is not an artefact of one rating culture.

Why it matters for course evaluation in practice

The ceiling effect quietly undermines the three things universities most often do with evaluation numbers:

  1. Comparing means. When distributions are skewed and compressed, the mean is sensitive to a handful of low ratings and insensitive to the dense top mass. Two instructors at 4.3 can differ wildly: one with a tight cluster at 4-5, another with most students at 5 and a vocal minority at 1-2. The single number hides exactly the information a reviewer needs.
  2. Ranking and cut-offs. On a ceilinged distribution, tiny absolute gaps near the top produce dramatic rank swings, so league tables of instructors are mostly sorting noise. Fixed or norm-referenced thresholds can fail large fractions of competent staff, as the 43% figure illustrates, and they fail them non-randomly — disproportionately in quantitative and demanding disciplines.
  3. Standard deviations and "consistency". Because the ceiling compresses variance, a low standard deviation does not necessarily mean a consistent experience; it can simply mean the scale ran out of room. Treating SD as a clean measure of agreement is unsafe.

The practical upshot: the skew is not a nuisance to be normalised away — it is information. A negatively skewed course with a heavy lower tail is telling you that a real subgroup had a poor experience, which a mean near 4.4 will bury.

A concrete illustration makes the danger vivid. Imagine two seminar groups of the same course. In Group A, 28 of 30 students rate the course 5 and two rate it 1, giving a mean of about 4.7. In Group B, all 30 students rate it 4, giving a mean of 4.0. A league table ranks Group A's instructor far above Group B's — yet Group B delivered a uniformly solid experience with nobody dissatisfied, while Group A left two students deeply unhappy. The mean rewards the ceiling-hugging distribution and punishes the consistent one. Reading the distributions instead of the averages reverses the naive conclusion and points attention to exactly the two students in Group A who actually need follow-up.

Limitations and honest caveats

A careful reader should hold several caveats:

  • Skew is not bias by itself. A ceilinged distribution can still be valid for what it measures; the problem is the statistics applied to it, not the existence of high ratings. The fix is better summarisation, not discarding the data.
  • The 43% figure is standard-dependent. It reflects one specific norm-referenced rule on one dataset; the precise number will vary with the cut-off and population. The robust claim is directional: norm-referenced cut-offs on skewed SET data mislabel many competent instructors, not that 43% is universal.
  • Public-rating data has its own selection issues. Uttl and Smibert used publicly posted evaluations, which are not a random sample of all teaching; the institutional-survey distributions may differ in level while sharing the skew. The corroborating institutional and cross-national evidence is what lets us generalise the shape, even if exact percentages do not transfer.
  • Non-parametric does not mean assumption-free. Medians, ranks, and distributional displays solve some problems but introduce their own (ties on a 5-point scale, interpretation of rank gaps). The honest position is that no single statistic rescues a ceilinged ordinal scale; transparency about the whole distribution does.

How Koji incorporates this

Koji is designed so that a ceilinged scale is never the end of the story.

  • Distribution-first reporting. Rather than headlining a one-decimal mean, Koji surfaces the full response distribution, the median, and the size of the lower tail, so a reviewer sees the shape Stark and Freishtat ask for instead of a falsely precise average. The presence and size of a negative tail is highlighted, because that tail is where the actionable signal lives.
  • Beyond the Likert number. The deeper mitigation is not statistical but methodological: Koji's AI-moderated conversational interviews probe past the ceiling. When a student selects the top point on a scale question, Koji can follow up with an open_ended prompt that asks what specifically worked and what did not, recovering discrimination that the saturated scale cannot. This turns "everyone says 5" into differentiated, codable feedback.
  • Structured question variety to spread the response. Because agree-disagree and broad satisfaction items ceiling fastest, Koji supports item types that resist it — ranking and best-worst style choices that force differentiation, single_choice and multiple_choice behaviour items about concrete classroom practices, and yes_no checks — so not every signal is funnelled through one saturated 5-point scale.
  • Bias-aware comparison. Koji frames cross-instructor and cross-course comparison with the distribution and respondent count visible, and avoids presenting small mean gaps near the ceiling as rankings, which is designed to mitigate the misclassification Uttl and Smibert describe. It does not claim to remove ceiling effects from a 5-point scale — no instrument can — but to stop them from being read as if they were not there.

The same reasoning carries to commercial work: Koji's core research platform at koji.so applies the same conversational, distribution-aware approach to product and customer satisfaction surveys, where ceilinged CSAT and top-box scores mislead teams in exactly the way ceilinged SET misleads QA committees.

Related Resources

References

  • Uttl, B., & Smibert, D. (2017). Student evaluations of teaching: teaching quantitative courses can be hazardous to one's career. PeerJ, 5, e3299. https://doi.org/10.7717/peerj.3299
  • Ho, A. D., & Yu, C. C. (2015). Descriptive statistics for modern test score distributions: Skewness, kurtosis, discreteness, and ceiling effects. Educational and Psychological Measurement, 75(3), 365-388. https://doi.org/10.1177/0013164414548576
  • Stark, P. B., & Freishtat, R. (2014). An evaluation of course evaluations. ScienceOpen Research. https://doi.org/10.14293/S2199-1006.1.SOR-EDU.AOFRQA.v1
  • Uttl, B., Smibert, D., Santos, M. C., & Cnudde, K. (2022). Spain is not different: teaching quantitative courses can also be hazardous to one's career (at least in undergraduate courses). PeerJ, 10, e13456. https://doi.org/10.7717/peerj.13456

Related articles

analysis-reporting

Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees

The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.

analysis-reporting

The 4.2 vs 4.4 Trap: Why Small Differences in Course-Evaluation Means Are Usually Noise

Faculty and administrators routinely read meaning into tiny gaps between course-evaluation means. Boysen (2015) and Boysen et al. (2014) show this happens even when confidence intervals say the difference is nothing — and that warnings barely help. How to report uncertainty honestly.

analysis-reporting

An Evaluation of Course Evaluations: What the Berkeley Statisticians Concluded About Reporting Student Ratings

Two Berkeley statisticians argued that averaging student-rating scores to rank instructors for promotion and tenure should be abandoned for both statistical and substantive reasons. What Stark & Freishtat (2014) actually recommend — and how to report ratings responsibly instead.

analysis-reporting

Should You Average a 1–5 Course-Evaluation Scale at All? The Ordinal-vs-Interval Debate

Is it legitimate to compute a mean from a five-point course-evaluation scale? The fifty-year ordinal-versus-interval debate, from Jamieson to Carifio and Perla to Norman, and what it means for reporting student ratings.