Should You Average a 1–5 Course-Evaluation Scale at All? The Ordinal-vs-Interval Debate
Is it legitimate to compute a mean from a five-point course-evaluation scale? The fifty-year ordinal-versus-interval debate, from Jamieson to Carifio and Perla to Norman, and what it means for reporting student ratings.
Koji Education Team
Product
In short: Reporting a course as "4.2 out of 5" treats an ordinal rating scale as if its intervals were equal — the gap between "agree" and "strongly agree" is assumed identical to the gap between "neutral" and "agree." Jamieson (2004) warned that averaging ordinal data is illegitimate; Carifio and Perla (2008) and Norman (2010) replied that, for summated multi-item scales analysed in the aggregate, parametric statistics are robust enough to be safe. The defensible position for course evaluation is narrow but practical: a mean of a well-constructed multi-item scale is usually acceptable for analysis, but a mean of a single ordinal item, reported to one decimal place and used to rank instructors, is exactly the case the critics are right about. Report distributions, not just means.
What the research says
The question of whether you may legitimately calculate a mean from a Likert response is one of the longest-running methodological disputes in the social sciences, and it lands directly on the desk of anyone who reports course evaluations. The dispute turns on levels of measurement. A five-point scale from "strongly disagree" to "strongly agree" is, on its face, ordinal: the categories are ordered, but nothing guarantees the psychological distance between adjacent points is equal. The arithmetic mean, by contrast, is an interval statistic — it only carries meaning if a one-unit step means the same thing everywhere on the scale. Averaging ordinal data therefore imports an assumption the scale itself does not justify.
Susan Jamieson crystallised the objection in a much-cited 2004 editorial in Medical Education, "Likert scales: how to (ab)use them." Her argument was direct: Likert items yield ordinal data, the intervals between points cannot be presumed equal, and so the appropriate descriptive statistics are the median and mode, with non-parametric tests (Mann–Whitney, Kruskal–Wallis, chi-square) for inference. Computing a mean and standard deviation, and then building rankings on them, risks "coming to the wrong conclusion" because the numbers attach a false precision to ordered categories.
The counter-position arrived most forcefully from James Carifio and Rocco Perla. In "Resolving the 50-year debate around using and misusing Likert scales" (2008, Medical Education) and related papers, they drew a sharp distinction that the debate had blurred: the difference between a single Likert item and a Likert scale — a summated composite of several items measuring one underlying construct. A single item is genuinely ordinal and coarse. But the sum or average of many items, they argued, behaves much more like an interval variable, has more possible values, and is what Likert himself (1932) prescribed analysing. For such composites, parametric analysis is appropriate.
Geoff Norman supplied the empirical close in "Likert scales, levels of measurement and the 'laws' of statistics" (2010, Advances in Health Sciences Education). Using both real and simulated data, Norman showed that parametric procedures — Pearson correlation, t-tests, regression, ANOVA — are remarkably robust to violations of the interval and normality assumptions, even with small samples and badly skewed distributions. His blunt conclusion: parametric statistics can be used with Likert data without fear of arriving at the wrong answer, and the level-of-measurement prohibition is, in practice, overstated. The contemporary methodological consensus sits closer to Norman than to a purist reading of Jamieson — but with an important rider that the course-evaluation context makes vital.
Why it matters for course evaluation in practice
The rider is this: the robustness results that rescue the mean apply most cleanly to multi-item summated scales analysed across reasonable samples. Course evaluation routinely violates both conditions. Institutions frequently report a single global item — "Overall, this was an excellent course" — as a one-decimal mean, and they do so for small classes where a seminar of twelve students produces a "4.3" built on a handful of ordered ticks. That is the precise scenario in which Jamieson's caution bites hardest: a single ordinal item, low n, treated as if "4.3" and "4.5" were comparable interval quantities, and then used to rank colleagues or trigger review.
Two practical errors follow from ignoring the debate. The first is false precision: presenting "4.18" implies a resolution the underlying ordered categories cannot support, inviting readers to distinguish instructors who are statistically indistinguishable. The second is distributional blindness: a mean conceals shape. A bimodal class — half the students delighted, half alienated — can produce the same "3.0" as a class where everyone was lukewarm, yet these describe completely different teaching situations and demand different responses. The mean is not wrong so much as insufficient; it discards exactly the information an evaluation committee needs.
The defensible reporting practice that follows from the literature is therefore not "never average" but report the distribution alongside any summary: show the full frequency of responses, report the median and the spread, and reserve means for multi-item composites where they are best justified — never as the sole basis for a consequential comparison.
A simple operating rule captures the defensible middle ground. Use the mean only where the assumptions are least strained — multi-item composites, adequate sample sizes — and always publish it next to the full response distribution and the median. Never rank colleagues by differences in the second decimal place of a single-item average, and never let a small-class "4.6" stand unqualified beside a large-class "4.4" as though the comparison were sound. The mean is a summary, not a verdict; treat it as one view of the data among several, and let the distribution carry the weight the average cannot.
Limitations and honest caveats
The debate has not been settled by fiat, and intellectual honesty requires noting where each side overreaches. Norman's robustness demonstrations are real, but "robust" is a statement about Type I error rates under specific conditions, not a licence to treat any ordinal number as interval in any context; with extreme skew, ceiling effects, and very small n — all common in course evaluation — robustness can erode. Conversely, the purist non-parametric position can be impractical and can discard genuine information, and Carifio and Perla are right that a well-built summated scale is not the same animal as a single tick-box.
There is also a deeper point the arithmetic argument cannot reach: even a perfectly interval-scaled, normally distributed mean can be substantively meaningless if the item is biased (see the bias literature on gender, accent, and difficulty) or if respondents are unrepresentative. Level of measurement is necessary, not sufficient, for a valid inference. Finally, the whole dispute concerns how to summarise numbers, and says nothing about whether the numbers measure teaching quality in the first place — a question on which the validity literature is far more sobering than any choice between mean and median.
How Koji incorporates this
Koji for Education treats the headline mean as the start of an analysis, never the conclusion. Its design reflects the resolved core of this debate — that distribution and meaning matter more than a single averaged digit:
- Distribution-first reporting. Koji is designed to surface the full response distribution, median, and spread for every scale item, not just a mean — so a bimodal "3.0" is visibly different from a uniformly lukewarm "3.0," and false precision on small classes is harder to fall for.
- Beyond the number, to the reason. Where the ordinal debate is about how to summarise a tick, Koji's AI-moderated conversational interviews ask why a student chose it. Structured items (scale, single_choice, ranking) are paired with open_ended probes and automatic thematic analysis, so interpretation rests on explained reasons rather than a contested average. Koji's core research platform at koji.so applies the same engine to product and customer research, where the single-metric trap is identical.
- Best-worst and ranking alternatives. For prioritisation questions where five-point means compress everything into "4-point-something," Koji supports ranking and choice formats that sidestep the interval assumption entirely.
- Small-sample guardrails. Koji's reporting is designed to flag when a class is too small for a mean difference to be meaningful, discouraging the rank-by-decimal habit the critics warn against.
- Multi-item construct support. When a construct is measured with several items, Koji supports composite reporting — the case where averaging is most defensible — while keeping single global items visibly separate.
Koji is designed to mitigate the false-precision and distributional-blindness failures the ordinal debate exposes; it does not pretend that any reporting choice can turn a biased or unrepresentative item into a valid one.
Related Resources
- The 4.2 vs 4.4 Trap: Why Small Differences in Course-Evaluation Means Are Usually Noise
- How Many Scale Points Should a Course-Evaluation Question Have?
- Do the Numbers on Your Rating Scale Change the Score? The Evidence on Numeric Anchors
- Cronbach's Alpha and Course-Evaluation Reliability: What It Does and Doesn't Tell You
- Interpreting and Reporting Student Ratings Responsibly
References
- Jamieson, S. (2004). Likert scales: how to (ab)use them. Medical Education, 38(12), 1217–1218. https://doi.org/10.1111/j.1365-2929.2004.02012.x
- Carifio, J., & Perla, R. (2008). Resolving the 50-year debate around using and misusing Likert scales. Medical Education, 42(12), 1150–1152. https://doi.org/10.1111/j.1365-2923.2008.03172.x
- Norman, G. (2010). Likert scales, levels of measurement and the "laws" of statistics. Advances in Health Sciences Education, 15(5), 625–632. https://doi.org/10.1007/s10459-010-9222-y
- Likert, R. (1932). A technique for the measurement of attitudes. Archives of Psychology, 22(140), 1–55.
Related articles
Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees
The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.
Cronbach's Alpha and Course-Evaluation Reliability: What It Does and Doesn't Tell You
A high Cronbach's alpha on your course-evaluation instrument is widely read as proof of a "reliable" survey. The psychometric literature says that reading is largely wrong. What alpha actually measures, where it misleads, and what to report instead.
The 4.2 vs 4.4 Trap: Why Small Differences in Course-Evaluation Means Are Usually Noise
Faculty and administrators routinely read meaning into tiny gaps between course-evaluation means. Boysen (2015) and Boysen et al. (2014) show this happens even when confidence intervals say the difference is nothing — and that warnings barely help. How to report uncertainty honestly.
Some Raters Are Harsh, Some Items Are Hard: Why Rasch and Many-Facet Models Beat Averaging Course-Evaluation Scores
Averaging Likert scores treats every student as an equally calibrated measuring instrument and every item as equally hard. Rasch and Many-Facet Rasch models do not. We explain the method, the evidence that rater severity and item difficulty distort raw means, the limitations, and how Koji applies the same logic.