Why Averaging Likert Scores Misleads in Course Evaluation
Reporting a course evaluation as "4.2 out of 5" treats ordinal data as if it were a measured quantity. It is not. Here is why the average is the wrong statistic, what the methodology literature says, and what to report instead.
Koji for Education
Research & Editorial Team · June 1, 2026
The short answer: A Likert response ("strongly disagree" to "strongly agree") is ordinal — the categories have a rank order, but the distances between them are not known to be equal. Averaging ordinal codes into a mean like "4.2 out of 5" assigns arithmetic to labels that do not support it, then ranks instructors and programmes on differences that are often smaller than the measurement noise. The defensible alternatives are to report distributions, use rank-appropriate statistics, and — most importantly — collect feedback rich enough that you do not have to lean on a single decimal to make decisions.
The error hiding in plain sight
Almost every legacy course-evaluation system does the same thing: it asks students to tick a box from 1 to 5, converts the ticks to numbers, and reports the mean. A department then compares "4.2" against "3.9", or against a faculty benchmark of "4.0", and draws conclusions. The move feels obviously fine — we average numbers all the time. But the numbers here are not measurements. They are codes standing in for ordered categories, and the arithmetic mean quietly assumes something the data never established.
Why ordinal data resists the mean
The clearest statement of the problem comes from the measurement literature. As Jamieson (2004) summarised, Likert scales sit at the ordinal level of measurement: the response options can be ranked, but "the intervals between values cannot be presumed equal." The psychological distance from strongly disagree to disagree is not guaranteed to equal the distance from agree to strongly agree. Because the intervals are unknown, the mean and standard deviation "lack validity" as a basis for establishing rank orders, and ordinal data should generally be analysed with non-parametric methods (Jamieson, 2004, Medical Education 38:1217–1218).
Put concretely: if you cannot assume the gap between a 1 and a 2 is the same "size" as the gap between a 4 and a 5, then adding the codes and dividing produces a figure whose units are undefined. "4.2" is not 4.2 of anything. It is a summary that looks quantitative but is built on an unjustified interval assumption.
False precision: the second-decimal illusion
The averaging habit compounds into a precision illusion. Two instructors separated by "4.21 vs 4.08" look meaningfully different on a dashboard. But with the typical class sizes and response counts in higher education, that difference is usually well inside the range of sampling noise — it could flip with two or three different responses. Berkeley statistician Philip Stark and colleagues have made this case forcefully: in their analysis of more than 23,000 evaluations, student ratings were more sensitive to students' gender bias and grade expectations than to teaching effectiveness, and small differences in averages do not reliably indicate real differences in teaching (Boring, Ottoboni & Stark, 2016, ScienceOpen Research). Ranking staff to the second decimal place dresses up noise as signal.
And the average is not even measuring what we think
Even if the arithmetic were sound, the construct is shaky. The meta-analysis by Uttl, White and Gonzalez (2017) re-examined the multisection studies long cited as proof that highly-rated instructors produce more learning, and found that SET ratings explain at most about 1% of the variability in student learning — effectively, the ratings and learning are unrelated (Uttl et al., 2017, Studies in Educational Evaluation 54:22–42). So the headline average is an interval-assuming summary of an ordinal scale that barely correlates with the outcome it is supposed to proxy. Three problems stacked on top of one another.
The strongest counterargument — taken seriously
"Plenty of respected research treats Likert data as interval and gets sensible results — isn't the ordinal objection pedantic?" This is a fair and genuinely debated point. With many items combined into a multi-item scale, roughly symmetric distributions, and large samples, parametric statistics are often robust, and simulation studies show means can behave reasonably. We should not pretend the methodological community is unanimous that means are always forbidden.
But notice how little of that applies to course evaluation as practised. Institutions rarely combine many items into a validated composite; they report single-item means per question. Distributions of teaching ratings are typically skewed and ceiling-bunched, not symmetric. Response counts per course are often small. And the stakes — promotion, renewal — are high, which is precisely when false precision does real damage. The defensible position is therefore not "means are always invalid" but "means are the wrong tool for this use: single ordinal items, skewed small samples, high-stakes ranking." The robustness arguments that rescue the mean in well-designed survey research simply do not hold here.
What to report instead
If the mean is the wrong summary, what is right?
- Report the full distribution. Show how many students chose each category. A course where ratings split between "excellent" and "poor" tells a completely different story from one where everyone chose "good" — yet both can average to the same number.
- Use the median and mode. These are appropriate for ordinal data and far more honest about a single decimal of apparent difference.
- Use rank-appropriate tests when comparing groups, rather than t-tests on category codes.
- Stop ranking to the decimal. If a difference can flip with two responses, it is not a ranking — it is noise with a number attached.
- Shift weight to qualitative feedback, which carries the diagnostic information a number cannot: why a course worked or did not.
Where Koji fits
The deeper fix is to stop forcing a rich human judgement through a five-point funnel in the first place. Koji for Education runs AI-moderated conversational interviews that ask students to explain their experience, then follow up — capturing the reasoning a Likert tick discards. Where scales are genuinely useful, Koji still supports them: its six structured question types include scale, ranking, single-choice, multiple-choice, yes/no and open-ended, so you can use the right format for each question rather than flattening everything to 1–5.
Crucially, Koji's automatic thematic analysis turns large volumes of open-text responses into structured, reportable themes, so qualitative feedback becomes something you can act on at scale instead of an unread free-text box. Its quality scoring flags low-effort or non-substantive responses, and programme- and institution-level reporting lets you see distributions and themes rather than a single contested average — all on a GDPR/AVG-compliant footing. To be precise: this does not magically make ordinal data interval; it removes the dependence on a misleading average by giving you better evidence to begin with.
Research and insights teams who run customer or user studies will recognise the same engine behind the main Koji platform, which applies the identical interview approach outside the classroom.
A worked example: same average, opposite stories
Consider two seminars of twenty students each. In Course A, ten students rate the teaching "excellent" (5) and ten rate it "poor" (1). In Course B, all twenty rate it "good" (4) — wait, let us make the averages match exactly: in Course A, ten students give a 5 and ten give a 3, for a mean of 4.0; in Course B, every student gives a 4, also a mean of 4.0. On the dashboard these courses are identical: both "4.0 out of 5." In reality they could hardly be more different. Course A is polarising — half the room is thriving and half is alienated, which usually signals a specific, fixable problem (pace, prerequisites, a divisive teaching choice). Course B is uniformly fine — no crisis, no standout. The average erases precisely the information a programme director needs to act. The median fares a little better (4 vs 4 here, but it would separate many other cases), yet even the median cannot tell you why Course A split. Only the distribution flags the polarisation, and only the open-text reasoning explains it. This is the everyday cost of the averaging habit: not an exotic statistical error, but the routine destruction of the most actionable signal in the data.
The bottom line
"4.2 out of 5" is a comfortable fiction: it converts ordered labels into a quantity, ranks people on differences smaller than the noise, and proxies an outcome it barely predicts. The methodology literature has flagged the ordinal problem for two decades, and the SET-specific evidence shows the practical damage. Report distributions and medians, refuse to rank to the decimal, and rebuild evaluation around feedback rich enough that you never had to trust the average in the first place.