The Mean Hides the Mutiny: Why Spread and Bimodality Matter More Than the Average in Course Evaluations
Two courses can share an identical 3.0 average while one delighted everyone moderately and the other split the room into love and hate. The mean cannot tell them apart. Why dispersion, shape and bimodality are the signal — and why reporting only the average throws away your most actionable information.
Koji for Education
Research & Editorial Team · June 22, 2026
Bottom line up front: A mean is a single point estimate of a distribution's centre. It says nothing about spread or shape — and in course evaluation, the spread and the shape are often where the real story lives. A 3.0 average can come from a class that found the course uniformly mediocre, or from a class that split sharply into students who loved it and students who hated it. Those are completely different teaching situations demanding completely different responses, yet they produce the identical headline number. If your evaluation reports show means without distributions, you are discarding the most actionable signal you collect.
One number, many distributions
Begin with the basic statistical point that course-evaluation reporting routinely ignores. As a 2021 analysis of evaluation data puts it bluntly, evaluators "often rely solely on mean score ratings," which is problematic because "a mean score alone does not illustrate the underlying score distribution, which in turn could completely alter the meaning of the data" (On Misinterpretation of Course and Instruction Evaluation Data, 2021). The same source notes that "in nearly all instances, the interpretation of a mean score of 3.0 could be very different depending on its underlying distribution of scores."
Consider three classes on a 1–5 scale, each with a mean near 3.0:
- Class A — most students answered 3. The course is reliably, unremarkably middling. The improvement task is incremental: lift a solid-but-flat experience.
- Class B — answers split between 1 and 5, with almost nobody at 3. This is a bimodal distribution: the course is polarising. Something works brilliantly for one group and fails another — perhaps the pace suits prepared students and loses the rest, or the assessment rewards one learning style. The mean of 3.0 describes no actual student's experience.
- Class C — a long left tail: mostly 4s with a cluster of 1s. A minority had a genuinely bad time. That tail might be an accessibility failure, a scheduling clash, or a subgroup the course is leaving behind.
These demand three different interventions. The mean treats them as the same course. As the basics of descriptive statistics insist, you must examine a distribution's centre, its shape (skewness, modality) and its spread (variability, range) — not the centre alone.
Why bimodality is the most important shape to catch
A bimodal distribution "indicates students had polarising opinions of the course," and it is the pattern most likely to be flattened into invisibility by an average. It is also diagnostically rich. Polarisation often signals a course that is implicitly designed for one kind of student — and, troublingly, the split can fall along demographic or preparedness lines. The misinterpretation analysis above notes that a bimodal pattern can arise when subgroups apply different expectations "depending on whether students match the instructor demographically or not." A polarised distribution is therefore not just noise to be averaged away; it can be the first visible symptom of an equity problem or a hidden prerequisite gap. Average it into a 3.0 and the symptom vanishes from the report.
Spread is also a reliability signal
Dispersion is not only about diagnosing the course; it is about how much you should trust the number at all. A tight distribution around 4.2 means students broadly agree — the rating is a reliable summary. A wide distribution around 4.2 means students disagree, and the mean is a fragile artefact that a handful of responses could swing. This connects to the points we make about effect sizes and whether a 0.3 difference is real and about the statistics of uncertainty in small classes: the standard deviation and the response count govern how seriously any mean difference should be taken. Reporting a mean without its spread is reporting an answer without its error bars.
The arbitrators agree: report the distribution
This is not merely a statistician's preference. When an arbitrator ordered Ryerson University to reform its use of student evaluations in 2018, the remedy was specific: replace the numerical ranking system and present results "in the form of a frequency distribution together with response rates" (Inside Higher Ed, 2018). An independent labour tribunal, reasoning about fairness, reached the same conclusion as the descriptive-statistics literature: the distribution, not the mean, is the honest unit of reporting.
"But decision-makers need one number"
The strongest objection is practical. Deans, accreditation panels and dashboards want a scalar they can sort, benchmark and put in a table. A full histogram per course does not paste neatly into a comparison grid, and asking busy committees to eyeball distributions invites its own inconsistency.
Two responses. First, if the goal is a fair summary, the mean is often the wrong scalar anyway — the median resists the long tails that skew evaluation data, and reporting the proportion of responses in the top two and bottom two categories ("% favourable / % unfavourable") preserves far more of the shape than a mean while staying sortable. Second, the demand for one number is exactly the pressure that produces the misreadings above; the right response to "we need something scannable" is a small set of complementary statistics — a measure of centre, a measure of spread, and a flag for bimodality — not a single mean pretending the distribution does not exist. As we argue in why averaging Likert scores misleads, the convenience of the mean is bought with the suppression of the very information that makes feedback useful.
A second objection: isn't the distribution just noise on small classes?
For very small classes, yes — a histogram of eight responses is jumpy and over-interpretable. But this is an argument for honest uncertainty, not for retreating to the mean, which is even more misleading on small n because it hides how few people and how much disagreement produced it. The right move on small classes is to widen the interpretive bands and lean harder on the qualitative comments, not to compress everything into one precise-looking average that is in fact the least stable summary available.
How Koji surfaces shape, not just centre
Koji for Education is designed so that the distribution is never the thing you have to remember to look at — it is the default unit of analysis. Its scale-type questions retain the full response distribution rather than collapsing to a mean, and its programme- and institution-level reporting is built to present shape and spread, in line with both the descriptive-statistics literature and the Ryerson remedy. More fundamentally, Koji does not stop at the numbers. Its AI-moderated conversational interviews ask why, and its automatic thematic analysis of open-text feedback turns a polarised distribution into a readable explanation — telling you not just that a course split the room, but which two rooms it split into and why. A bimodal scale score and the themes that explain it are far more actionable together than either alone. To be precise about the claim: Koji does not invent signal that is not there; on a genuinely tiny class the uncertainty remains. What it does is stop the reporting layer from throwing away the spread, the shape and the reasons — the parts a mean silently deletes.
Product and UX teams running general research meet the identical "the average looks fine but the distribution is bimodal" problem in satisfaction data; the main Koji platform runs the same interview engine for that work.
The takeaway
The mean is a convenient lie of omission. Two courses with the same average can be having opposite experiences — one uniformly fine, one in open revolt — and only the spread and shape can tell them apart. Bimodality flags polarisation and possible inequity; dispersion governs how much you should trust the number at all; and both the descriptive-statistics literature and a labour arbitrator independently conclude that the frequency distribution, not the mean, is the honest way to report. Show the distribution, flag the bimodal cases, read the comments that explain them — and stop letting the average hide the mutiny.
Koji for Education reports the full distribution and the themes behind it, not a lone average that flattens polarised feedback into the middle. See how Koji approaches evaluation reporting.