Top-Box vs Mean: How to Report Course Evaluation Scores Without Throwing Away Information
Reporting the percentage of students who chose the top box feels intuitive, but collapsing a scale to favorable/unfavorable discards information. Here is what the measurement evidence says and how to report responsibly.
Koji Education Team
Product
In brief
"Percent favorable" or "top-box" reporting — collapsing a 1–5 scale into the share of students who picked the top one or two options — is popular because it is easy to read and sounds decisive. But collapsing a graded scale into two buckets is a form of dichotomization, and the measurement literature is blunt about its cost: MacCallum, Zhang, Preacher and Rucker (2002, Psychological Methods) show that dichotomizing a quantitative variable is "rarely defensible," typically reduces statistical power, shrinks observed effect sizes, and can produce misleading results. Top-box is not wrong to show — it communicates well — but it should never be the only number, and it should never be used for the fine-grained instructor comparisons that dichotomization is least equipped to support. Report the full distribution, an appropriate central-tendency or ordinal statistic, and a measure of uncertainty alongside any top-box figure.
What the research says
The core evidence is about what happens when you convert a graded measurement into a coarse one. MacCallum and colleagues (2002) systematically examined dichotomization of quantitative variables — splitting a continuous or multi-category scale into two groups (for example, "favorable" vs "not favorable," or a median split). Their conclusion is that dichotomization discards information about individual differences within each bucket, which generally attenuates correlations and effect sizes, lowers statistical power, and in some configurations can even inflate or reverse relationships, making it "rarely defensible." A student who rated a course 3 and one who rated it 1 are treated as identical under a top-box cut, even though they are three scale points apart.
This is not an isolated claim. Cohen's (1983) classic paper, "The cost of dichotomization" (Applied Psychological Measurement), quantified the loss: dichotomizing a normally distributed variable at its midpoint is equivalent to discarding a substantial fraction of the sample's information — the reduction in a correlation can be on the order of 20%. The two papers together form the standard citation base for the position that turning graded data into binary favorable/unfavorable categories throws away hard-won measurement precision.
There is a countervailing consideration from the ordinal-versus-interval debate. Likert responses are ordinal, so computing a mean also involves an assumption (that the distance from 1→2 equals 4→5) that some methodologists reject. This is a genuine tension: the mean makes an interval assumption top-box avoids, while top-box makes a coarsening sacrifice the mean avoids. The resolution the reporting literature converges on is not "pick one" but "show the distribution" — the full frequency of each scale point — because the distribution makes no collapsing assumption at all and lets the reader see skew, bimodality, and ceiling effects that both a single mean and a single top-box percentage hide.
Finally, top-box-style metrics have a well-known commercial cousin — the Net Promoter Score, which collapses an 11-point scale into promoters, passives and detractors. NPS is widely used precisely because it is simple, and widely criticised in the measurement literature for exactly the dichotomization reasons above: it discards information and is unstable in small samples.
Why it matters for course evaluation in practice
The reporting format is not a cosmetic choice; it changes what decisions the data can bear.
- Small classes. In a seminar of twelve students, a top-box percentage swings wildly: one student moving from "agree" to "strongly agree" can shift "percent favorable" by 8 points. Dichotomization amplifies sampling noise in exactly the small-N settings where course evaluation is most fragile.
- Instructor comparison and personnel decisions. Because dichotomization attenuates real differences and adds noise, ranking instructors by top-box percentage is even less defensible than ranking by means — and ranking by means is already discouraged for small samples. Coarse metrics give false precision to high-stakes comparisons.
- Ceiling effects hide. Evaluation distributions are typically left-skewed and bunched near the top. A top-box figure of "82% favorable" can conceal whether the other 18% were neutral or genuinely dissatisfied — a distinction that matters enormously for action.
- Communication vs inference. Top-box genuinely communicates well to non-technical audiences (students, external reviewers). The right move is to use it as a headline for communication while basing inference and decisions on the full distribution plus uncertainty.
Limitations and honest caveats
Intellectual honesty requires acknowledging the case for top-box and the limits of the critique.
- The mean is not innocent. Criticising top-box for its assumptions while defaulting to the arithmetic mean is inconsistent, because averaging ordinal Likert data makes its own contested interval assumption. Neither statistic is assumption-free.
- Dichotomization is sometimes appropriate. When a genuine threshold exists — for example, a quality-assurance policy that flags any course where more than 30% of students disagree that assessment criteria were clear — a categorical cut reflects a real decision rule, not an arbitrary split. The MacCallum critique targets arbitrary dichotomization of continuous relationships, not principled, pre-registered thresholds.
- Audience matters. For lay audiences, a full frequency distribution can be harder to parse than a single favorable percentage. Responsible reporting balances fidelity against comprehensibility rather than maximising either.
- Effect-size attenuation is context-dependent. The exact information loss from dichotomization depends on the underlying distribution and cut point; the 20%-loss figure is illustrative for a midpoint split of a normal variable, not a universal constant.
The defensible position is therefore nuanced: top-box is a legitimate communication device and a legitimate threshold flag, but a poor primary analytic statistic — and it must be accompanied by the full distribution and an honest indicator of uncertainty.
How Koji incorporates this
Koji for Education is designed so that reporting choices do not quietly destroy information:
- Full-distribution reporting by default. For every
scaleandsingle_choiceitem, Koji surfaces the complete frequency distribution — not just a mean or a top-box percentage — so ceiling effects, bimodality and the size of the dissatisfied tail are visible, honouring the "show the distribution" resolution to the mean-vs-top-box debate. - Uncertainty is shown alongside point estimates. Because dichotomization is most misleading in small classes, Koji is built to accompany summary figures with sample size and interval-style uncertainty, discouraging over-reading a swingy top-box number in a twelve-person seminar.
- Thresholds are explicit, not implicit. Where an institution genuinely wants a categorical flag (for example, "flag if >X% disagree that feedback was timely"), Koji lets that be configured as a stated decision rule — the principled dichotomization the measurement literature permits — rather than smuggling a top-box cut in as the only headline.
- Bias-aware, comparison-cautious reporting. Koji is designed to discourage fragile instructor rankings built on coarse metrics, foregrounding distributions and confidence rather than a single league-table percentage.
- Open-text triangulation. AI-moderated interviews and automatic thematic analysis explain why a distribution looks the way it does, so a modest top-box figure is interpreted against students' actual reasons rather than treated as a bare verdict.
As always these are framed as mechanisms designed to support responsible reporting, not to guarantee correct decisions — the statistic still has to be read in context. Koji's core research platform at koji.so applies the same distribution-first reporting philosophy to product and customer research, where top-box and NPS-style metrics carry the same trade-offs.
A reporting template that keeps the information
A defensible course-evaluation report layers three views of the same item rather than choosing one. First, the full frequency distribution — the count or percentage at each scale point — which reveals skew, bimodality and the size of the dissatisfied tail without any collapsing assumption. Second, a summary statistic with uncertainty: for ordinal data the median or an appropriate location estimate, accompanied by the sample size and an interval so readers can see how much the number could plausibly move. Third, an optional communication headline — the top-box or percent-favorable figure — clearly labelled as a summary for lay audiences and never used as the basis for ranking. Presenting all three costs little space and prevents the two failure modes that dominate practice: over-reading a swingy favorable percentage in a small class, and hiding a meaningful minority of dissatisfied students behind an impressive-looking headline. For any comparison that carries consequences for an individual, add the explicit caveat that small samples and coarse metrics cannot support fine distinctions between instructors. The same template travels across an institution: when every department reports the distribution plus uncertainty plus a labelled headline, cross-unit comparison becomes honest, because reviewers see the shape of the data rather than a single number engineered to look favourable.
Related Resources
- Should you average Likert data? The ordinal-interval debate
- Numeric values and rating-scale labels
- Norm-referenced vs criterion-referenced reporting
- Small mean differences and confidence intervals
- Interpreting and reporting student ratings responsibly
- Net Promoter Score in higher education evaluation
References
- MacCallum, R. C., Zhang, S., Preacher, K. J., & Rucker, D. D. (2002). On the practice of dichotomization of quantitative variables. Psychological Methods, 7(1), 19–40. https://doi.org/10.1037/1082-989X.7.1.19
- Cohen, J. (1983). The cost of dichotomization. Applied Psychological Measurement, 7(3), 249–253. https://doi.org/10.1177/014662168300700301
- Carifio, J., & Perla, R. J. (2007). Ten common misunderstandings, misconceptions, persistent myths and urban legends about Likert scales and Likert response formats and their antidotes. Journal of Social Sciences, 3(3), 106–116. https://doi.org/10.3844/jssp.2007.106.116
- Fisher, N. I., & Kordupleski, R. E. (2019). Good and bad market research: A critical review of Net Promoter Score. Applied Stochastic Models in Business and Industry, 35(1), 138–151. https://doi.org/10.1002/asmb.2417
Related articles
Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees
The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.
Should You Use Net Promoter Score for Courses? The "Would You Recommend" Question in Higher Education
Net Promoter Score is migrating from customer experience into student feedback. What Reichheld (2003) actually claimed, why Keiningham et al. (2007) failed to replicate its superiority, and whether a single recommend-question belongs in course evaluation.
The 4.2 vs 4.4 Trap: Why Small Differences in Course-Evaluation Means Are Usually Noise
Faculty and administrators routinely read meaning into tiny gaps between course-evaluation means. Boysen (2015) and Boysen et al. (2014) show this happens even when confidence intervals say the difference is nothing — and that warnings barely help. How to report uncertainty honestly.
Should You Average a 1–5 Course-Evaluation Scale at All? The Ordinal-vs-Interval Debate
Is it legitimate to compute a mean from a five-point course-evaluation scale? The fifty-year ordinal-versus-interval debate, from Jamieson to Carifio and Perla to Norman, and what it means for reporting student ratings.