New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
analysis-reporting10 min read

Ordinal Regression for Course-Evaluation Data: Why Cumulative-Link Models Beat Averaging Likert Scores

Averaging Likert-scale course-evaluation items treats ordered categories as if they were equal-interval numbers. What the research says about the errors this causes, and why cumulative-link (ordinal) regression is the defensible analysis method.

Koji Education Team

Product

The short answer

Almost everyone reports course evaluations by averaging Likert items — treating "strongly agree = 5, agree = 4, ..." as if the distance from 4 to 5 equals the distance from 2 to 3. It rarely does. Ordinal (cumulative-link) regression models the response for what it is — an ordered category — and estimates the probability of each rating rather than pretending the numbers are a continuous measurement. When the underlying distribution is skewed or the category spacing is uneven (both are normal for evaluation data), averaging can not only distort effect sizes but reverse the direction of a difference. If a score influences a personnel or programme decision, the ordinal model is the defensible choice.

BLUF: Averaging Likert course-evaluation items assumes equal spacing between response categories, which is usually false. Liddell and Kruschke (2018) show that metric (mean/OLS) analysis of ordinal data can inflate error rates and even flip the sign of an effect. Cumulative-link ordinal regression models the ordered categories directly and avoids these artefacts. Koji is designed to support this by preserving raw category distributions and offering ordinal-appropriate summaries, so instructor and course comparisons rest on the response distribution rather than a fragile mean.

What the research says

The most direct evidence is Torrin Liddell and John Kruschke''s Journal of Experimental Social Psychology article, provocatively titled "Analyzing ordinal data with metric models: What could possibly go wrong?" (2018). They first document the scale of the problem: in a survey of empirical articles, essentially 100% of the studies that used ordinal (Likert-type) data analysed it with metric models — means, t-tests, ANOVA, OLS regression — which assume interval-level, normally distributed data. They then run systematic simulations comparing metric models against the ordinal alternative (an ordered-probit / cumulative-link model). The findings are stark: treating ordinal data as metric can produce inflated false-alarm (Type I) rates, reduced power (Type II errors), and — most alarmingly — inversions, where the metric model reports a difference in the opposite direction from the true ordinal difference. These inversions arise precisely when distributions are skewed or when groups differ in variance or in how they use the endpoints of the scale — conditions that are the norm, not the exception, in ceiling-heavy evaluation data.

Paul-Christian Bürkner and Matti Vuorre, in Advances in Methods and Practices in Psychological Science (2019), provide the constructive counterpart: a tutorial on ordinal regression models. They lay out the family — the cumulative model (the workhorse for Likert data), plus adjacent-category, sequential, and category-specific-effects variants — and argue that these models are both more appropriate and more informative than metric approximations, because they estimate the probability of each response category and can model how predictors shift the whole distribution, including changes in variability rather than just the mean. Crucially, the cumulative model rests on a latent-variable interpretation: the observed rating is a coarse, thresholded view of a continuous latent attitude, and the model estimates both the regression effects and the (unequal) thresholds between categories — the very spacing that averaging assumes away.

Spencer Harpe''s widely cited practitioner guide in Currents in Pharmacy Teaching and Learning (2015) offers a more measured, applied bridge. He distinguishes a single Likert item (ordinal, best summarised with medians/modes and frequencies, and analysed non-parametrically or with ordinal models) from a Likert scale — a sum or mean of many items designed to measure one construct — which under some conditions can reasonably be treated as continuous. This nuance matters: the fiercest objection to averaging applies to single items and short, heterogeneous evaluation forms; a long, validated, unidimensional scale sits on firmer ground. But most institutional course-evaluation reporting averages individual items (or even single global items), which is exactly the case the ordinal-model literature warns against.

Why it matters for course evaluation in practice

Course-evaluation data has the two features that make averaging dangerous. It is highly skewed — ratings pile up at the top of the scale (ceiling effects), so distributions are non-normal and asymmetric. And the categories are not equally spaced — the psychological gap between "good" and "excellent" is not obviously the same as between "poor" and "fair," and students use scale endpoints differently across cultures and disciplines. Under these conditions, two instructors with identical means can have very different distributions (one polarising, one uniformly middling), and a mean hides the difference. Worse, per Liddell and Kruschke, a mean can rank instructor A above instructor B when the ordinal structure says the reverse.

The stakes rise when scores feed decisions. Ranking instructors, flagging courses for review, or supplying evidence for promotion on the basis of a third-decimal-place difference in means is exactly where a metric artefact becomes an injustice. Ordinal regression also unlocks fairer adjusted comparisons: because it is a regression, you can include covariates (class size, level, discipline, elective vs required) and model nested data, estimating differences net of known confounds instead of comparing raw averages. And it reports results in an interpretable currency — the probability that a course receives each rating — which is more honest than an average that no student actually gave.

Limitations and honest caveats

Several caveats keep this from being dogma. First, the practical difference between ordinal and metric analysis is often small for symmetric, non-ceilinged data with five or more well-behaved categories; the horror cases (inversions) require skew or unequal variances. Analysts should not imply that every historical mean is wrong. Second, ordinal models are more complex to fit, explain, and communicate to non-technical committees; the cumulative model also carries the proportional-odds assumption (predictor effects are constant across thresholds), which must be checked and sometimes relaxed with partial-proportional-odds or category-specific effects. Third, Harpe''s distinction is real: validated multi-item scales may justify metric treatment, and there is a long, unresolved methodological debate (the "ordinal versus interval" question) rather than a settled verdict. Fourth, ordinal regression fixes the measurement-level problem but not the validity problems — bias, non-response, and confounding are untouched by a better link function. Finally, sample-size demands are higher for stable estimation of multiple thresholds, so very small classes gain less. The defensible claim is narrow and strong: for skewed, single-item, decision-relevant evaluation data, cumulative-link models avoid documented artefacts that averaging can introduce — not that means are always wrong.

How Koji incorporates this

Koji is designed so that the analysis method can match the data type, and so that reviewers are never forced to reduce an ordered distribution to a single fragile number.

  • Raw category distributions are preserved and shown. Koji reports the full frequency distribution for every scale item — the share choosing each category — rather than collapsing to a mean by default. This is the prerequisite for any ordinal-appropriate interpretation and lets reviewers see skew, polarisation, and ceiling effects directly.
  • Ordinal-appropriate summaries. Alongside distributions, Koji surfaces medians, top-box/percent-favourable, and category probabilities — summaries that respect the ordered-but-not-interval nature of the data — instead of privileging the arithmetic mean that the research warns against.
  • Distribution-aware comparison. When comparing courses, cohorts, or time points, Koji foregrounds shifts in the whole distribution rather than differences in means alone, which is exactly what a cumulative-link view captures — and it flags small, potentially artefactual gaps rather than ranking on them.
  • Structured question types that keep ordinality explicit. Koji''s scale questions record the ordered categories as categories, and single_choice / ranking types capture genuinely ordinal or ordinal-adjacent preferences without silently coercing them to interval numbers.
  • Triangulation with conversational open text. Because Koji''s AI-moderated interviews probe the reasons behind a rating, a skewed or polarised distribution can be interpreted rather than merely averaged — the qualitative layer explains the shape the ordinal model reveals.

None of this claims to run every institution''s inferential model for them; it is designed to support ordinal-correct analysis by preserving the distribution and offering the right summaries, so downstream comparisons do not silently inherit the equal-spacing assumption. Koji''s core research platform at koji.so applies the same distribution-first reporting to product and customer research, where averaged satisfaction scores mislead in identical ways.

Frequently asked questions

Why is averaging Likert course-evaluation scores a problem? Averaging treats ordered categories as equal-interval numbers — assuming the gap from "agree" to "strongly agree" equals the gap from "disagree" to "neutral." That assumption is usually false for evaluation data, and Liddell and Kruschke (2018) show it can inflate error rates and even reverse the direction of a difference.

What is a cumulative-link (ordinal) regression model? It models the probability of a response falling at or below each ordered category, treating the observed rating as a coarse view of a continuous latent attitude. It estimates both predictor effects and the (unequal) thresholds between categories, so it respects the ordering without assuming equal spacing.

Is it ever acceptable to treat Likert data as continuous? Sometimes. Harpe (2015) notes that a validated, unidimensional multi-item scale (a mean of many items) can reasonably be treated as continuous under some conditions. The strongest objection applies to single items and short heterogeneous forms — which is most institutional course-evaluation reporting.

Does ordinal regression fix bias in course evaluations? No. It fixes the measurement-level mismatch. Bias, non-response, and confounding are separate problems that require their own controls; a better link function does not make an invalid instrument valid.

When does the choice of method matter most? When distributions are skewed or ceilinged (typical for evaluations), when groups differ in variance or endpoint use, and when a small difference in scores drives a real decision such as ranking or promotion. Those are exactly the conditions under which averaging can produce artefacts.

Related resources

References