Your Course Evaluation Treats a Likert Scale Like a Ruler. Measurement Theory Says It Is Not One.
A five-point rating is a rank, not a measured distance. Rasch measurement and item response theory are the established methods for turning ordinal course-evaluation ratings into a scale you can legitimately average, compare, and trust.
Koji Education Team
Product ·
The number on a course-evaluation form is a rank dressed up as a measurement. When a student selects "4" on a five-point scale, that 4 tells you the response outranks a 3 and trails a 5. It does not tell you that the psychological distance from 3 to 4 is the same as the distance from 4 to 5. Yet the moment an institution averages those numbers, ranks lecturers by the mean, and treats a 0.2-point gap as meaningful, it has quietly assumed exactly that equal-interval property — an assumption almost no university has ever tested. Rasch measurement and item response theory (IRT) are the established psychometric methods for checking that assumption and for converting ordinal ratings into a scale you can defensibly add, average, and compare.
This is not a pedantic footnote. It changes which lecturers land "below threshold" and which course changes register as real improvements.
Ordinal numbers, interval arithmetic
The problem is a category error about levels of measurement. S. S. Stevens' classic 1946 taxonomy distinguishes nominal, ordinal, interval, and ratio scales, and the arithmetic each one permits. Ordinal data support order statistics — medians, ranks, frequencies — but not the addition and averaging that assume equal spacing between points.
A Likert-type response scale is ordinal. As Susan Jamieson argued in her widely cited note "Likert scales: how to (ab)use them" (Medical Education, 2004, 38:1217–1218), the intervals between "strongly disagree", "disagree", "neutral" and so on cannot be presumed equal, so means and standard deviations computed on them rest on an untested assumption. Cohen and colleagues put it bluntly: it is illegitimate to infer that the intensity of feeling between "strongly disagree" and "disagree" equals the intensity between any other adjacent pair.
The consequence for course evaluation is direct. Every ranked list of instructors, every "action required below 3.5" rule, every year-on-year comparison of a mean silently treats an ordinal scale as if it were a tape measure.
What a Rasch model actually does
Rasch measurement, named for the Danish statistician Georg Rasch, and the broader family of IRT models offer a principled way out. Rather than assuming the raw scores are interval, they estimate a latent trait — how positively a student is disposed toward the course — on a genuine interval metric (logits), from the pattern of responses across many items and many students.
Three things fall out of the model that a mean cannot give you:
- Item difficulty, or endorsability. Some statements are easy to agree with ("the lecturer was audible") and some are hard ("the course changed how I think"). Rasch places items and persons on the same scale, so a "4" on an easy item is not scored as if it were a "4" on a demanding one.
- A unidimensionality test. The model checks whether your ten questions actually measure one underlying construct, or several tangled together. If they do not, averaging them into a single course score is meaningless — a point that connects directly to the halo effect and to the jingle-jangle problem of vaguely defined constructs.
- Fit statistics and invariance. Rasch flags items and respondents that behave erratically, and it lets you test whether the scale works the same way for different student groups — the same measurement-invariance question that determines whether any cross-group comparison is even valid.
Bond and Fox's standard text, Applying the Rasch Model: Fundamental Measurement in the Human Sciences, makes the core claim plainly: constructing an interval scale is something you do and verify, not something you assume by printing numbers on a form.
What averaging hides
Consider two lecturers with identical mean scores of 4.1. In one class the responses cluster tightly at 4; in the other they split between enthusiastic 5s and a substantial minority of 2s. The mean erases the difference, but the Rasch person-distribution and the dispersion do not. One course is working for everyone; the other is a bimodal warning sign.
Or consider an item everyone endorses. In raw scoring it pushes every course's average up. In a Rasch framework it is recognised as carrying almost no discriminating information — it cannot separate a good course from a poor one, because both max it out. This is the measurement-theory sibling of the ceiling-effect problem.
"But surely this is overkill" — the strongest counterargument
The most serious rebuttal comes from Geoff Norman's frequently cited paper "Likert scales, levels of measurement and the 'laws' of statistics" (Advances in Health Sciences Education, 2010, 15:625–632). Norman marshals evidence that parametric statistics are remarkably robust to violations of the interval assumption: t-tests and ANOVA on Likert data usually reach the same conclusions as their non-parametric counterparts, even with skewed, small samples. On this view, agonising over ordinality is a distraction.
This deserves a fair hearing, and it is partly right. If your only question is "did the mean move", robustness studies suggest you will rarely be badly misled. But three points survive the rebuttal:
- Robustness of a test is not the same as validity of a difference. Norman defends the inference procedure, not the interpretation that a 0.2-point gap represents a fixed, comparable quantity of teaching quality. High-stakes personnel and tenure decisions lean on exactly that interpretation.
- The payoff of Rasch is diagnostic, not just the mean. Unidimensionality checks, item fit, and differential item functioning are things a t-test never provides, and they are where course-evaluation instruments most often fail.
- Small differences and small classes are where robustness thins. For a class of twelve, the reassurance evaporates.
The honest position is not "always run a Rasch analysis". It is: if you are going to rank people and trigger consequences on fractional differences, you owe them a scale that has been shown to be interval, not one merely assumed to be.
Where Koji fits
Koji does not claim that a conversation repeals measurement theory. Scale items are still ordinal, and if you want defensible interval scores you still model them. What Koji changes is the raw material the numbers are supposed to summarise.
A logit tells you where a student sits on a latent scale; it never tells you what a "3" meant to them. Koji's AI-moderated conversational interviews probe that meaning directly — why a rating landed where it did — and its automatic thematic analysis turns thousands of open responses into structured, weighted themes rather than a single contested average. Its six structured question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) let you keep the scale items you need for modelling while capturing the qualitative evidence that gives them meaning. And because Koji reports distribution and themes, not just a mean, the bimodal course that averaging would hide stays visible.
Legacy SET tools were built to compute and print that single average. The measurement-theory critique is, in the end, a critique of treating a printed mean as the finding. Koji treats it as the starting question.
Teams outside the registrar's office hit the same ordinal trap: product and user-research groups routinely average satisfaction scales that were never interval. The same AI interview engine powers koji.so for that general research work — worth a look if your organisation over-trusts a mean elsewhere too.
If your evaluation system ranks lecturers on fractions of a point, it is worth asking what those fractions actually measure. See how Koji for Education approaches course evaluation.
FAQ
Is a Likert scale ordinal or interval? A single Likert-type item is ordinal: its categories have a rank order, but the distances between them are not guaranteed to be equal. Summed multi-item scales are sometimes treated as approximately interval, but that is an assumption to test — with Rasch or IRT — not a given.
What is the difference between Rasch analysis and simply averaging scores? Averaging assumes every point on the scale is equally spaced and every item equally informative. Rasch analysis estimates an interval-level latent score, models how easy each item is to endorse, and tests whether the items measure a single construct — none of which averaging can do.
Does the ordinal-versus-interval debate actually change course-evaluation decisions? It can. When decisions hinge on small differences in means, on rankings, or on fixed cut-off scores, the untested interval assumption is doing real work. For descriptive, low-stakes feedback the practical risk is smaller.
Is running a Rasch model realistic for a university? It is more feasible than it sounds — established software and validated instruments exist — but it is not free. A pragmatic path is to Rasch-validate your core instrument once, then monitor it, rather than modelling every module every term.
Does Koji use Rasch modelling? Koji focuses on capturing richer evidence — conversational depth, thematic analysis, dispersion — rather than distilling everything to one average. Scale items it collects can be modelled with Rasch or IRT; the platform's distinctive contribution is the qualitative meaning behind the numbers, not a new scoring formula.