Beyond the Likert Scale: Adaptive Comparative Judgement in Evaluation
Humans are better at judging "which of these two is better" than at assigning an absolute number. Adaptive Comparative Judgement (Pollitt 2012), built on Thurstone's law of comparative judgement, turns pairwise choices into a reliable scale — with lessons for course evaluation.
Koji Education Team
Product
In brief: People are far more reliable at saying which of two things is better than at assigning an absolute score on a scale. Adaptive Comparative Judgement (ACJ) — formalised by Pollitt (2012) and rooted in Thurstone's (1927) law of comparative judgement — exploits this: judges make a series of holistic pairwise comparisons, and a statistical model converts those choices into a reliable rank order and interval scale, often reaching reliabilities above 0.90. For course evaluation, the deeper lesson is that the response format itself is a source of error: rigid Likert numbers carry rater-severity and scale-use noise that comparative methods largely sidestep.
What the research says
Most course evaluation rests on an assumption we rarely examine: that a student (or a peer reviewer, or an accreditation panel) can take a complex object — a course, a teaching performance, a piece of student work — and compress it into an absolute number on a 1-to-5 or 1-to-7 scale, consistently, and comparably with everyone else doing the same. Nearly a century of psychophysics says this assumption is shaky.
Thurstone (1927), in "A law of comparative judgement," showed that human judgement of magnitude is far more dependable when expressed as a comparison than as an absolute rating. The comparative judgement of two objects depends lawfully on the magnitude of their difference in quality, and a series of such comparisons can be modelled to place objects on an interval scale — without anyone ever assigning a number.
Pollitt (2012), "The method of Adaptive Comparative Judgement," in Assessment in Education: Principles, Policy & Practice, modernised Thurstone's idea for educational assessment. The method developed within a research project at Goldsmiths (2004–2010): instead of marking work against a rubric, an expert is shown two pieces and simply decides which is better. Across many such judgements by many judges, a Bradley–Terry/Rasch-style model estimates a quality parameter for each item and builds a scale. ACJ adds adaptivity in scoring (not testing): later pairings are chosen to be informative — pitting items of similar estimated quality against each other — so the scale sharpens quickly. In a large study of roughly 1,000 student writing samples, Pollitt reported a robust, valid rank order and strong support from educators, who often preferred holistic comparison to rubric marking.
The reliability evidence is the headline. Comparative-judgement exercises routinely report Scale Separation Reliability above 0.90, frequently exceeding the inter-rater reliability achievable with rubric-based absolute marking of the same work. Independent scrutiny has refined the claim: Cambridge Assessment researchers (e.g., Bramley) noted that the adaptive algorithm can inflate reliability statistics if not handled carefully, and a 2021 examination in the International Journal of Technology and Design Education re-checked ACJ reliability across settings. The mature consensus is that comparative judgement is genuinely reliable and valid, while the adaptive variant needs careful design to avoid over-stating its precision.
The connection to course evaluation is conceptual but direct. The well-known many-facet Rasch treatment of student ratings (Linacre) models rater severity precisely because absolute Likert ratings are contaminated by how harshly or leniently each rater uses the scale. Comparative judgement attacks the same problem at the source: if you never ask for an absolute number, you cannot suffer from one judge calling "good" a 4 and another calling the identical thing a 6.
Why it matters for course evaluation in practice
Course evaluation uses absolute scales almost everywhere — students rating a course, peers rating a teaching observation, panels scoring a programme against criteria. Three practical implications follow from the comparative-judgement literature.
-
Format is a measurement choice, not a neutral container. The decision to use a 5-point Likert item is itself a source of error: ceiling effects, central-tendency bias, and divergent scale use all live in the format. Recognising this is the first step; it reframes "are our items well written?" into the larger "is an absolute rating even the right instrument for this judgement?"
-
Some evaluation tasks are better suited to comparison. Where experts judge complex artefacts — peer review of teaching portfolios, moderation of student work for programme-level assessment, ranking exemplars to set standards — comparative judgement can produce more reliable, defensible orderings than rubric scoring, and it makes standard-setting (what does "good" look like?) explicit through real exemplars rather than abstract descriptors.
-
Student course evaluation is a different case — and that matters. ACJ shines when a manageable number of expert judges compare a fixed pool of artefacts. A mass student survey is the opposite: thousands of novice raters, each experiencing only their own course. You cannot ask a student to compare two courses they did not both take. So the lesson for SET is not "replace Likert with pairwise voting"; it is more subtle — stop over-trusting absolute means, report comparatively and with uncertainty, and use comparative methods where the structure fits (e.g., asking students to rank aspects of their own course by importance, or a quality panel to compare exemplar feedback). MaxDiff/best-worst scaling is the survey-friendly cousin of this idea.
Limitations and honest caveats
Comparative judgement is powerful but not a universal upgrade, and the honest caveats are important.
- Reliability statistics can be inflated by adaptivity. This is the most serious technical caveat: the same adaptive pairing that sharpens the scale can artificially boost the Scale Separation Reliability coefficient. Reported reliabilities above 0.95 should be read with the design in mind, not taken at face value.
- It produces a rank order, not a diagnosis. Comparative judgement tells you that A is better than B; it does not, by itself, tell you why or what to fix. For formative course improvement, you still need qualitative reasons.
- Judge time and pool constraints. Building a reliable scale needs many pairwise judgements (often 10+ per item), which is feasible for expert moderation of a finite pool but impractical for mass, single-experience student surveys.
- Not a fit for self-experience ratings. Students cannot validly compare courses they did not both take; comparative judgement applies to evaluating artefacts or performances, not to aggregating individual lived experience.
- Transparency and acceptance. Stakeholders used to numbers and rubrics may distrust a "black-box" model that produces a scale from votes; defensibility for high-stakes use requires clear documentation.
The balanced reading: comparative judgement is an excellent tool for expert evaluation of complex work and standard-setting, a useful corrective to naive faith in absolute Likert means everywhere, and a poor fit for the specific job of aggregating thousands of students' individual course experiences.
How Koji incorporates this
Koji does not pretend that a mass student survey should become a pairwise-voting exercise — that would misapply the method. Instead, Koji takes the underlying insight (absolute single numbers are noisy; comparison and uncertainty are your friends) and builds it into how evaluation is collected and reported.
- Comparative and ranking question types where they fit. Koji supports
rankingitems, letting students order aspects of their own course (clarity, pace, assessment, feedback) by importance or quality — a comparison they can validly make — which yields more decision-useful priorities than a row of near-identical Likert means. This is the survey-appropriate expression of the comparative-judgement principle (closely related to best-worst scaling / MaxDiff). - AI-moderated interviews that elicit reasons, not just rankings. Because comparison gives order but not diagnosis, Koji's conversational interviewer probes why one aspect ranked above another — recovering the formative "what to fix" that a bare ranking omits.
- Uncertainty-aware reporting. Koji reports distributions and the imprecision of estimates rather than encouraging false confidence in a third-decimal-place mean — the same humility that the comparative-judgement literature urges when interpreting any single absolute score.
- A natural fit for expert moderation panels. For programme-level quality work — comparing exemplar courses, moderating open-text feedback quality, setting standards — Koji's structured data and thematic outputs give panels a consistent base to make the holistic comparative judgements that this research shows experts make well.
Framed honestly: Koji uses comparative and ranking formats to mitigate the noise of absolute single-item ratings where the structure allows, not to claim it has turned a course survey into a Thurstonian scale. Koji's core research platform at koji.so applies the same ranking and AI-interview tooling to product and customer research, where "which of these matters most" beats "rate each from 1 to 5" for prioritisation.
Related resources
- Best-Worst Scaling and MaxDiff for Course Evaluation Priorities
- Rasch Many-Facet Measurement and Rater Severity in Course Evaluation
- The Optimal Number of Rating Scale Points
- Should You Average Likert Scores? The Ordinal-Interval Debate
- Single-Item vs Multi-Item Global Ratings
- Slider vs Radio-Button Response Formats
References
- Pollitt, A. (2012). The method of Adaptive Comparative Judgement. Assessment in Education: Principles, Policy & Practice, 19(3), 281–300. https://doi.org/10.1080/0969594X.2012.665354
- Thurstone, L. L. (1927). A law of comparative judgment. Psychological Review, 34(4), 273–286. https://doi.org/10.1037/h0070288
- Kimbell, R. (2021). Examining the reliability of Adaptive Comparative Judgement (ACJ) as an assessment tool in educational settings. International Journal of Technology and Design Education, 31. https://doi.org/10.1007/s10798-021-09654-w
- Bramley, T. (2015). Investigating the reliability of Adaptive Comparative Judgment. Cambridge Assessment Research Report. https://www.cambridgeassessment.org.uk/Images/232694-investigating-the-reliability-of-adaptive-comparative-judgment.pdf
- Linacre, J. M. (1989). Many-facet Rasch measurement. MESA Press.
Related articles
Some Raters Are Harsh, Some Items Are Hard: Why Rasch and Many-Facet Models Beat Averaging Course-Evaluation Scores
Averaging Likert scores treats every student as an equally calibrated measuring instrument and every item as equally hard. Rasch and Many-Facet Rasch models do not. We explain the method, the evidence that rater severity and item difficulty distort raw means, the limitations, and how Koji applies the same logic.
Should You Average a 1–5 Course-Evaluation Scale at All? The Ordinal-vs-Interval Debate
Is it legitimate to compute a mean from a five-point course-evaluation scale? The fifty-year ordinal-versus-interval debate, from Jamieson to Carifio and Perla to Norman, and what it means for reporting student ratings.
When Everything Scores 4/5: Best-Worst Scaling (MaxDiff) for Course-Evaluation Priorities
Likert ratings on course evaluations cluster near the top and cannot tell you what matters most to students. Best-worst scaling (Louviere, Flynn & Marley) forces trade-offs that reveal genuine priorities. What the method is, its limits, and how it fits a Koji study.
Sliders, Visual-Analogue, or Radio Buttons? The Evidence on Course-Evaluation Response Formats
Slider widgets look modern, but the survey-methodology evidence says they cost you data. Funke (2016), Couper et al. (2006) and Bosch et al. (2019) on choosing a response widget for online course evaluations.