New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods10 min read

Beyond the Likert Scale: Adaptive Comparative Judgement in Evaluation

Humans are better at judging "which of these two is better" than at assigning an absolute number. Adaptive Comparative Judgement (Pollitt 2012), built on Thurstone's law of comparative judgement, turns pairwise choices into a reliable scale — with lessons for course evaluation.

Koji Education Team

Product

In brief: People are far more reliable at saying which of two things is better than at assigning an absolute score on a scale. Adaptive Comparative Judgement (ACJ) — formalised by Pollitt (2012) and rooted in Thurstone's (1927) law of comparative judgement — exploits this: judges make a series of holistic pairwise comparisons, and a statistical model converts those choices into a reliable rank order and interval scale, often reaching reliabilities above 0.90. For course evaluation, the deeper lesson is that the response format itself is a source of error: rigid Likert numbers carry rater-severity and scale-use noise that comparative methods largely sidestep.

What the research says

Most course evaluation rests on an assumption we rarely examine: that a student (or a peer reviewer, or an accreditation panel) can take a complex object — a course, a teaching performance, a piece of student work — and compress it into an absolute number on a 1-to-5 or 1-to-7 scale, consistently, and comparably with everyone else doing the same. Nearly a century of psychophysics says this assumption is shaky.

Thurstone (1927), in "A law of comparative judgement," showed that human judgement of magnitude is far more dependable when expressed as a comparison than as an absolute rating. The comparative judgement of two objects depends lawfully on the magnitude of their difference in quality, and a series of such comparisons can be modelled to place objects on an interval scale — without anyone ever assigning a number.

Pollitt (2012), "The method of Adaptive Comparative Judgement," in Assessment in Education: Principles, Policy & Practice, modernised Thurstone's idea for educational assessment. The method developed within a research project at Goldsmiths (2004–2010): instead of marking work against a rubric, an expert is shown two pieces and simply decides which is better. Across many such judgements by many judges, a Bradley–Terry/Rasch-style model estimates a quality parameter for each item and builds a scale. ACJ adds adaptivity in scoring (not testing): later pairings are chosen to be informative — pitting items of similar estimated quality against each other — so the scale sharpens quickly. In a large study of roughly 1,000 student writing samples, Pollitt reported a robust, valid rank order and strong support from educators, who often preferred holistic comparison to rubric marking.

The reliability evidence is the headline. Comparative-judgement exercises routinely report Scale Separation Reliability above 0.90, frequently exceeding the inter-rater reliability achievable with rubric-based absolute marking of the same work. Independent scrutiny has refined the claim: Cambridge Assessment researchers (e.g., Bramley) noted that the adaptive algorithm can inflate reliability statistics if not handled carefully, and a 2021 examination in the International Journal of Technology and Design Education re-checked ACJ reliability across settings. The mature consensus is that comparative judgement is genuinely reliable and valid, while the adaptive variant needs careful design to avoid over-stating its precision.

The connection to course evaluation is conceptual but direct. The well-known many-facet Rasch treatment of student ratings (Linacre) models rater severity precisely because absolute Likert ratings are contaminated by how harshly or leniently each rater uses the scale. Comparative judgement attacks the same problem at the source: if you never ask for an absolute number, you cannot suffer from one judge calling "good" a 4 and another calling the identical thing a 6.

Why it matters for course evaluation in practice

Course evaluation uses absolute scales almost everywhere — students rating a course, peers rating a teaching observation, panels scoring a programme against criteria. Three practical implications follow from the comparative-judgement literature.

  1. Format is a measurement choice, not a neutral container. The decision to use a 5-point Likert item is itself a source of error: ceiling effects, central-tendency bias, and divergent scale use all live in the format. Recognising this is the first step; it reframes "are our items well written?" into the larger "is an absolute rating even the right instrument for this judgement?"

  2. Some evaluation tasks are better suited to comparison. Where experts judge complex artefacts — peer review of teaching portfolios, moderation of student work for programme-level assessment, ranking exemplars to set standards — comparative judgement can produce more reliable, defensible orderings than rubric scoring, and it makes standard-setting (what does "good" look like?) explicit through real exemplars rather than abstract descriptors.

  3. Student course evaluation is a different case — and that matters. ACJ shines when a manageable number of expert judges compare a fixed pool of artefacts. A mass student survey is the opposite: thousands of novice raters, each experiencing only their own course. You cannot ask a student to compare two courses they did not both take. So the lesson for SET is not "replace Likert with pairwise voting"; it is more subtle — stop over-trusting absolute means, report comparatively and with uncertainty, and use comparative methods where the structure fits (e.g., asking students to rank aspects of their own course by importance, or a quality panel to compare exemplar feedback). MaxDiff/best-worst scaling is the survey-friendly cousin of this idea.

Limitations and honest caveats

Comparative judgement is powerful but not a universal upgrade, and the honest caveats are important.

  • Reliability statistics can be inflated by adaptivity. This is the most serious technical caveat: the same adaptive pairing that sharpens the scale can artificially boost the Scale Separation Reliability coefficient. Reported reliabilities above 0.95 should be read with the design in mind, not taken at face value.
  • It produces a rank order, not a diagnosis. Comparative judgement tells you that A is better than B; it does not, by itself, tell you why or what to fix. For formative course improvement, you still need qualitative reasons.
  • Judge time and pool constraints. Building a reliable scale needs many pairwise judgements (often 10+ per item), which is feasible for expert moderation of a finite pool but impractical for mass, single-experience student surveys.
  • Not a fit for self-experience ratings. Students cannot validly compare courses they did not both take; comparative judgement applies to evaluating artefacts or performances, not to aggregating individual lived experience.
  • Transparency and acceptance. Stakeholders used to numbers and rubrics may distrust a "black-box" model that produces a scale from votes; defensibility for high-stakes use requires clear documentation.

The balanced reading: comparative judgement is an excellent tool for expert evaluation of complex work and standard-setting, a useful corrective to naive faith in absolute Likert means everywhere, and a poor fit for the specific job of aggregating thousands of students' individual course experiences.

How Koji incorporates this

Koji does not pretend that a mass student survey should become a pairwise-voting exercise — that would misapply the method. Instead, Koji takes the underlying insight (absolute single numbers are noisy; comparison and uncertainty are your friends) and builds it into how evaluation is collected and reported.

  • Comparative and ranking question types where they fit. Koji supports ranking items, letting students order aspects of their own course (clarity, pace, assessment, feedback) by importance or quality — a comparison they can validly make — which yields more decision-useful priorities than a row of near-identical Likert means. This is the survey-appropriate expression of the comparative-judgement principle (closely related to best-worst scaling / MaxDiff).
  • AI-moderated interviews that elicit reasons, not just rankings. Because comparison gives order but not diagnosis, Koji's conversational interviewer probes why one aspect ranked above another — recovering the formative "what to fix" that a bare ranking omits.
  • Uncertainty-aware reporting. Koji reports distributions and the imprecision of estimates rather than encouraging false confidence in a third-decimal-place mean — the same humility that the comparative-judgement literature urges when interpreting any single absolute score.
  • A natural fit for expert moderation panels. For programme-level quality work — comparing exemplar courses, moderating open-text feedback quality, setting standards — Koji's structured data and thematic outputs give panels a consistent base to make the holistic comparative judgements that this research shows experts make well.

Framed honestly: Koji uses comparative and ranking formats to mitigate the noise of absolute single-item ratings where the structure allows, not to claim it has turned a course survey into a Thurstonian scale. Koji's core research platform at koji.so applies the same ranking and AI-interview tooling to product and customer research, where "which of these matters most" beats "rate each from 1 to 5" for prioritisation.

Related resources

References