Everything Scored 4/5: Best-Worst Scaling and the Priorities Your Likert Course Evaluation Cannot Reveal
When students rate every item on your course evaluation as important, the ranking you need to act on disappears into a cluster of 4.1s. Best-worst scaling forces the trade-offs a rating scale never does — and tells you what to fix first.
Koji Education Team
Product ·
The short answer: Likert rating scales let students mark every aspect of a course "important" or "good," which is why end-of-term reports so often show a flat band of 3.9–4.3 with no usable priority order. Best-worst scaling (also called MaxDiff, or maximum-difference scaling) removes that escape route by asking students to pick the most and least important item from small sets, forcing genuine trade-offs. The result is a discriminating, interval-like priority ranking that is largely immune to the scale-use biases — acquiescence, extreme-response, central-tendency — that flatten conventional evaluations. It is not a replacement for every question, but for the "what matters most to you and what should we fix first" decisions that course evaluation exists to inform, it is a measurable improvement over the average.
The problem: a rating scale gives everyone permission to say "yes"
Ask students to rate ten aspects of a module on a five-point scale — clarity of explanation, usefulness of feedback, pace, assessment fairness, reading load, and so on — and a predictable thing happens. Most items land between 3.8 and 4.3. A programme director reading the report cannot tell whether "feedback usefulness" at 4.0 is a genuine priority relative to "pace" at 4.1, because the difference is within noise, and because nothing in the question format ever required the student to say which one mattered more.
This is not student carelessness. It is a structural feature of independent rating scales: each item is judged in isolation, so a respondent can — quite honestly — agree that all of them are good or important. Survey methodologists have documented the consequences for decades under the headings of acquiescence bias (the tendency to agree), extreme-response style, and central-tendency bias (clustering on the midpoint). We have written about each of these individually in Acquiescence, Straightlining and the Data-Quality Problem and Central Tendency Bias. Best-worst scaling attacks all of them at once, by changing the question rather than correcting the answer after the fact.
What best-worst scaling actually is
Best-worst scaling (BWS) was formalised by marketing scientist Jordan Louviere in 1987 and has since become a standard discrete-choice method across market research, health-preference studies, and increasingly education. Instead of rating items one by one, respondents see a small subset — typically four or five items at a time — and are asked two questions: which of these is most important? and which is least important? Across a designed sequence of these subsets, every item appears several times against different competitors.
Because each choice extracts the maximum-difference pair (the best and the worst) from a set, a single BWS task yields more information than a single rating. Aggregated across respondents and choice sets, the method produces a ratio- or interval-scaled ranking of relative importance — a clean ordering with meaningful distances between items, not a pile of near-identical means. In practice you learn not just that "assessment fairness" outranks "reading load," but roughly by how much.
Why it beats the Likert scale for prioritisation
Three advantages matter for course evaluation specifically:
1. It forces trade-offs. A rating scale lets a student call everything a 4. A best-worst task does not; picking one item as "most important" is simultaneously a statement that the others are less so. This is the whole point — real decisions are trade-offs, and a priority instrument should elicit them.
2. It is robust to scale-use heterogeneity. Different students interpret "4 out of 5" differently; some never use the extremes, others never use the middle. This is the measurement invariance problem we cover in Does a "4" Mean the Same Thing to Everyone?. Because BWS records only choices within a set rather than absolute scores, it sidesteps most of that heterogeneity: a choice of "most important" means the same act regardless of how a respondent would have calibrated a number.
3. It discriminates. Humans are demonstrably better at judging extremes than at placing middling options on an abstract scale. BWS exploits this, producing sharper separation between items than either rating or full ranking, with lower respondent fatigue than ranking a long list.
"But doesn't this just move the problem around?" — the honest counterargument
A methodologically literate reader will push back, and rightly. Three objections deserve a straight answer.
"Best-worst only gives you relative importance, not absolute levels." Correct, and important. BWS tells you students value clear feedback more than they value a lighter reading load; it does not tell you whether feedback is currently good or bad. It is a prioritisation instrument, not a performance measure. You still need a separate — ideally qualitative — read on how each priority item is actually performing. The right design pairs a best-worst prioritisation with open-ended probing, not one instead of the other.
"It adds respondent burden and design complexity." Also fair. A BWS block requires an experimental design (balanced, orthogonal appearance of items) and more clicks than a rating grid. For a ten-item list this is modest; for a fifty-item list it is not, and rating scales or a hybrid may be the pragmatic choice. BWS earns its cost precisely when you have a shortlist of competing priorities and a real decision to make about them.
"You are trading one bias for the artefacts of forced choice." Forcing a choice among items a student considers genuinely equal introduces its own noise. This is why BWS should be reserved for questions where prioritisation is the actual goal, and why the "least important" pick is as informative as the "most" — it lets weakly-held items reveal themselves across repeated sets.
None of these caveats is fatal. They define the instrument's scope: BWS is the right tool for "rank what matters," the wrong tool for "measure how good each thing is," and best used alongside, not instead of, qualitative feedback.
Where Koji fits
Koji for Education is built around six structured question types, and ranking is one of them — the native primitive for prioritisation rather than isolated rating. More importantly, Koji does not stop at the ranking. Its AI-moderated conversational interviews take a student's top and bottom priorities and probe them: when a student marks "assessment feedback" as the single most important aspect, the interviewer asks why, and what specifically about the feedback helped or failed. That is the pairing the methodology demands — a forced-choice priority signal, immediately enriched with the reasoning behind it, analysed thematically at scale rather than left as a bar chart.
This matters for the action gap we describe in Importance-Performance Analysis: knowing what students prioritise is only half of a decision. Koji surfaces the priority and the performance narrative in one conversation, so a teaching team sees not just "feedback ranks first" but what to change about it. Because the moderation is standardised, every student is probed consistently — no human-interviewer drift — and because it is GDPR/AVG-compliant and EU-appropriate in its data handling, the richer data does not come at a governance cost.
The same conversational interview engine powers general customer and user research on the main Koji platform — the prioritisation-plus-probing pattern is not unique to education; it is how good discovery works anywhere you need to know what matters most and why.
Running a best-worst block in practice
You do not need a large item pool to benefit. A practical design shortlists the eight to twelve course attributes that genuinely compete for a teaching team's attention — clarity, feedback quality, pace, assessment fairness, workload, resources, support, and so on — and shows each student a handful of balanced subsets, so every attribute appears the same number of times against varied competitors. Four or five items per screen and five to six screens is usually enough to estimate a stable ranking without fatiguing respondents. The output is a single ordered list with relative importance scores, ready to plot against how each attribute is currently performing. Keep the pool honest: including an attribute you have no power to change wastes a choice and muddies the priorities you can act on. And resist the temptation to run best-worst on everything — reserve it for the genuine prioritisation decision, and let ordinary questions handle the rest.
The takeaway
If your course-evaluation report is a wall of 4.1s, the problem is not your students — it is that you asked them to rate everything independently and they obliged. Best-worst scaling asks the question course evaluation is actually for: of these things, which matters most, and which least? Used within its scope — prioritisation, not performance measurement, and paired with qualitative depth — it turns a flat average into an ordered, actionable list. That is the difference between a survey you file and a survey you act on.