New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods9 min read

When Everything Scores 4/5: Best-Worst Scaling (MaxDiff) for Course-Evaluation Priorities

Likert ratings on course evaluations cluster near the top and cannot tell you what matters most to students. Best-worst scaling (Louviere, Flynn & Marley) forces trade-offs that reveal genuine priorities. What the method is, its limits, and how it fits a Koji study.

Koji Education Team

Product

In brief: When most course-evaluation items score 4 or 5 out of 5, the ratings stop discriminating — you cannot tell which improvements students actually want. Best-worst scaling (BWS), also called MaxDiff, fixes this by asking students to choose the best and worst item from small sets, forcing the trade-offs that flat Likert scales hide. Developed by Louviere and formalised in Louviere, Flynn and Marley (2015), BWS produces clear, ratio-like priority rankings and sidesteps several response-style biases. It is a powerful complement to — not a replacement for — diagnostic and open-text feedback.

The ceiling-effect problem

Anyone who reads course-evaluation data knows the pattern: nearly every item sits between 4.0 and 4.6 on a five-point scale. Lectures, materials, assessment, feedback, organisation — all "good". This ceiling effect is partly genuine goodwill and partly the well-documented tendency of rating scales to compress at the top. The consequence is that ratings fail at the one job a quality team most needs them to do: telling you where to spend limited improvement effort. If feedback scores 4.2 and assessment scores 4.3, that 0.1 is noise, not a priority order.

Best-worst scaling attacks this directly by changing the response task from rating each item in isolation to choosing among items.

What the research says

Best-worst scaling was invented by Jordan Louviere in 1987 and developed over the following decades, with the definitive treatment in Louviere, Flynn and Marley (2015), Best-Worst Scaling: Theory, Methods and Applications (Cambridge University Press). The method shows respondents a subset of items (at least three) and asks them to pick the most and the least important — the best and the worst. Across a balanced set of such choices, each item earns a score derived from how often it was chosen best versus worst. The authors show BWS is grounded in random-utility theory and can supersede conventional ratings-based and discrete-choice surveys for measuring the relative importance of a set of objects.

The mechanism matters. Because respondents make a relative judgement (this is more important than that), BWS removes the scale-use biases that plague Likert data: acquiescence (agreeing with everything), and response styles where some respondents or cultures cluster at the extremes while others stay near the middle. As Louviere and colleagues argue, a respondent cannot rate everything "5" in a best-worst task; the format compels discrimination.

The approach has a strong applied track record outside marketing. Flynn, Louviere, Peters and Coast (2007), "Best-worst scaling: What it can do for health care research and how to do it" (Journal of Health Economics, 26(1), 171-189), demonstrated BWS for eliciting patient priorities and provided a practical how-to that the survey-methods community widely adopted. Health, education, and public-policy researchers have since used BWS precisely where Likert ratings ceiling out and where cross-group comparability is needed.

Why it matters for course evaluation in practice

For a programme director, BWS answers a question end-of-term Likert grids cannot: of all the things we could improve, which do students themselves rank highest? A best-worst block might present students with rotating sets drawn from items such as "clarity of assessment criteria", "speed of feedback", "quality of learning materials", "pace of teaching", "availability of staff", and "relevance to career". Forced to pick the most and least important in each set, students reveal a genuine priority ordering rather than a wall of 4s.

This has several practical payoffs:

  • Actionable prioritisation. The output is a ranked, well-separated list of what to fix first — ideal for closing-the-loop planning and for showing accreditors evidence-based prioritisation.
  • Cross-cohort and cross-cultural comparability. Because BWS dampens response-style differences, comparing priorities between, say, domestic and international cohorts is more defensible than comparing their Likert means, a problem we discuss under cross-cultural response styles.
  • Resistance to straightlining. The choice format makes the satisficing shortcut of clicking straight down the middle impossible, improving data quality relative to long Likert grids.

BWS works best as a focused priorities module — six to twelve items, a handful of choice sets — sitting alongside diagnostic ratings and open text, not as the whole instrument.

Limitations and honest caveats

BWS is not a free lunch, and a methodologically literate reader will want the caveats. It measures relative, not absolute, standing. BWS tells you assessment feedback is the top priority relative to the other items you listed; it does not tell you whether feedback is objectively poor or merely least-good among strong options. A programme where everything is genuinely excellent and one where everything is weak can produce identical BWS rankings. This is the mirror image of the Likert ceiling problem and the reason BWS should accompany, not replace, absolute measures.

The item set is the study. Results are entirely conditioned on which attributes you include; omit "mental-health support" and it cannot rank. Item selection therefore demands the same care — ideally cognitive pretesting — as any questionnaire.

Design and analysis are more demanding. Balanced incomplete block designs and the modelling of best-worst counts (from simple count analysis to multinomial logit) require more expertise than computing Likert means, and respondents may find the repeated choices more effortful, which can raise burden if overused. Finally, the basic ("Case 1", object-level) BWS used here is simpler than the profile and multi-profile variants; conflating them in analysis is a common error.

How Koji incorporates this

Koji's structured-question toolkit and conversational engine make best-worst-style prioritisation practical inside a normal evaluation.

  • Ranking and choice question types. Koji supports ranking, single_choice, and multiple_choice items, which let an evaluation designer build best-worst-style prioritisation tasks that force trade-offs rather than collecting another grid of 4s — directly targeting the ceiling effect.
  • The conversational follow-up explains the ranking. Where classic BWS gives you an ordering but not a reason, Koji's AI-moderated interview can probe why a student ranked feedback worst, and automatic thematic analysis aggregates those reasons — pairing the relative priority with diagnostic, absolute context that BWS alone lacks.
  • Comparable priorities across cohorts. Because the choice format dampens response-style noise, Koji's cohort comparisons of priorities are more defensible than comparisons of raw means, supporting fair cross-group and cross-cultural reporting.
  • Closing the loop on what matters most. Feeding a clear, well-separated priority list into action tracking lets institutions act on the top student-identified issues and evidence that action for accreditation.

Koji frames BWS as a complement designed to mitigate the ceiling and response-style limits of Likert data, not a cure-all — absolute diagnostic measures and open text remain essential. The same prioritisation logic powers the core Koji research platform at koji.so, where product teams use best-worst tasks to rank feature priorities instead of drowning in uniformly high satisfaction scores.

A practical design: building a six-item priorities block

Consider a programme team that wants to know where to direct a limited improvement budget. They define six candidate priorities: clarity of assessment criteria, speed of feedback, quality of learning materials, pace of teaching, availability of staff, and relevance to career. A best-worst block presents these in rotating subsets of three or four, asking students to mark the most and least important in each set. A balanced design ensures every item appears an equal number of times and against every other item, so no attribute is advantaged by position. Even simple count analysis — the share of times each item is chosen best minus the share chosen worst — yields a clear, well-separated ranking, and more formal multinomial-logit estimation produces interval-scaled importance weights.

The contrast with a Likert grid is stark. On a five-point grid the same six attributes might score 4.4, 4.3, 4.5, 4.2, 4.4 and 4.1 — a spread of 0.4 that is mostly noise and offers no honest basis for prioritisation. The best-worst output might instead reveal that speed of feedback is chosen "worst" far more often than anything else while relevance to career dominates "best", giving the team an unambiguous mandate. Because every student is forced to discriminate, the result is robust to the ceiling effect and to the acquiescent respondent who would otherwise rate everything highly.

Two design cautions complete the picture. Keep the block short — six to twelve items across a handful of sets — to control the cognitive burden the repeated choices impose. And select the items with care, ideally cognitively pretested, because anything left off the list simply cannot be ranked; the method is only as complete as the attributes you choose to include.

Related Resources

References

  • Louviere, J. J., Flynn, T. N., & Marley, A. A. J. (2015). Best-Worst Scaling: Theory, Methods and Applications. Cambridge: Cambridge University Press. https://doi.org/10.1017/CBO9781107337855
  • Flynn, T. N., Louviere, J. J., Peters, T. J., & Coast, J. (2007). Best-worst scaling: What it can do for health care research and how to do it. Journal of Health Economics, 26(1), 171-189. https://doi.org/10.1016/j.jhealeco.2006.04.002
  • Louviere, J. J., & Woodworth, G. G. (1991). Best-Worst Scaling: A Model for Largest Difference Judgments. Working Paper, University of Alberta.

Related articles

research-methods

Response Styles and Likert Scales: Why Cross-Cultural Evaluation Needs More Than Numbers

Acquiescence and extreme response styles vary systematically by culture (Harzing, 2006; Baumgartner & Steenkamp, 2001), which means raw Likert averages are not directly comparable across nationalities in Europe's multinational classrooms. This article explains the evidence and how to evaluate fairly across diverse cohorts.

research-methods

How Many Scale Points Should a Course-Evaluation Question Have?

What the measurement literature — Preston & Colman (2000), Weng (2004), Dawes (2008) — says about the optimal number of response categories on rating scales, and why the answer for course evaluation is not just a number but a question about what a Likert item can and cannot capture.

research-methods

What Do Student Evaluations Actually Measure? Marsh, the SEEQ, and the Case for Multidimensional Feedback

Herbert Marsh spent decades showing that well-built student evaluations are multidimensional, reliable, and stable — and that a single global score throws away most of what they can tell you. What his work establishes, where critics push back, and how to design feedback that is actually usable.

research-methods

Why Students Click Straight Down the Middle: Satisficing in Course Evaluations

A large share of students do not answer evaluation questions carefully — they straightline, speed through, and pick the easy option. Barge and Gehlbach (2012) showed this "satisficing" doesn't just add noise; it inflates your reliability and validity statistics, making bad data look good.