Stop Asking Students to Rate Everything Highly: Discrete Choice Experiments Reveal What They Will Trade Off
Discrete choice experiments (DCEs) ask students to choose between course scenarios rather than rate each feature, forcing real trade-offs and yielding a random-utility estimate of what actually drives their preferences.
Koji Education Team
Product
In brief
Ask students to rate feedback timeliness, lecture clarity, workload and assessment fairness on a 1–5 scale and they will often rate everything important. That tells you little about priorities. A discrete choice experiment (DCE) instead shows students two or three realistic course profiles that differ on several attributes and asks which they prefer. Because every choice forces a trade-off, the resulting data — analysed with random-utility (logit) models — reveals how much each attribute actually drives preference, and even the rate at which students would trade one good feature for another. DCEs are the established preference-elicitation method in health economics and marketing, and they transfer cleanly to course design and evaluation.
What the research says
The theoretical engine is random utility theory: a person chooses the option whose total utility — a weighted sum of its attribute levels plus a random component — is highest. Daniel McFadden's conditional logit model (McFadden, 1974, Conditional logit analysis of qualitative choice behavior, in Zarembka, ed., Frontiers in Econometrics, Academic Press, 105–142) gave this a tractable estimator and later a Nobel Prize; it remains the workhorse for analysing choice data. Each estimated coefficient is the marginal utility of an attribute level, and ratios of coefficients express willingness to trade (in health economics, willingness to pay; in course design, willingness to accept more workload for better feedback, say).
The most-cited practical guide is Emily Lancsar and Jordan Louviere's Conducting discrete choice experiments to inform healthcare decision making: A user's guide (Lancsar & Louviere, 2008, PharmacoEconomics, 26(8), 661–677, DOI 10.2165/00019053-200826080-00004). It lays out the full workflow: define attributes and levels, build a statistically efficient experimental design, field the choice sets, and estimate preferences with logit-family models. Their central methodological point is that a good DCE lives or dies on the attribute set and the experimental design, not on sample size alone.
Esther de Bekker-Grob, Mandy Ryan and Karen Gerard reviewed how the method has actually been used: Discrete choice experiments in health economics: A review of the literature (de Bekker-Grob, Ryan & Gerard, 2012, Health Economics, 21(2), 145–172, DOI 10.1002/hec.1697) synthesised 114 DCEs and documented the field's maturing standards for design, estimation (mixed logit to capture preference heterogeneity) and validity testing. The heterogeneity point matters for education: students are not one preference type, and modern DCE analysis estimates a distribution of preferences rather than a single average.
Within course evaluation, DCEs are distinct from the priority tools already in the corpus. They differ from best-worst scaling (MaxDiff), which ranks single items within one list; from adaptive comparative judgement, which pairs whole artefacts to build a quality scale; and from importance-performance analysis, which plots self-reported importance against satisfaction. A DCE is the only one of these that models multi-attribute trade-offs under a formal utility framework and can quantify the exchange rate between attributes.
Why it matters for course evaluation in practice
Standard evaluations suffer from ceiling effects and "everything matters" responses. A DCE breaks the ceiling by design: students cannot mark every attribute as essential, because each choice sacrifices one thing for another. This yields three practically useful outputs.
First, a defensible priority ranking for course redesign. If students consistently choose profiles with prompt, specific feedback even at the cost of a heavier reading load, that is far stronger evidence than a 4.6 average on a "feedback was useful" item. It connects directly to the recurring finding that assessment and feedback score lowest — a DCE tells you whether students would actually accept trade-offs to fix it.
Second, willingness-to-accept estimates. Because coefficient ratios express trade-offs, a programme can state, for example, that students value a shift from end-of-term to weekly feedback about as much as a one-band improvement in lecture clarity. That is design-grade information a mean score cannot give.
Third, preference heterogeneity. Mixed-logit analysis (de Bekker-Grob et al., 2012) reveals whether a single course design fits everyone or whether distinct student segments want different things — the same segmentation question addressed descriptively by latent profile analysis, but here grounded in actual choices.
A fourth, often overlooked use is design simulation under constraint. Because a DCE yields a utility weight for each attribute level, a programme facing a fixed budget or an immovable timetable can simulate which combination of feasible changes would shift preference the most, and whether a proposed change is even large enough for students to notice. That is a form of policy simulation a set of rating means simply cannot support: it lets a course team choose among rival redesign options, and rule out cosmetic changes, before committing scarce teaching resources.
Limitations and honest caveats
- Hypothetical bias. Students state preferences over imagined course profiles; stated choices can diverge from what they would do when real effort and grades are at stake. The health-economics literature treats external validity as an open, actively debated question, not a settled one.
- Design complexity. A valid DCE needs a carefully constructed, statistically efficient design and defensible attribute levels. Get the attributes wrong — omit one that matters, or use levels students find implausible — and the estimates are biased regardless of sample size (Lancsar & Louviere, 2008).
- Cognitive burden. Choice tasks are harder than ticking a Likert scale. Too many attributes or choice sets and students satisfice, attend to only one attribute, or drop out — degrading the very trade-off information the method exists to capture.
- It measures preference, not learning. A DCE tells you what students want, which is not the same as what teaches them most; the feeling-of-learning gap applies here too. DCE evidence should inform design and satisfaction, not stand in for learning-outcome evidence.
- Analytic expertise. Logit and mixed-logit estimation, efficient design generation and interpretation of trade-off ratios require statistical skill that not every quality office has in-house.
How Koji incorporates this
Koji's question engine and AI moderation make choice-based elicitation practical inside a normal evaluation cycle, not just a specialist research study.
- Native choice and ranking items. Koji supports
single_choice,multiple_choiceandrankingquestion types, which are the building blocks of choice sets — students can be shown competing course-design profiles and asked to choose, with the responses captured in a clean, analysis-ready structure. - Conversational rationale capture. After a choice, Koji's AI moderator can ask why a profile was picked, converting a bare choice into an explained one. This helps diagnose whether students attended to all attributes or fixated on one — the key validity check the literature flags.
- Segment-aware reporting. Because Koji stores choice data at the respondent level, downstream mixed-logit or latent-class analysis of preference heterogeneity is possible, and Koji's reporting is designed to surface distinct student segments rather than a single average preference.
- Honest framing of trade-offs. Koji's reports frame DCE outputs as revealed priorities and trade-offs, deliberately kept separate from learning-outcome and satisfaction metrics, so a committee does not read "students prefer X" as "X teaches better".
Koji's core research platform at koji.so runs the same choice-based and conversational elicitation for product and customer research, where DCEs and conjoint analysis are standard tools for feature prioritisation.
The honest position: Koji makes it feasible to field choice tasks and capture the reasoning behind them; the efficient experimental design and logit estimation are still a methods task for the institution's analysts, and DCE evidence answers "what do students prioritise", not "what did they learn".
Frequently asked questions
How is a DCE different from just asking students to rank features?
A ranking asks students to order a single list. A DCE shows competing whole-course profiles that vary on several attributes at once and asks which they would choose, so it captures how attributes trade off against each other — information a simple ranking cannot give.
How is it different from MaxDiff or importance-performance analysis?
MaxDiff picks best and worst from one list of items; importance-performance plots self-rated importance against satisfaction. Only a DCE models multi-attribute trade-offs under a formal utility model and can estimate the exchange rate between attributes.
Does a DCE tell me what teaches students best?
No. It reveals what students prefer and will trade off, which is about satisfaction and design, not learning. Keep DCE results separate from learning-outcome evidence.
How many students do I need?
There is no single number; adequacy depends on the design efficiency, the number of attributes and levels, and whether you model preference heterogeneity. A well-designed choice set can be informative with moderate samples, but a poorly specified attribute set is not rescued by a large one.
Is hypothetical bias a dealbreaker?
It is a real limitation: stated choices can differ from real behaviour. Treat DCE outputs as strong evidence about priorities, validate against actual behaviour where possible, and avoid over-precise willingness-to-accept claims.
References
- McFadden, D. (1974). Conditional logit analysis of qualitative choice behavior. In P. Zarembka (Ed.), Frontiers in Econometrics (pp. 105–142). Academic Press.
- Lancsar, E., & Louviere, J. (2008). Conducting discrete choice experiments to inform healthcare decision making: A user's guide. PharmacoEconomics, 26(8), 661–677. https://doi.org/10.2165/00019053-200826080-00004
- de Bekker-Grob, E. W., Ryan, M., & Gerard, K. (2012). Discrete choice experiments in health economics: A review of the literature. Health Economics, 21(2), 145–172. https://doi.org/10.1002/hec.1697
Related resources
- Best-Worst Scaling (MaxDiff) for Course-Evaluation Priorities
- Adaptive Comparative Judgement in Evaluation
- Importance-Performance Analysis: Turning Scores into a Priority Map
- The Kano Model and the Asymmetry of Student Satisfaction
- Should You Use Net Promoter Score for Courses?
- The Repertory Grid Technique for Course Evaluation
Related articles
Should You Use Net Promoter Score for Courses? The "Would You Recommend" Question in Higher Education
Net Promoter Score is migrating from customer experience into student feedback. What Reichheld (2003) actually claimed, why Keiningham et al. (2007) failed to replicate its superiority, and whether a single recommend-question belongs in course evaluation.
Importance-Performance Analysis: Turning Course-Evaluation Scores into a Priority Map
A course-evaluation report that lists twenty item means tells you nothing about where to act first. Importance-Performance Analysis (IPA) plots each attribute by how much it matters to students against how well you did, producing a four-quadrant map that separates urgent fixes from wasted effort.
When Everything Scores 4/5: Best-Worst Scaling (MaxDiff) for Course-Evaluation Priorities
Likert ratings on course evaluations cluster near the top and cannot tell you what matters most to students. Best-worst scaling (Louviere, Flynn & Marley) forces trade-offs that reveal genuine priorities. What the method is, its limits, and how it fits a Koji study.
Beyond the Likert Scale: Adaptive Comparative Judgement in Evaluation
Humans are better at judging "which of these two is better" than at assigning an absolute number. Adaptive Comparative Judgement (Pollitt 2012), built on Thurstone's law of comparative judgement, turns pairwise choices into a reliable scale — with lessons for course evaluation.