Can Forced-Choice Items Beat Response Bias in Course Evaluations? The Thurstonian IRT Evidence
Forced-choice (ipsative) formats were designed to suppress acquiescence, halo, and social-desirability response styles that contaminate ordinary Likert course evaluations. Brown and Maydeu-Olivares'' Thurstonian IRT model solves the classic ipsative-data problem — but the format is costly to build and not a free lunch. Here is the evidence and what it means for a university QA office.
Koji Education Team
Product
In brief
Forced-choice items — where students rank statements within a block instead of rating each one on a Likert scale — were designed to suppress the response styles (acquiescence, halo, central tendency, social desirability) that inflate or flatten ordinary course-evaluation data. The historical objection was that forced-choice scoring produces ipsative data (each respondent''s scores sum to a constant), which breaks normal statistics. Brown and Maydeu-Olivares'' Thurstonian Item Response Theory (TIRT) model removed that objection: it recovers genuine, between-person comparable trait estimates from forced-choice blocks. The catch is that forced-choice questionnaires are expensive to construct, harder for respondents, and only worth it when response-style contamination is a documented, material problem. For most course evaluations, a well-built Likert instrument plus targeted screening is sufficient; forced-choice is a specialist tool, not a default.
What the research says
A response style is a tendency to use a rating scale in a way unrelated to the content being rated. The best-documented are acquiescence (agreeing regardless of item), extreme responding and its mirror central tendency, the halo effect (one global impression colouring every item), and socially desirable responding. These styles are not random noise — they are systematic, person-specific, and correlated with culture, personality, and the stakes of the survey, so they bias group comparisons in ways that simple averaging cannot remove.
The forced-choice format attacks the problem at the source. Instead of asking "Rate how clearly the lecturer explained concepts (1–5)" and a dozen similar agree/disagree items — all of which an acquiescent or halo-driven respondent will answer near-identically — a forced-choice block presents several roughly equally desirable statements and asks the student to pick the one that is most (and sometimes least) like the course. Because the statements are matched on desirability, you cannot satisfy a social-desirability pull by endorsing all of them, and because the response is comparative, a uniform "agree with everything" strategy is impossible.
The historical price was ipsativity. Classical forced-choice scoring counts how often each dimension "wins," so a respondent who rates one dimension higher must, mechanically, rate another lower; the scores sum to a constant. Ipsative scores cannot be compared between people in the usual way, distort factor analysis and reliability estimates, and make normative interpretation invalid. For decades this was the decisive argument against forced choice.
Brown and Maydeu-Olivares (2011, Educational and Psychological Measurement) resolved this. Building on Thurstone''s 1927 law of comparative judgement, they modelled each binary outcome of a comparison (item A preferred to item B) as a function of the items'' latent utilities, which in turn load on the underlying traits. The result — the Thurstonian IRT model — is a second-order factor model with binary indicators that yields normative, between-person-comparable trait estimates from purely comparative data. Their follow-up (Brown & Maydeu-Olivares, 2012, Behavior Research Methods) gave applied researchers a concrete recipe for fitting the model in Mplus, which is what moved the method from theory into practice. Thurstonian-based forced-choice assessments are now used at scale in high-stakes personnel selection — precisely the context where faking and social desirability are worst — across dozens of countries and languages.
Two corroborating strands round out the picture. First, the broad response-styles literature (reviewed by Van Vaerenbergh & Thomas, 2013, in the International Journal of Public Opinion Research) confirms that response styles are pervasive, stable within persons, and a genuine threat to the comparability of Likert data — establishing the problem forced choice is meant to solve. Second, simulation and integration studies (e.g., the 2017 Frontiers in Psychology simulation integrating forced-choice and Likert formats) show that forced-choice TIRT scores can approach the accuracy of Likert scores when the blocks are well designed and contain a mix of positively and negatively keyed items — and degrade when they are not. The method works, but it is sensitive to construction quality.
Why it matters for course evaluation in practice
Course evaluations are exactly the kind of low-stakes-for-students, high-stakes-for-staff survey where response styles do real damage. The halo effect is the most consequential: a student who liked the instructor tends to rate clarity, fairness, organisation, and workload all uniformly high, collapsing a multidimensional instrument into a single "did I like this?" signal (see our note on the halo effect). Acquiescence quietly lifts every agree/disagree item (our acquiescence-bias article covers the reverse-wording remedy). Cross-cultural response styles make raw Likert means non-comparable across international cohorts (see response styles and Likert scales). And central-tendency straightlining shows up as the satisficing we describe in why students click straight down the middle.
Forced choice promises to neutralise all four at once, because none of the uniform-response strategies survive a matched-desirability ranking task. For a programme director comparing teaching dimensions within a course — "is the problem clarity, or assessment, or pace?" — a forced-choice block that ranks those dimensions against each other can be more diagnostic than four separate 4.3-vs-4.4 Likert means that all sit at the ceiling. This connects directly to the best-worst scaling / MaxDiff approach, which is a close cousin of forced choice for prioritisation.
But the practical bar is high. Forced choice only earns its keep when (a) you have evidence that response-style contamination is materially distorting your data, (b) you can invest in proper block construction with desirability-matched, mixed-keyed statements, and (c) your reporting pipeline can handle TIRT scoring rather than counting. Absent those, you risk replacing a familiar bias with an unfamiliar measurement headache.
Limitations and honest caveats
A PhD reader will rightly push back on several points:
- Construction cost and expertise. TIRT requires desirability-matched statement pools, blocks balancing positively and negatively keyed items, and model fitting in software like Mplus or the R
thurstonianIRTpackage. This is a research project, not a form redesign. Poorly matched blocks reintroduce the very desirability bias they were meant to remove. - Respondent burden and reactivity. Ranking statements is cognitively heavier than ticking a scale. On the small screens most students use (see device effects), this can raise breakoff and frustration — potentially trading measurement bias for nonresponse bias.
- Recovery is not perfect. TIRT recovers normative information well only when designs are rich enough (enough blocks, mixed keying). With short questionnaires — and course evaluations are usually short — trait estimates can be less reliable than a good Likert scale. The format does not manufacture information that the design did not collect.
- Generalisability. Most validation evidence comes from personality and selection settings, not course evaluation. Extrapolating effect sizes to a 10-item teaching survey is an inference, not a demonstration.
- Interpretability for committees. Evaluation committees already misread Likert means (see Boysen on small mean differences). Latent TIRT scores are harder still to explain to a non-technical promotion panel, which can undermine the legitimacy of the whole exercise.
The honest summary: forced choice is a validated solution to a real problem, but it is a specialist instrument whose costs are easy to underestimate and whose benefits depend entirely on doing the construction properly.
How Koji incorporates this
Koji for Education does not ask universities to rebuild their instruments around ipsative blocks. Instead it targets the underlying problem — response styles that flatten and contaminate fixed-scale data — using its AI-moderated conversational design:
- Probing beyond the Likert number. Where a forced-choice block makes the halo effect mechanically impossible, Koji''s AI-moderated interview makes it behaviourally unattractive: after a high overall rating, the interviewer asks the student to name the single weakest aspect of the course and explain it. This is a conversational analogue of "pick the least-like" forced choice — it breaks the uniform-positive response set without imposing ipsative scoring.
- Mixed question types in one flow. Koji supports
scale,single_choice,multiple_choice,ranking,yes_no, andopen_endeditems. Therankingtype lets a programme run a genuine forced-choice-style prioritisation ("rank these five aspects of the module from most to least in need of improvement") alongside conventional scales, so teams can adopt the diagnostic strength of comparative judgement where it helps without converting the whole survey. - Bias-aware reporting. Koji''s reporting flags straightlining and suspiciously uniform response patterns (the central-tendency and acquiescence signatures) rather than averaging over them, and triangulates the numeric scale against the student''s own open-text reasoning.
- Designed to mitigate, not eliminate. Koji frames these as mitigations: conversational probing reduces, but does not abolish, halo and social-desirability effects, and the platform reports them transparently rather than claiming a bias-free score.
The same AI-moderated interview engine powers Koji''s core research platform at koji.so, where comparative and probing question designs are used for product and customer research — the underlying methodology travels across both domains.
Related Resources
- Agree/Disagree or Item-Specific? The Scale Choice That Quietly Degrades Data
- Acquiescence Bias and Reverse-Worded Items
- Response Styles and Likert Scales in Cross-Cultural Evaluation
- Best-Worst Scaling (MaxDiff) for Course-Evaluation Priorities
- Why Students Click Straight Down the Middle: Satisficing
- The Halo Effect in Course Evaluations
References
- Brown, A., & Maydeu-Olivares, A. (2011). Item Response Modeling of Forced-Choice Questionnaires. Educational and Psychological Measurement, 71(3), 460–502. https://doi.org/10.1177/0013164410375112
- Brown, A., & Maydeu-Olivares, A. (2012). Fitting a Thurstonian IRT model to forced-choice data using Mplus. Behavior Research Methods, 44(4), 1135–1147. https://doi.org/10.3758/s13428-012-0217-x
- Van Vaerenbergh, Y., & Thomas, T. D. (2013). Response Styles in Survey Research: A Literature Review of Antecedents, Consequences, and Remedies. International Journal of Public Opinion Research, 25(2), 195–217. https://doi.org/10.1093/ijpor/eds021
- Cao, M., & Drasgow, F. (2019). Does forcing reduce faking? A meta-analytic review of forced-choice personality measures in high-stakes situations. Journal of Applied Psychology, 104(11), 1347–1368. https://doi.org/10.1037/apl0000414
- Sass, R., et al. (2017). Integration of the Forced-Choice Questionnaire and the Likert Scale: A Simulation Study. Frontiers in Psychology, 8, 806. https://doi.org/10.3389/fpsyg.2017.00806
Related articles
Why Students Click Straight Down the Middle: Satisficing in Course Evaluations
A large share of students do not answer evaluation questions carefully — they straightline, speed through, and pick the easy option. Barge and Gehlbach (2012) showed this "satisficing" doesn't just add noise; it inflates your reliability and validity statistics, making bad data look good.
Response Styles and Likert Scales: Why Cross-Cultural Evaluation Needs More Than Numbers
Acquiescence and extreme response styles vary systematically by culture (Harzing, 2006; Baumgartner & Steenkamp, 2001), which means raw Likert averages are not directly comparable across nationalities in Europe's multinational classrooms. This article explains the evidence and how to evaluate fairly across diverse cohorts.
The Halo Effect in Course Evaluations: When One Impression Colours Every Rating
When students like an instructor, that single global impression bleeds into their ratings of unrelated specifics. From Nisbett & Wilson (1977) to Feistauer & Richter (2018) and Cannon & Cipriani (2021), here is how the halo effect distorts itemised course evaluations — and what to do about it.
When Everything Scores 4/5: Best-Worst Scaling (MaxDiff) for Course-Evaluation Priorities
Likert ratings on course evaluations cluster near the top and cannot tell you what matters most to students. Best-worst scaling (Louviere, Flynn & Marley) forces trade-offs that reveal genuine priorities. What the method is, its limits, and how it fits a Koji study.