Why "The Pace Was About Right" Breaks Your Scale: Ideal-Point Unfolding Models for Course Evaluation
Most rating models assume more of a trait always means more agreement. For "appropriateness" items — workload, pace, difficulty — that assumption is false, and it quietly wrecks the scale. Ideal-point unfolding models fix it.
Koji Education Team
Product
In brief
Nearly every model used to score course evaluations — averaging Likert items, Rasch, Mokken, cumulative-link regression — assumes a dominance response process: the more of the underlying trait a person has, the more likely they are to endorse an item, monotonically. That assumption is correct for most teaching items ("the lecturer explained clearly") but false for a whole class of items about appropriateness: "the workload was about right", "the pace suited me", "the difficulty was appropriate". Here endorsement follows an ideal-point process — a student agrees most when the course matches their optimum and disagrees when it is either too much or too little. Fitting a dominance model to such items discards them as "poorly discriminating" and can invert their meaning. Ideal-point unfolding models, of which the Generalized Graded Unfolding Model (GGUM) is the standard, are built for exactly these items.
What the research says
The distinction traces to Clyde Coombs is A Theory of Data (1964), which formalised unfolding. Coombs is picture is a folded ruler: each person and each item occupies a position on a latent continuum, and a person is preference for an item declines with the distance between them, in either direction. Ask people to rank statements about an attitude and you can "unfold" their responses to recover both the item positions and each person is ideal point. The signature is non-monotonic: a moderate person endorses a moderate item, and disagrees with items at both extremes — a pattern a dominance model, which expects agreement to rise steadily with the trait, cannot represent.
James Roberts, John Donoghue and James Laughlin (2000, Applied Psychological Measurement 24(1):3-32, doi:10.1177/01466216000241001) turned this into a modern, estimable item-response model: the Generalized Graded Unfolding Model. GGUM gives each item a location on the trait and models the probability of each graded response (strongly disagree through strongly agree) as a function of the distance between the person and the item, allowing the characteristic curve to be single-peaked rather than monotone. A key GGUM insight is that a single observed "disagree" is ambiguous — it can come from a person below the item or above it — and the model resolves the ambiguity using the full response pattern. David Andrich had shown the same non-monotone logic applied to attitude scales built from Likert-style questionnaires (1988, Applied Psychological Measurement 12, 33-51), demonstrating empirically that intermediate attitude items misbehave under dominance scaling.
The provocative synthesis is Fritz Drasgow, Oleksandr Chernyshenko and Stephen Stark is 2010 focal article 75 Years After Likert: Thurstone Was Right! (Industrial and Organizational Psychology 3(4):465-476, doi:10.1111/j.1754-9434.2010.01273.x). They argue that three-quarters of a century of rating-scale practice has defaulted to a dominance (Likert) assumption that is often wrong, and that ideal-point methods descending from Thurstone better represent how people actually respond to attitude and preference statements — with direct consequences for scale construction and item selection. The claim drew published debate, particularly over how far it extends to personality measurement, and that contest is part of the honest picture. Tooling has since caught up: Jorge Tendeiro and Sebastian Castro-Alvarez is 2019 GGUM package (Applied Psychological Measurement, doi:10.1177/0146621618772290) made fitting these models routine.
Why it matters for course evaluation in practice
Course-evaluation instruments are quietly full of ideal-point items, and standard scaling mishandles every one. "The workload was appropriate", "the pace was about right", "the level of challenge suited me", "the amount of group work was reasonable" are all appropriateness judgements with an optimum in the middle. A student who found the workload perfect will disagree with "the workload was too heavy" and also with "the workload was too light" — and may sit in the middle on a bipolar "appropriate" item for reasons a dominance model reads as noise. When you drop these items into a Rasch or Mokken analysis, they show low discrimination or poor fit and get flagged for deletion — not because they are bad items, but because the model is wrong for them.
Three practical consequences follow. Item retention decisions get corrupted. A psychometrician cleaning an instrument with Rasch or many-facet tools will systematically purge the appropriateness items, leaving an instrument that measures only the monotone "more is better" dimensions and is blind to calibration. Middle responses are misread. On an appropriateness item the midpoint often means "just right", the most informative answer — the opposite of the "no opinion / satisficing" reading the neutral-midpoint literature warns about for dominance items. Scores become non-comparable. Averaging a mix of dominance and ideal-point items assumes they all point the same way; they do not, and the composite is incoherent. Recognising which items are which — and modelling the appropriateness items with an unfolding model — is what keeps the ordinal-vs-interval debate from being moot before it starts.
Limitations and honest caveats
Ideal-point modelling is powerful but narrow, and overselling it would be a mistake. Most course-evaluation items really are dominance items, and for those a Rasch or cumulative-link model is correct; unfolding is for the appropriateness/optimum subset, not the whole instrument. Misclassifying a monotone item as ideal-point is as wrong as the reverse. Unfolding models are data-hungry and harder to estimate. GGUM needs larger samples and more careful estimation than dominance IRT; small classes will not support a stable item-level unfolding fit, so the method is a design-and-validation tool applied to pooled data, not something you run on one seminar. Identifiability and interpretation are subtler. Because a "disagree" is directionally ambiguous, results depend on the full pattern and on defensible anchoring; a careless fit can mislabel item locations. The Drasgow et al. thesis is contested beyond attitudes, and a responsible reader should treat "Thurstone was right" as a strong, useful hypothesis about appropriateness items rather than a settled law for all rating data. And it addresses response process, not response bias — an unfolding model correctly scaled can still sit atop a leniency- or acquiescence-contaminated instrument.
How Koji incorporates this
Koji is question design and analysis are built to respect the difference between "more is better" items and "just right" items rather than forcing both through one monotone scale. When a course-evaluation study includes appropriateness judgements — workload, pace, challenge, balance of activities — Koji is analysis layer is designed to treat a mid-scale "about right" as the informative optimum it is, not as fence-sitting, so a well-calibrated course is not penalised by a model that only rewards extreme agreement. Rather than lean on a single Likert number for these constructs, Koji is AI-moderated conversational interview is designed to probe directionally: when a student signals the pace was off, the follow-up establishes whether it was too fast or too slow, recovering exactly the direction that a lone unfolding item leaves ambiguous and that a flat rating discards entirely. That turns an appropriateness item from a fragile scale point into a two-sided, actionable finding for a programme director — too heavy for this cohort, too light for that one — which is far more useful than an averaged 3.4. For instrument validation across many cohorts, Koji is psychometric tooling is designed to flag when an "appropriateness" item is behaving non-monotonically so it is modelled and reported as an ideal-point construct instead of being silently deleted as a weak dominance item. Koji is core research platform at koji.so applies the same ideal-point awareness to product and customer research, where "the right amount" judgements — pricing, feature scope, frequency — are pervasive and equally mishandled by monotone scoring.
The disciplined message is that the response process must match the model. For most teaching qualities, dominance scaling is right. For every judgement with an optimum in the middle, it is wrong — and an ideal-point unfolding model, or a conversational probe that recovers direction, is what keeps those items meaningful.
Frequently asked questions
What is the difference between a dominance and an ideal-point item?
For a dominance item, endorsement rises monotonically with the trait — the clearer you found the teaching, the more you agree "the teaching was clear". For an ideal-point item, endorsement peaks when the course matches your optimum and falls off on both sides — "the workload was about right" is disagreed with by students who found it too heavy and those who found it too light.
Which course-evaluation items are ideal-point items?
Appropriateness and calibration judgements: workload, pace, difficulty, level of challenge, amount of group work or assessment — anything where "just right" sits in the middle and both extremes are undesirable. Quality judgements (clarity, helpfulness, organisation) are usually dominance items.
Why does fitting a Rasch model to these items cause problems?
Rasch and other dominance models expect agreement to increase steadily with the trait. An ideal-point item violates that, so it shows low discrimination or misfit and is often deleted during instrument cleaning — removing a valid item because the wrong model was applied to it.
Do I need a special model, or can I just reword the items?
Rewording into two unipolar items ("too heavy" / "too light") converts one ideal-point item into two dominance items and is often the simplest practical fix. When you must keep a bipolar "appropriate" item, an unfolding model such as GGUM scores it correctly; a conversational follow-up that recovers direction achieves the same end.
Is GGUM practical for a single course?
No. Unfolding models need substantial samples and careful estimation, so they are validation tools run on pooled, multi-cohort data, not per-seminar analyses. For a single small class, reworded unipolar items or a directional probe are the realistic route.
Does an ideal-point model remove rating biases?
No. It corrects the response process for appropriateness items so the scale is coherent, but it does not address leniency, acquiescence or halo. Use it alongside bias-aware analysis, not as a replacement.
References
- Coombs, C. H. (1964). A Theory of Data. New York: Wiley.
- Andrich, D. (1988). The application of an unfolding model of the PIRT type to the measurement of attitude. Applied Psychological Measurement, 12(1), 33-51.
- Roberts, J. S., Donoghue, J. R., & Laughlin, J. E. (2000). A general item response theory model for unfolding unidimensional polytomous responses. Applied Psychological Measurement, 24(1), 3-32. doi:10.1177/01466216000241001
- Drasgow, F., Chernyshenko, O. S., & Stark, S. (2010). 75 years after Likert: Thurstone was right! Industrial and Organizational Psychology, 3(4), 465-476. doi:10.1111/j.1754-9434.2010.01273.x
- Tendeiro, J. N., & Castro-Alvarez, S. (2019). GGUM: An R package for fitting the generalized graded unfolding model. Applied Psychological Measurement. doi:10.1177/0146621618772290
Related Resources
- Some Raters Are Harsh, Some Items Are Hard: Why Rasch and Many-Facet Models Beat Averaging Course-Evaluation Scores
- Do Your Evaluation Items Actually Form a Scale? Mokken Analysis and Nonparametric Item Response Theory
- Should You Average a 1–5 Course-Evaluation Scale at All? The Ordinal-vs-Interval Debate
- Should a Course-Evaluation Scale Have a Neutral Midpoint? What "Neither Agree Nor Disagree" Really Captures
- Can Forced-Choice Items Beat Response Bias in Course Evaluations? The Thurstonian IRT Evidence
- Acquiescence Bias and Reverse-Worded Items: Should Course Evaluations Flip the Question?
Related articles
Some Raters Are Harsh, Some Items Are Hard: Why Rasch and Many-Facet Models Beat Averaging Course-Evaluation Scores
Averaging Likert scores treats every student as an equally calibrated measuring instrument and every item as equally hard. Rasch and Many-Facet Rasch models do not. We explain the method, the evidence that rater severity and item difficulty distort raw means, the limitations, and how Koji applies the same logic.
Should You Average a 1–5 Course-Evaluation Scale at All? The Ordinal-vs-Interval Debate
Is it legitimate to compute a mean from a five-point course-evaluation scale? The fifty-year ordinal-versus-interval debate, from Jamieson to Carifio and Perla to Norman, and what it means for reporting student ratings.
Should a Course-Evaluation Scale Have a Neutral Midpoint? What "Neither Agree Nor Disagree" Really Captures
A neutral midpoint looks harmless, but the evidence shows it often functions as a hidden "don't know." Here is what the research says about including or omitting the middle option in course evaluations.
Can Forced-Choice Items Beat Response Bias in Course Evaluations? The Thurstonian IRT Evidence
Forced-choice (ipsative) formats were designed to suppress acquiescence, halo, and social-desirability response styles that contaminate ordinary Likert course evaluations. Brown and Maydeu-Olivares'' Thurstonian IRT model solves the classic ipsative-data problem — but the format is costly to build and not a free lunch. Here is the evidence and what it means for a university QA office.