One Score, Two Things: Item Response Tree Models Separate Opinion from Response Style
Item response tree (IRTree) models split each Likert answer into latent decisions — respond or stay neutral, agree or disagree, moderate or extreme — so a student's real opinion of a course can be estimated apart from their personal scale-use style.
Koji Education Team
Product
In brief
When a student marks a course item "4 out of 5", that single number blends two different things: what they actually think about the course, and how they personally use a rating scale. Item response tree (IRTree) models decompose each Likert response into a sequence of latent decisions — engage or stay neutral, agree or disagree, be moderate or extreme — and estimate a separate trait for each. That lets you recover a student's substantive opinion after partialling out response style, so a habitual mid-point picker and a habitual extreme responder can be compared on the same footing. IRTree models are members of the generalised linear mixed model family and can be fitted with standard software.
What the research says
The core idea is that a rating is not one act but several. Ulf Böckenholt formalised this in Modeling multiple response processes in judgment and choice (Böckenholt, 2012, Psychological Methods, 17(4), 665–678, DOI 10.1037/a0028111). He proposed that a response to a bipolar Likert item can be represented as a small decision tree: first a student decides whether to give a substantive answer or sit on the midpoint (a midpoint / indifference process), then a direction decision (agree vs disagree, the trait of interest), then an extremity decision (how far from the centre to go). Each node is governed by its own latent variable, so "content" is disentangled from two of the most common response styles — midpoint responding and extreme responding.
Paul De Boeck and Ivailo Partchev showed that these models are not exotic: IRTrees: Tree-Based Item Response Models of the GLMM Family (De Boeck & Partchev, 2012, Journal of Statistical Software, 48(c01), 1–28, DOI 10.18637/jss.v048.c01) demonstrates that every node in the tree is a binary (or ordinal) sub-model that can be estimated jointly with the glmer function in R's lme4 package. Because each node is an item response sub-model, item difficulties, discriminations and person abilities are all recoverable per process.
Methodological guidance has matured alongside the models. Hansjörg Plieninger's Developing and Applying IR-Tree Models: Guidelines, Caveats, and an Extension to Multiple Groups (Plieninger, 2021, Organizational Research Methods, 24(3), 655–685, DOI 10.1177/1094428120911096) walks through how to specify the tree, how to interpret the style dimensions, and — importantly — where the approach can mislead if the tree structure is wrong. This tradition sits alongside the broader literature on modelling rather than merely describing response styles, and it is the "how do we actually separate them" answer to the descriptive account in our companion note on response styles and Likert scales.
The practical payoff, repeatedly reported, is that response-style variance is not small. Extreme and midpoint response styles can account for a meaningful share of the variance in raw Likert scores, and — critically for evaluation — those styles are correlated with respondent characteristics such as culture, language and education. Ignore them and you are partly ranking instructors on the response habits of whoever happened to take their class.
Why it matters for course evaluation in practice
Course evaluation lives and dies on small differences between mean scores. If Instructor A averages 4.3 and Instructor B averages 4.1, a committee may treat that gap as signal. But if Instructor A's cohort contains more students who habitually pick the extremes and Instructor B's contains more midpoint-huggers, the gap can be an artefact of scale use, not teaching. IRTree models attack this directly by estimating a student's direction (opinion) trait net of their extremity and midpoint traits.
Three consequences follow. First, fairer cross-group comparison: international cohorts, mature students and different disciplines are known to differ systematically in response style; an IRTree analysis puts them on a comparable metric before aggregation, complementing formal measurement-invariance checks. Second, richer diagnostics: the style dimensions are not just nuisance — a sudden rise in midpoint responding can flag disengagement or satisficing, which is useful quality information in its own right. Third, honest reliability: because the direction trait is separated from style, its reliability is estimated on the substantive signal rather than on style-contaminated totals, which pairs naturally with a move to ordinal cumulative-link regression for reporting.
Limitations and honest caveats
A PhD reader will raise several objections, and they are fair.
- The tree is an assumption, not a fact. Böckenholt's midpoint-then-direction-then-extremity ordering is one plausible processing account; others exist (for example, an extremity-first tree). If the assumed tree is wrong, the separated traits are misestimated. Plieninger (2021) is explicit that tree misspecification is the central risk, and model comparison across candidate trees is essential.
- Style is confounded with content. A genuinely lukewarm student and a habitual midpoint responder can look alike on any single item. IRTree models rely on multiple items to identify the styles; with a short 5-item evaluation, the style dimensions are weakly identified and estimates are unstable.
- Response styles may not be stable traits. Treating extremity as a fixed personal trait assumes consistency across items and time; some evidence suggests styles shift with item content and mood, which weakens the "partial it out once" logic.
- Interpretability and communication. A dean wants a number, not three latent dimensions. Translating "direction trait adjusted for extremity" into a defensible report takes care, and over-adjustment can itself introduce bias if style and content are truly correlated for substantive reasons.
- Sample size. These are mixed models with several variance components; small classes (the norm in higher education) provide thin information, and the models are best applied at the department or programme level, not the individual 12-student seminar.
None of this makes IRTree models optional theatre — it makes them a tool to be used where the data can support them, with model comparison and transparent reporting.
How Koji incorporates this
Koji is built to generate exactly the kind of structured, multi-item, multi-process data that IRTree models need, and to reduce the response-style problem at the source.
- Structured question types that expose the process. Because Koji collects
scale,single_choice,yes_noandrankingitems with clean, consistent metadata, the response matrix is well formed for node-by-node modelling rather than being a bag of free-form ratings. - AI-moderated conversational follow-up. The largest defence against response style is not a statistical correction but a better measurement. When a student parks on the midpoint, Koji's AI moderator can ask a neutral probe — "what would have moved this higher or lower?" — turning an uninformative midpoint into substantive content. This is designed to mitigate midpoint responding at collection time, not to eliminate it.
- Bias-aware reporting. Koji's reporting layer is designed to flag when score differences are within the range plausibly explained by response-style or sampling variation, so committees are steered away from over-reading a 0.2-point gap — the same caution behind our 4.2 vs 4.4 note.
- Triangulation across cohorts. Koji retains item-level data across cohorts, which is what a group-extension IRTree analysis (Plieninger, 2021) requires to compare style-adjusted opinion across international or disciplinary groups.
Koji's core research platform at koji.so applies the same AI-moderated interview engine to product and customer research, where extreme-vs-midpoint response styles distort NPS and satisfaction scores in exactly the same way.
The honest framing: Koji does not "remove" response style with a switch. It collects the structured, probe-enriched, multi-item data that makes style modellable, and it presents scores with the uncertainty that style and small samples imply.
Frequently asked questions
What is an IRTree model in plain terms?
It is a way of treating a single Likert answer as the outcome of several yes/no-style decisions in sequence — first whether to be neutral, then which direction, then how extreme — and estimating a separate score for each decision, so a student's real opinion is separated from their scale-use habits.
How is this different from just using ordinal regression?
Ordinal regression respects the ordered, categorical nature of a Likert item but still treats the answer as one process. An IRTree model additionally splits that answer into distinct latent processes (midpoint use, direction, extremity), which ordinal regression alone does not do. The two are complementary.
Do I need huge samples to use IRTrees?
You need enough items and respondents to identify the style dimensions. A 5-item survey in a 12-person seminar is too thin; these models are best at programme or department scale, or pooled across cohorts.
Is separating out response style always the right thing to do?
No. If a group genuinely tends to use extremes for substantive reasons, partialling out extremity can remove real signal. Model comparison and theory should justify the tree before you adjust.
Can Koji fit an IRTree model for me automatically?
Koji focuses on collecting clean, structured, multi-item data and on reducing midpoint responding through conversational probes. The exported item-level data is formatted so an institutional-research team can fit IRTree (or ordinal) models in R or comparable tools.
References
- Böckenholt, U. (2012). Modeling multiple response processes in judgment and choice. Psychological Methods, 17(4), 665–678. https://doi.org/10.1037/a0028111
- De Boeck, P., & Partchev, I. (2012). IRTrees: Tree-based item response models of the GLMM family. Journal of Statistical Software, 48(Code Snippet 1), 1–28. https://doi.org/10.18637/jss.v048.c01
- Plieninger, H. (2021). Developing and applying IR-tree models: Guidelines, caveats, and an extension to multiple groups. Organizational Research Methods, 24(3), 655–685. https://doi.org/10.1177/1094428120911096
Related resources
- Response Styles and Likert Scales: Why Cross-Cultural Evaluation Needs More Than Numbers
- Can Forced-Choice Items Beat Response Bias? The Thurstonian IRT Evidence
- Rasch and Many-Facet Models Beat Averaging Course-Evaluation Scores
- Acquiescence Bias and Reverse-Worded Items
- Should a Course-Evaluation Scale Have a Neutral Midpoint?
- Ordinal Regression for Course-Evaluation Data
Related articles
Some Raters Are Harsh, Some Items Are Hard: Why Rasch and Many-Facet Models Beat Averaging Course-Evaluation Scores
Averaging Likert scores treats every student as an equally calibrated measuring instrument and every item as equally hard. Rasch and Many-Facet Rasch models do not. We explain the method, the evidence that rater severity and item difficulty distort raw means, the limitations, and how Koji applies the same logic.
Ordinal Regression for Course-Evaluation Data: Why Cumulative-Link Models Beat Averaging Likert Scores
Averaging Likert-scale course-evaluation items treats ordered categories as if they were equal-interval numbers. What the research says about the errors this causes, and why cumulative-link (ordinal) regression is the defensible analysis method.
Response Styles and Likert Scales: Why Cross-Cultural Evaluation Needs More Than Numbers
Acquiescence and extreme response styles vary systematically by culture (Harzing, 2006; Baumgartner & Steenkamp, 2001), which means raw Likert averages are not directly comparable across nationalities in Europe's multinational classrooms. This article explains the evidence and how to evaluate fairly across diverse cohorts.
Should a Course-Evaluation Scale Have a Neutral Midpoint? What "Neither Agree Nor Disagree" Really Captures
A neutral midpoint looks harmless, but the evidence shows it often functions as a hidden "don't know." Here is what the research says about including or omitting the middle option in course evaluations.