New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods10 min read

Do Your Evaluation Items Actually Form a Scale? Mokken Analysis and Nonparametric Item Response Theory

Before you average a set of course-evaluation items into a subscale score, you should check that they form a scale at all. Mokken scale analysis tests that with far weaker assumptions than Rasch or factor analysis — and tells you whether the items even order students the same way.

Koji Education Team

Product

In brief

Averaging several course-evaluation items into a "teaching quality" or "workload" subscale silently assumes those items belong together and measure one underlying thing in the same direction. Mokken scale analysis (MSA) is a nonparametric item response theory method that tests exactly that assumption using only weak, checkable conditions — monotonicity and the ordering of items — rather than the strong parametric form a Rasch model or a linear factor model imposes. Its central statistic, the scalability coefficient H, tells you whether a set of items forms a usable scale, which items to drop, and whether the items order every student in the same way.

For a quality-assurance office deciding whether to report a composite score or its separate items, MSA is the honest first check. It is complementary to the Rasch and many-facet tradition and to the dimensionality debate: where those impose or test a parametric structure, Mokken asks the more basic question first — is there a scale here at all?

What the research says

Mokken scale analysis originates with Robert Mokken's (1971) A Theory and Procedure of Scale Analysis, which set out a nonparametric alternative to the parametric IRT models (such as Rasch) then emerging. The modern reference treatments are Sijtsma and Molenaar (2002), Introduction to Nonparametric Item Response Theory, and the applied tutorial by Sijtsma and Van der Ark (2017), A Tutorial on How to Do a Mokken Scale Analysis on Your Test and Questionnaire Data (British Journal of Mathematical and Statistical Psychology, 70(1), 137-158), which walks through the procedure for both dichotomous and polytomous (Likert) items — the format course evaluations actually use.

MSA rests on two nested nonparametric models:

  • The monotone homogeneity model (MHM) assumes (a) one underlying latent trait, (b) local independence, and (c) item response functions that are monotonically non-decreasing in the trait. If the MHM holds, the total score orders respondents stochastically on the latent trait — which is the property that justifies using a summed or averaged score to rank students at all.
  • The double monotonicity model (DMM) adds that the item response functions do not cross, which implies invariant item ordering: the items keep the same difficulty order for every respondent. When that holds, "the item most students agree with" is the same item for weak and strong groups alike.

The workhorse statistics are Loevinger's scalability coefficients H. There is an H for each item pair, an H for each item (its scalability against the rest), and an H for the whole scale. Conventional rules of thumb, given in Sijtsma and Van der Ark (2017): item and scale H below 0.3 means the items do not form a scale; 0.3-0.4 is a weak scale; 0.4-0.5 medium; above 0.5 strong. An automated item selection procedure partitions a pool of items into one or more Mokken scales, which is why MSA doubles as a dimensionality check: if your 12 evaluation items split into two Mokken scales, you have two constructs, not one.

Crucially, MSA makes no assumption of a specific parametric form and no assumption of interval-level, normally distributed data. This is the appeal for Likert course-evaluation data, which are ordinal and typically skewed — the same data properties that motivate cumulative-link ordinal regression and cast doubt on naive averaging. Reviews of nonparametric IRT in quality-of-life and health measurement (a field with the same ordinal-scale problem) report MSA as a robust scale-construction tool precisely because it survives when parametric assumptions fail.

Why it matters for course evaluation in practice

Most institutions report subscale composites — an "assessment and feedback" score, a "teaching" score — built by averaging several items. Three practical questions MSA answers before you trust such a composite:

  1. Do these items form a scale at all? If the scale H is below 0.3, averaging the items produces a number without a coherent underlying construct. Reporting it as "the teaching-quality score" is measurement theatre. MSA gives you the empirical basis to keep, drop, or split items.

  2. Do the items order students the same way (invariant item ordering)? This is the quiet assumption behind any claim that one item is "the hardest to agree with". If item ordering is not invariant — the DMM fails — then the item profile means different things for different students, and item-level benchmarking across cohorts is on shaky ground. MSA tests it directly with the HT coefficient.

  3. How many dimensions are really there? The automated item-selection algorithm is a data-driven dimensionality tool that does not require you to pre-specify a factor structure. It pairs naturally with parallel analysis as a cross-check: if both a parametric factor-retention method and a nonparametric scaling method agree the items split into two clusters, that conclusion is robust to the method chosen.

For an instrument like the SEEQ or CEQ, where the multidimensionality of student ratings is the whole design philosophy, MSA offers an assumption-light way to verify that the intended dimensions actually behave as scales in your population — not just in the instrument's original validation sample.

Limitations and honest caveats

  • Weak assumptions still permit weak conclusions. MSA confirms the ordering property of the total score; it does not deliver an interval-level measure. If you need genuine interval measurement — for example, to claim the gap between 4.0 and 4.5 equals the gap between 3.0 and 3.5 — a fitted Rasch model, where it holds, gives you more. MSA tells you a scale exists; it does not upgrade your ordinal data to interval.

  • H thresholds are conventions, not laws. The 0.3/0.4/0.5 bands are rules of thumb from Mokken and reaffirmed by Sijtsma and Van der Ark; they are not significance tests. Range-preserving confidence intervals for H exist and should be reported, because with small class sizes the point estimate of H is unstable.

  • Local independence is assumed, not guaranteed. If two items are near-duplicates ("The lecturer was clear" and "The lecturer explained clearly"), they violate local independence and inflate H artificially, mimicking a strong scale that is really redundancy. Content review must accompany the statistics.

  • Small samples and short scales. MSA's item-selection procedure can capitalise on chance in small samples, finding "scales" that do not replicate. Cross-validation on a hold-out cohort, or at minimum a bootstrap, is advisable before an institution commits to a reported subscale structure.

  • It describes structure, not bias. A set of items can form a strong Mokken scale and still be systematically biased by course difficulty, discipline, or instructor gender. Scalability is a measurement property; it is silent about the confounds this corpus documents. Do not read a high H as evidence that the score is fair.

How Koji incorporates this

Koji for Education treats "should these items be combined?" as an empirical question rather than a design assumption:

  • Structured item types that MSA can analyse. Koji's structured questions — scale, single_choice, multiple_choice, ranking, yes_no — produce exactly the ordinal item-level data Mokken analysis needs. Because responses are captured item by item with clean metadata, an institution can run scalability analysis on its own instrument rather than trusting a vendor's claim that the items "form a scale".

  • Dimension checks before composite reporting. Where Koji reports a composite or thematic score, the design intent is that the grouping reflects how items actually behave in that population, cross-checked against dimensionality evidence, rather than being hard-coded. This is designed to prevent the "measurement theatre" of averaging items that do not cohere.

  • Beyond the Likert number. The deeper point of MSA — that a single number hides whether items even mean the same thing to different students — is the same motivation behind Koji's AI-moderated conversational interviews. When a scalability check flags an item that behaves oddly, the conversational layer can probe why a construct is not holding together, surfacing the reason a Likert score alone cannot. Koji's core research platform at koji.so applies the same engine to product and customer surveys, where composite indices raise the identical "do these items scale?" question.

  • Honest, uncertainty-aware reporting. Koji is designed to present scale-quality evidence (including the instability of coefficients in small classes) rather than a bare composite, so a quality committee can see whether a subscale is a strong scale or a weak aggregate before acting on it.

Related resources

References

  • Mokken, R. J. (1971). A Theory and Procedure of Scale Analysis: With Applications in Political Research. Mouton / De Gruyter. https://doi.org/10.1515/9783110813203
  • Sijtsma, K., & Molenaar, I. W. (2002). Introduction to Nonparametric Item Response Theory. Sage. https://doi.org/10.4135/9781412984676
  • Sijtsma, K., & Van der Ark, L. A. (2017). A tutorial on how to do a Mokken scale analysis on your test and questionnaire data. British Journal of Mathematical and Statistical Psychology, 70(1), 137-158. https://doi.org/10.1111/bmsp.12078
  • Straat, J. H., Van der Ark, L. A., & Sijtsma, K. (2013). Comparing optimization algorithms for item selection in Mokken scale analysis. Journal of Classification, 30(1), 75-99. https://doi.org/10.1007/s00357-013-9122-y
  • Wind, S. A. (2017). An instructional module on Mokken scale analysis. Educational Measurement: Issues and Practice, 36(2), 50-66. https://doi.org/10.1111/emip.12153

Frequently asked questions

How is Mokken scale analysis different from factor analysis?

Factor analysis assumes a linear relationship between items and a continuous latent factor and typically treats ordinal Likert data as interval. Mokken analysis is nonparametric: it assumes only that item response functions are monotonic in the latent trait, with no linear form and no interval-scale assumption. Both can be used to check dimensionality, but Mokken survives when the parametric assumptions of factor analysis are violated, which is common for skewed ordinal course-evaluation data.

What does the scalability coefficient H actually tell me?

H measures how consistently items order respondents on the latent trait. Sijtsma and Van der Ark's conventions: below 0.3 the items do not form a scale; 0.3-0.4 is weak, 0.4-0.5 medium, and above 0.5 a strong scale. There is an H for each item and an H for the whole scale, so you can identify and drop individual items that weaken the scale.

What is invariant item ordering and why should I care?

Invariant item ordering means the items keep the same difficulty order for every respondent — the item most people agree with is the same item for weak and strong groups. It is the property tested by the double monotonicity model. If it fails, comparing item profiles across cohorts is unsafe, because an item does not have a stable meaning across the range of students.

Is Mokken analysis better than a Rasch model?

Neither is universally better; they answer different questions. Mokken tests whether a scale exists under weak assumptions and is robust when data are ordinal and non-normal. A Rasch model, when it fits, delivers interval-level measurement and additional properties, at the cost of stronger assumptions. A sensible workflow uses Mokken as the assumption-light first check and Rasch where interval measurement is genuinely required.

Can I run Mokken analysis on a small class?

With caution. The item-selection procedure can find scales that do not replicate in small samples, and the point estimate of H is unstable. Report confidence intervals for H and, where possible, cross-validate the scale on a separate cohort before committing to a reported subscale structure.

Does a strong Mokken scale mean my evaluation is fair?

No. Scalability is a measurement property: it tells you the items cohere and order students consistently. It says nothing about whether the score is biased by course difficulty, discipline, class size, or instructor characteristics. A well-scaling instrument can still be systematically unfair, so scalability evidence must be combined with confounding analysis, not treated as a substitute for it.

Related articles

analysis-reporting

Some Raters Are Harsh, Some Items Are Hard: Why Rasch and Many-Facet Models Beat Averaging Course-Evaluation Scores

Averaging Likert scores treats every student as an equally calibrated measuring instrument and every item as equally hard. Rasch and Many-Facet Rasch models do not. We explain the method, the evidence that rater severity and item difficulty distort raw means, the limitations, and how Koji applies the same logic.

analysis-reporting

How Many Dimensions Does Your Evaluation Really Measure? Parallel Analysis for Factor Retention

Deciding how many factors your evaluation instrument measures with the eigenvalue-greater-than-one rule or the scree plot routinely gets the wrong answer. Parallel analysis (Horn, 1965) compares your data against random noise — and is one of the most accurate methods available.

research-methods

What Do Student Evaluations Actually Measure? Marsh, the SEEQ, and the Case for Multidimensional Feedback

Herbert Marsh spent decades showing that well-built student evaluations are multidimensional, reliable, and stable — and that a single global score throws away most of what they can tell you. What his work establishes, where critics push back, and how to design feedback that is actually usable.

research-methods

How Long Should a Course Evaluation Be? Questionnaire Length, Breakoff, and Answer Quality

What the survey-methodology evidence says about questionnaire length: longer instruments depress participation and degrade answers to later questions, but ruthless shortening is not automatically the answer. A research-grounded guide for designing course evaluations.