Measurement Invariance: Can You Compare Course-Evaluation Scores Across Groups at All?
Comparing average ratings across departments, languages or online vs paper assumes the questionnaire means the same thing to every group. Measurement invariance is the test that assumption usually fails — and almost nobody runs it.
Koji Education Team
Product
In short: Comparing average course-evaluation scores across groups — departments, languages, online vs paper, cohorts — silently assumes the questionnaire measures the same construct on the same scale for every group. That assumption is called measurement invariance, and unless at least scalar (strong) invariance holds, an observed difference in means can be an artefact of the instrument rather than a real difference in teaching. Most institutions never test it. When it has been tested carefully — for example the Romanian SEEQ across online and paper administration (2024) — invariance can hold, but that is something you demonstrate, not assume.
Why this is the most-ignored problem in evaluation reporting
Almost every quality dashboard ranks or compares mean scores: this department vs that one, this term vs last, online sections vs in-person, the English-language programme vs the national-language one. Every one of those comparisons rests on a hidden psychometric premise — that a "4 out of 5" means the same thing to a first-year engineer and a final-year humanities student, in two languages, on two devices. Measurement invariance is the formal name for that premise, and the literature shows it cannot be taken for granted.
What the research says
The conceptual backbone comes from Vandenberg and Lance (2000) in Organizational Research Methods, whose review of the measurement-invariance literature established the now-standard hierarchy of tests, each a stricter condition than the last:
- Configural invariance — the same items load on the same factors in every group; the questionnaire has the same basic structure. This is the weakest level.
- Metric (weak) invariance — the factor loadings are equal across groups, so a one-unit change in the underlying construct produces the same change in item scores everywhere. Needed before you can compare relationships (correlations, regressions) across groups.
- Scalar (strong) invariance — the item intercepts are also equal across groups. This is the level you must reach before comparing observed or latent means across groups. Without it, a difference in average scores is confounded with a difference in how groups use the scale.
When scalar invariance fails at the level of an individual item, that item is said to exhibit Differential Item Functioning (DIF): respondents who stand at the same level on the underlying trait — equally well-taught, say — but who belong to different groups have systematically different expected answers to that item. DIF is measurement invariance failing, item by item. It is the psychometric machinery underneath every intuitive worry about bias: if international students interpret "the lecturer was approachable" differently from domestic students, that item has DIF, and averaging across the two groups mixes a real signal with a measurement artefact.
The encouraging news is that invariance can hold when an instrument is well built. A 2024 study in Studies in Educational Evaluation — "Student Evaluation of Teaching: The analysis of measurement invariance across online and paper-based administration procedures of the Romanian version of Marsh's Student Evaluations of Educational Quality scale" — tested exactly this on a sample of 809 students, one group completing the SEEQ on paper and another completing it online for the same teachers and courses. The Romanian SEEQ passed internal-validity and reliability checks and demonstrated configural, metric and scalar invariance across the two administration modes. In that specific case, therefore, comparing online and paper scores was psychometrically justified — because the authors tested it rather than assuming it.
That result also connects to the original instrument. Marsh (1982), introducing the SEEQ in the British Journal of Educational Psychology, argued that a defensible evaluation instrument must be multidimensional and demonstrate stable measurement properties across the settings in which it is used. Invariance testing is simply the modern, formal way of holding an instrument to Marsh's standard before its numbers are compared.
Why it matters for course evaluation in practice
The practical lesson is uncomfortable: most cross-group comparisons in routine evaluation reporting have never been checked for invariance, which means a chunk of the "differences" that drive decisions may be measurement artefacts. Three high-risk comparisons deserve particular caution:
- Cross-language comparisons. Multilingual European institutions routinely compare programmes taught and evaluated in different languages. Translation rarely preserves item intercepts exactly; DIF is the rule, not the exception, and uncorrected mean comparisons across languages are the least defensible of all.
- Mode comparisons (online vs paper, mobile vs desktop). The Romanian SEEQ study shows invariance can hold across mode — but it shows it by testing, and other instruments may not clear the bar.
- Cross-disciplinary or cross-cohort league tables. Ranking departments on mean scores presumes scalar invariance across very different student populations and course types; that presumption is almost never examined.
The constructive response is not to abandon comparison but to earn it: keep instruments stable, test invariance before publishing cross-group rankings, and where invariance fails, report group results separately or focus on within-group change over time (which is far more robust to DIF than between-group level comparisons).
Limitations and honest caveats
A rigorous reader should not over-rotate on invariance either:
- Invariance testing needs sample size and a latent-variable model. Confirmatory factor analysis or item-response models require reasonable n per group; small modules cannot support the test, which is itself a limitation on what can be compared.
- Partial invariance is common and contested. Real instruments often achieve invariance on most but not all items. There are defensible procedures for partial-invariance comparisons, but they involve judgement calls reasonable methodologists disagree about.
- A single study does not generalise. The Romanian SEEQ result is specific to that instrument, language and sample; it licenses the practice of testing, not a blanket claim that SET is mode-invariant everywhere.
- Invariance is necessary, not sufficient. An instrument can be perfectly invariant and still measure the wrong thing (for example, charisma rather than teaching quality). Invariance protects comparisons; it does not validate the construct.
How Koji incorporates this
Koji is an AI-native course-evaluation platform, and measurement invariance shapes how it is designed to support defensible comparison rather than encourage misleading league tables — framed, as always, as designed to mitigate rather than eliminate:
- Stable, structured instruments. Koji keeps
scaleand other structured item types consistent across cohorts and modes, which is the precondition for any invariance argument; ad-hoc question changes between terms quietly destroy comparability. - Probing where DIF is most likely. Where numeric items are most exposed to differential interpretation — across languages, international cohorts, or unusual course formats — the AI moderator collects an
open_endedexplanation alongside the rating, so a committee can see whether two groups mean the same thing by the same number instead of assuming it. - Thematic analysis instead of naïve averaging. By clustering open-text responses, Koji surfaces why a group scored as it did, which is exactly the information a raw cross-group mean hides when DIF is present.
- Bias-aware reporting that resists false precision. Reports are built to favour within-group trends and triangulated evidence over headline cross-group rankings, the comparisons most fragile to invariance failure.
The same AI-moderated interview engine powers Koji's core product and customer-research platform at koji.so, where comparing segments fairly raises the identical measurement-invariance question; the education product applies that discipline to course quality.
A concrete DIF scenario
Make it tangible. Suppose an item reads "The lecturer was approachable." In a culture with high power-distance norms, students may reserve the top of the scale for staff regardless of behaviour, while students from a low power-distance background use the full range. Two cohorts with identically approachable lecturers then return different average scores on that item — not because the teaching differed, but because the item functions differently across the groups. That is DIF, and it is invisible in the mean. Detecting it requires a latent-variable model (confirmatory factor analysis or item-response theory), which in turn needs a workable sample per group — a practical reason small modules cannot support cross-group claims at all. Where formal testing is impossible, the defensible fallbacks are: report groups separately rather than pooling, prefer within-group change over time to between-group level comparisons, and pair any cross-group number with qualitative evidence that reveals whether the groups mean the same thing. The discipline is not statistical perfectionism; it is refusing to publish a ranking the data cannot bear.
Related resources
- Online vs Paper Course Evaluations: Do Lower Response Rates Mean Worse Data?
- Response Styles and Likert Scales: Why Cross-Cultural Evaluation Needs More Than Numbers
- Generalizability Theory and the Reliability of Student Ratings
- What Do Student Evaluations Actually Measure? Marsh and the SEEQ
- How Many Responses Do You Need for a Reliable Course Evaluation?
- Can You Fairly Rank Instructors by Their Course-Evaluation Scores?
References
- Vandenberg, R. J., & Lance, C. E. (2000). A review and synthesis of the measurement invariance literature: Suggestions, practices, and recommendations for organizational research. Organizational Research Methods, 3(1), 4–70. https://doi.org/10.1177/109442810031002
- Student Evaluation of Teaching: The analysis of measurement invariance across online and paper-based administration procedures of the Romanian version of Marsh's Student Evaluations of Educational Quality scale (2024). Studies in Educational Evaluation, 80, 101331. https://doi.org/10.1016/j.stueduc.2024.101331 (https://www.sciencedirect.com/science/article/pii/S0191491X24000191)
- Marsh, H. W. (1982). SEEQ: A reliable, valid, and useful instrument for collecting students' evaluations of university teaching. British Journal of Educational Psychology, 52(1), 77–95. https://doi.org/10.1111/j.2044-8279.1982.tb02505.x
Related articles
How Many Responses Do You Need for a Reliable Course Evaluation?
Reliability of a class-average rating rises steeply with the number of students who respond. Generalizability theory — from Marsh''s syntheses to Dzakadzie (2026) — shows raters matter more than items, and a class of 12 cannot be evaluated like a class of 120.
Can You Fairly Rank Instructors by Their Course-Evaluation Scores? What Esarey & Valdes (2020) Proved
Even if course evaluations were unbiased, reliable, and valid, ranking instructors by their scores would still misclassify many good teachers. A walk through the Esarey & Valdes (2020) simulation and what it means for how you report and use evaluation data.
Generalizability Theory and the Reliability of Student Ratings: Why One Class Is Not Enough
Reliability is not one number. Generalizability theory (Gillmore, Kane & Naccarato 1978; Marsh 1984) decomposes the variance in student ratings into student, teacher, course and occasion components — and shows that a single class can be reliable for the course yet a poor estimate of the teacher. What that means for fair evaluation.
Response Styles and Likert Scales: Why Cross-Cultural Evaluation Needs More Than Numbers
Acquiescence and extreme response styles vary systematically by culture (Harzing, 2006; Baumgartner & Steenkamp, 2001), which means raw Likert averages are not directly comparable across nationalities in Europe's multinational classrooms. This article explains the evidence and how to evaluate fairly across diverse cohorts.