New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
analysis-reporting11 min read

Some Raters Are Harsh, Some Items Are Hard: Why Rasch and Many-Facet Models Beat Averaging Course-Evaluation Scores

Averaging Likert scores treats every student as an equally calibrated measuring instrument and every item as equally hard. Rasch and Many-Facet Rasch models do not. We explain the method, the evidence that rater severity and item difficulty distort raw means, the limitations, and how Koji applies the same logic.

Koji Education Team

Product

In brief: When you average a course's Likert ratings, you assume every student rates on the same scale and every item is equally easy to endorse. Both assumptions are false. The Rasch measurement model — and its extension, the Many-Facet Rasch Model (MFRM) — separately estimates how severe or lenient each rater is, how difficult each item is to endorse, and how high the instructor's underlying quality is, placing all three on a single interval scale. A study of 2,553 student evaluations across 145 courses at the University of Cape Coast (Quansah, 2022) found roughly 30% of students could not discriminate between performance levels or showed halo effects, that items were not of equal difficulty, and that some students' ratings were simply invalid. The practical conclusion: raw arithmetic means can misrank instructors, and a measurement model is the disciplined alternative — but it is data-hungry and is itself only as good as its assumptions.

What the research says

Classical Test Theory (CTT) — the framework behind a raw mean and Cronbach's alpha — treats a course-evaluation score as the sum or average of item responses, implicitly assuming that all items are interchangeable and that the rating scale behaves identically for every respondent. The Rasch model, from the work of Georg Rasch and elaborated for rating scales by Andrich (1978), A rating formulation for ordered response categories (Psychometrika, 43(4), 561–573), takes a different stance. It models the probability that a respondent endorses a given category of an item as a function of two things on the same logit (log-odds) scale: the respondent's (or object's) location on the latent trait, and the item's difficulty. If the data fit the model, you obtain genuinely interval measures rather than ordinal counts dressed up as numbers — and you can tell which items and which respondents misfit.

The Many-Facet Rasch Model (MFRM), developed by Linacre (1989, Many-Facet Rasch Measurement), generalises this to performance assessment, where a third (or fourth) facet enters: the rater. In a course-evaluation context the facets are typically the instructor (the object being measured), the student (the rater), and the item. MFRM estimates, simultaneously, how good each instructor is, how severe or lenient each student-rater is, and how hard each item is to endorse — and then produces fair measures of instructor quality that are adjusted for the particular mix of raters and items each instructor happened to receive. Two instructors with identical raw means can have different fair scores if one was rated by systematically harsher students. Bond and Fox (2015), Applying the Rasch Model, provide the standard applied treatment of how to fit, diagnose, and interpret these models.

The most directly relevant empirical anchor is Quansah (2022), Item and rater variabilities in students' evaluation of teaching in a university in Ghana: Application of Many-Facet Rasch Model (Heliyon, 8(12), e12548). Using 2,553 student responses across 145 courses at the University of Cape Coast, with a 13-item evaluation form analysed under a partial-credit MFRM, Quansah reported several findings that should worry anyone who averages raw scores:

  • Raters were not interchangeable. About 30.5% of participating students either could not discriminate between different performance levels or exhibited halo effects (rating every item nearly identically). Students showed overall leniency, with an observed mean rating of 3.63 against a model-expected 3.51.
  • Items were not equally difficult to endorse. A chi-square test rejected equal item difficulty (approximately χ²(12) = 72.1, p < .001); roughly 61% of items were problematic, with several showing misfit and a few overfitting.
  • Some scores were simply invalid. Because of inconsistent rating behaviour, non-functional rating-scale categories, and defective items, the study concluded that data from students' appraisal of lecturers should be used with caution — not aggregated naively into a personnel-grade number.

A separate methodological strand (e.g., IRT-based validity frameworks for teaching evaluation) makes the same point from the other direction: Rasch analysis can exclude the influence of rater severity and scale artefacts that are unrelated to a teacher's actual quality, allowing comparison of "true" abilities rather than contaminated averages.

Why it matters for course evaluation in practice

For an institutional researcher or evaluation committee, the message is concrete:

  • Raw means can misrank. If instructor A's students are systematically harsher than instructor B's, A's lower raw mean may reflect who rated them, not how they taught. MFRM fair scores adjust for this; arithmetic means cannot. This compounds the misclassification problems already documented for SET when scores are used for ranking.
  • The rating scale may not be working. Rasch category diagnostics reveal when adjacent scale points (say, "3" and "4") are not functioning as ordered, distinct categories — meaning your 5-point scale is effectively a 3-point scale and your decimals are noise.
  • Bad items become visible. Misfit statistics flag items that do not cohere with the rest of the instrument (often double-barrelled, ambiguous, or off-construct items), giving a principled basis for revising the form rather than guessing.
  • Halo and straight-lining are detectable. Respondents who rate every item identically contribute little information; MFRM can identify and down-weight or flag them, improving the signal in what remains.

Limitations and honest caveats

Rasch is a discipline, not a magic wand, and a sophisticated reader will raise several objections.

  • It is data-hungry. Stable facet estimates require enough ratings per instructor and per item. For a small seminar with eight respondents, MFRM estimates are imprecise — the very situation where reliable measurement is hardest. Rasch does not manufacture information that the data lack.
  • Fit is an assumption, not a guarantee. The interval-scale claim holds only if the data fit the model. When they do not (and course-evaluation data frequently misfit), you face a choice between modifying the model, dropping items/raters, or admitting the construct is messier than a unidimensional trait. Discarding misfitting data can itself introduce bias.
  • Unidimensionality may be wrong. Standard Rasch assumes one latent dimension. Teaching quality, as the multidimensionality literature (Marsh, d'Apollonia and Abrami) argues, is plausibly several dimensions; forcing it onto one logit scale can obscure the very profile information a programme needs.
  • "Adjusting away" rater severity is a modelling decision, not a fact. If harsh ratings reflect a real difference in what those students experienced (not mere stylistic severity), MFRM's adjustment could erase genuine signal. The model cannot tell stylistic severity from substantive difference without further assumptions.
  • Interpretability and trust. Logit-scaled fair scores are harder to explain to deans and faculty than a familiar "4.2 out of 5." Sophistication can reduce transparency, and transparency matters for legitimacy.

The honest framing: Rasch/MFRM is the right tool for diagnosing and partially correcting rater and item artefacts, and the wrong tool for pretending small, noisy course samples can yield precise instructor rankings.

How Koji incorporates this

Koji for Education does not ask evaluation committees to read logit tables, but it is built on the same measurement principle Rasch formalises: a rating is a function of the rater and the item, not just the object.

  • Designing items that behave. Many Rasch misfit problems are design problems — double-barrelled, vague, or off-construct items that no statistic can rescue after the fact. Koji's structured question types (scale, single_choice, multiple_choice, ranking, yes_no, open_ended) and its emphasis on single-barrelled, low-inference wording are designed to produce items that would fit a measurement model, reducing the misfit Quansah found in 61% of items.
  • Detecting halo and non-discriminating raters. The ~30% of students who rate every item identically are exactly the straight-liners and halo-raters MFRM flags. Koji's AI-moderated conversational interview is designed to interrupt straight-lining by asking for specifics ("You rated everything highly — was there anything that did not work as well?"), recovering discriminating information that a static grid loses.
  • Reporting uncertainty, not false precision. Consistent with the data-hungriness limitation, Koji is designed to report small-sample results with appropriate uncertainty and to discourage ranking instructors on tiny, noisy samples — echoing the "use with caution" conclusion rather than printing spurious decimals.
  • Triangulation over single-source adjustment. Because "adjusting away" rater severity can erase real signal, Koji's approach is to triangulate across cohorts, items, and qualitative evidence rather than rely on a single statistical correction to declare an instructor good or bad.
  • Bias-aware aggregation. Koji's reporting is designed to surface when an instructor's score is being driven by an unusual rater mix or a malfunctioning item, rather than presenting one mean as ground truth.

These are mechanisms designed to mitigate rater- and item-level distortion; they do not turn eight responses into a precise measurement. Koji's core research platform at koji.so applies the same measurement-aware question design to product and customer research, where rater severity and item difficulty distort raw means just as they do in the classroom.

Related Resources

References

  1. Quansah, F. (2022). Item and rater variabilities in students' evaluation of teaching in a university in Ghana: Application of Many-Facet Rasch Model. Heliyon, 8(12), e12548. https://doi.org/10.1016/j.heliyon.2022.e12548
  2. Andrich, D. (1978). A rating formulation for ordered response categories. Psychometrika, 43(4), 561–573. https://doi.org/10.1007/BF02293814
  3. Linacre, J. M. (1989). Many-Facet Rasch Measurement. Chicago: MESA Press.
  4. Bond, T. G., & Fox, C. M. (2015). Applying the Rasch Model: Fundamental Measurement in the Human Sciences (3rd ed.). New York: Routledge.

Related articles

analysis-reporting

Generalizability Theory and the Reliability of Student Ratings: Why One Class Is Not Enough

Reliability is not one number. Generalizability theory (Gillmore, Kane & Naccarato 1978; Marsh 1984) decomposes the variance in student ratings into student, teacher, course and occasion components — and shows that a single class can be reliable for the course yet a poor estimate of the teacher. What that means for fair evaluation.

analysis-reporting

Measurement Invariance: Can You Compare Course-Evaluation Scores Across Groups at All?

Comparing average ratings across departments, languages or online vs paper assumes the questionnaire means the same thing to every group. Measurement invariance is the test that assumption usually fails — and almost nobody runs it.

analysis-reporting

Cronbach's Alpha and Course-Evaluation Reliability: What It Does and Doesn't Tell You

A high Cronbach's alpha on your course-evaluation instrument is widely read as proof of a "reliable" survey. The psychometric literature says that reading is largely wrong. What alpha actually measures, where it misleads, and what to report instead.

analysis-reporting

The 4.2 vs 4.4 Trap: Why Small Differences in Course-Evaluation Means Are Usually Noise

Faculty and administrators routinely read meaning into tiny gaps between course-evaluation means. Boysen (2015) and Boysen et al. (2014) show this happens even when confidence intervals say the difference is nothing — and that warnings barely help. How to report uncertainty honestly.