New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods10 min read

What Do Student Evaluations Actually Measure? Marsh, the SEEQ, and the Case for Multidimensional Feedback

Herbert Marsh spent decades showing that well-built student evaluations are multidimensional, reliable, and stable — and that a single global score throws away most of what they can tell you. What his work establishes, where critics push back, and how to design feedback that is actually usable.

Koji Research Desk

Education Research

In short: A single "overall rating" is the least useful thing a course evaluation can produce. Herbert Marsh''s research programme — anchored by Marsh and Roche (1997) in American Psychologist and the Students'' Evaluations of Educational Quality (SEEQ) instrument — established that when evaluations are well constructed, they measure several distinct dimensions of teaching (clarity, enthusiasm, organisation, rapport, assessment, workload, and more), are highly reliable given enough respondents, and are stable over time. The practical lesson is not that evaluations are flawless — later critics dispute their freedom from bias — but that collapsing them to one number discards the diagnostic, improvement-oriented signal that makes them worth collecting at all.

The question this article answers

When a department reads "the instructor scored 4.1 out of 5", what has it actually learned? Marsh''s lifelong argument is: almost nothing actionable. Teaching is not one thing, and a competent evaluation should not pretend it is. This article sets out what the multidimensionality research shows, where the honest disputes lie, and how to build an instrument that yields improvement-ready evidence.

What the research says

Herbert W. Marsh, across a body of work spanning the 1980s and 1990s, developed and validated the SEEQ instrument and assembled the evidence that students'' evaluations of teaching, under appropriate conditions, have four properties worth stating precisely. In Marsh and Roche (1997) — part of a landmark American Psychologist special issue featuring both proponents and critics — these are framed as the questions of validity, bias, and utility.

  1. Multidimensionality. SEEQ distinguishes around nine factors, including Learning/Value, Instructor Enthusiasm, Organisation/Clarity, Group Interaction, Individual Rapport, Breadth of Coverage, Examinations/Grading, Assignments/Readings, and Workload/Difficulty. Crucially, these dimensions are separable: an instructor can be highly organised but low on rapport, or enthusiastic but assessed as setting an unclear workload. A global score blends them into mud.
  2. Reliability. With a sufficient number of student raters per class, SEEQ-type instruments reach high reliability — Marsh reported reliabilities around 0.95 with roughly 20–25 raters, falling sharply with small classes. Reliability is therefore a property of how many students respond, not an intrinsic guarantee.
  3. Stability and relative validity. Marsh showed evaluations are reasonably stable across time and that profiles of the same instructor are consistent. He argued they correlate with other indicators of effective teaching (e.g., student learning, instructor self-evaluation, trained-observer ratings) more than critics allowed.
  4. Relative robustness to some confounds. Marsh''s position was that, properly measured, the much-feared influences of grading leniency, class size, and workload were smaller than folklore assumed — a claim that remains contested (see caveats).

Marsh and Roche''s deeper point was about utility: evaluations only improve teaching if they are fed back diagnostically and paired with consultation. Handing an instructor a number changes nothing; handing them a dimensional profile, plus a conversation about what to do, measurably improves subsequent teaching.

Corroborating and contrasting work

The multidimensional structure is one of the better-replicated findings in the field. Spooren, Brockx and Mortelmans (2013), in their Review of Educational Research synthesis "On the validity of student evaluation of teaching: the state of the art", concluded that the construct validity of well-designed SET instruments is reasonably supported, while cautioning that consequential and criterion validity — whether scores track learning and survive high-stakes use — are far weaker. That is the productive tension: Marsh established what evaluations can measure reliably (dimensions of the teaching experience); later work questions whether those dimensions equal learning or justify summative use. Both can be true. Uttl, White and Gonzalez (2017) sharpen the contrast, finding the evaluation–learning correlation near zero once prior ability is controlled — a direct challenge to the strongest validity claims, though not to the multidimensionality finding itself.

Why it matters for course evaluation in practice

Three design implications follow directly:

  • Measure dimensions, not a verdict. If your instrument asks mostly global "rate this course/instructor overall" items, you are collecting the one signal that is least diagnostic and most vulnerable to the halo effect. Ask about clarity, pacing, feedback quality, workload appropriateness, and rapport separately.
  • Respondent count is a validity input, not an afterthought. Because reliability scales with raters, a beautiful instrument in a class of eight tells you little. Response-rate strategy is part of measurement quality, not just logistics.
  • Feedback without consultation wastes the data. Marsh and Roche''s utility finding means the evaluation is only half the system; the other half is structured follow-up that turns a profile into a change.

Limitations and honest caveats

A PhD reader will rightly note that Marsh is the field''s most prominent proponent, and his robustness claims are the most disputed part of his legacy.

  • The bias debate is unsettled. Experimental work since the 1990s — MacNell et al. (2015) on gender, Boring (2017), and others — finds bias effects that Marsh''s correlational, instrument-focused tradition tended to minimise. Multidimensionality being real does not make the dimensions bias-free.
  • Construct validity ≠ criterion validity. That evaluations reliably measure perceived organisation or enthusiasm does not establish that high scores mean more learning. The Uttl meta-analysis is the strongest counterweight here.
  • Reliability is conditional and often unmet. The 0.95 figure assumes many raters; real-world low response rates routinely leave classes below the threshold where the instrument behaves as validated.
  • Generalisability across cultures and disciplines. SEEQ was validated largely in specific Western contexts; response styles and disciplinary norms (see our cross-cultural and discipline-bias articles) complicate transport.

The fair synthesis: Marsh established that good evaluations are structured, multidimensional, and reliable measurement instruments of the teaching experience — which is exactly why reducing them to one number is indefensible — without settling the separate question of whether they measure learning or should drive personnel decisions.

How Koji incorporates this

Koji is built around the multidimensional, utility-first picture Marsh''s work supports.

  • Dimension-structured instruments. Koji studies combine structured question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) so that distinct facets of teaching — clarity, feedback, workload, rapport — are captured and reported separately, not blended into one average.
  • Probing beyond the number. Koji''s AI-moderated conversational interviews ask why a dimension was rated as it was, recovering the specific, actionable content a Likert item compresses away — the difference between "Organisation: 3.4" and "students could not find the assessment criteria until week 6".
  • Automatic thematic analysis. Open-text and interview responses are clustered into themes that map onto teaching dimensions, giving departments a profile rather than a verdict.
  • Reliability-aware reporting. Because Marsh''s reliability depends on respondent numbers, Koji is designed to surface response counts alongside results and to caution where a dimension rests on too few responses.
  • Closing the loop / consultation built in. Marsh and Roche''s utility finding — feedback only works with follow-up — maps onto Koji''s mid-cycle formative collection and action-tracking, so a dimensional weakness becomes a tracked change within the same cohort rather than a number filed away.

Koji is designed to mitigate the single-number problem Marsh identified, not to claim it resolves the deeper validity debates. Teams running broader research will find the same multidimensional, probe-the-why engine in Koji''s core platform at koji.so.

What a multidimensional instrument looks like in practice

Translating Marsh into a working questionnaire means resisting the temptation to ask "rate this course overall" and little else. A dimensional instrument asks separately about the facets students can actually observe and that map onto distinct teaching behaviours. Clarity and organisation: were objectives, structure, and assessment criteria clear and available early? Intellectual stimulation and enthusiasm: did the teaching engage interest and convey the value of the material? Feedback and assessment: was feedback timely, specific, and aligned with what was taught? Workload and pacing: was the effort demanded appropriate and well distributed — Marsh''s "good" versus "bad" workload distinction, where demanding-but-valuable differs sharply from merely heavy? Rapport and support: were students able to get help and feel their questions were welcome?

Each facet should be reported on its own, because the improvement actions they imply are different: a clarity problem is fixed by rewriting a syllabus, a feedback problem by changing assessment turnaround, a rapport problem by restructuring contact time. A blended score points to none of these. The discipline, then, is to design the instrument backwards from the actions you might take, and to treat any global "overall satisfaction" item as a summary headline — useful for trend-spotting, useless for knowing what to change. This is exactly where Marsh''s multidimensionality finding stops being an academic nicety and becomes the difference between an evaluation that improves teaching and one that merely scores it.

Related Resources

References

  • Marsh, H. W., & Roche, L. A. (1997). Making students'' evaluations of teaching effectiveness effective: The critical issues of validity, bias, and utility. American Psychologist, 52(11), 1187–1197. https://doi.org/10.1037/0003-066X.52.11.1187
  • Marsh, H. W. (1987). Students'' evaluations of university teaching: Research findings, methodological issues, and directions for future research. International Journal of Educational Research, 11(3), 253–388. https://doi.org/10.1016/0883-0355(87)90001-2
  • Spooren, P., Brockx, B., & Mortelmans, D. (2013). On the validity of student evaluation of teaching: The state of the art. Review of Educational Research, 83(4), 598–642. https://doi.org/10.3102/0034654313496870
  • Uttl, B., White, C. A., & Gonzalez, D. W. (2017). Meta-analysis of faculty''s teaching effectiveness: Student evaluation of teaching ratings and student learning are not related. Studies in Educational Evaluation, 54, 22–42. https://doi.org/10.1016/j.stueduc.2016.08.007

Related articles

research-methods

Response Styles and Likert Scales: Why Cross-Cultural Evaluation Needs More Than Numbers

Acquiescence and extreme response styles vary systematically by culture (Harzing, 2006; Baumgartner & Steenkamp, 2001), which means raw Likert averages are not directly comparable across nationalities in Europe's multinational classrooms. This article explains the evidence and how to evaluate fairly across diverse cohorts.

research-methods

The Halo Effect in Course Evaluations: When One Impression Colours Every Rating

When students like an instructor, that single global impression bleeds into their ratings of unrelated specifics. From Nisbett & Wilson (1977) to Feistauer & Richter (2018) and Cannon & Cipriani (2021), here is how the halo effect distorts itemised course evaluations — and what to do about it.

research-methods

What Can Open-Text Student Comments Tell You That Likert Scores Cannot?

A research-grounded guide to open-ended student comments in course evaluation: what Alhija and Fresko found about who writes them and what they contain, how thematic analysis surfaces issues numbers miss (Stupans et al.), how rare abusive comments actually are (Tucker), and how to turn free text into reliable evidence.

research-methods

How Many Scale Points Should a Course-Evaluation Question Have?

What the measurement literature — Preston & Colman (2000), Weng (2004), Dawes (2008) — says about the optimal number of response categories on rating scales, and why the answer for course evaluation is not just a number but a question about what a Likert item can and cannot capture.