Is Student Evaluation of Teaching Valid? What the State-of-the-Art Review Concludes
A close reading of Spooren, Brockx and Mortelmans'' 2013 Review of Educational Research synthesis of SET validity — what it actually concludes, the evidence behind it, and what it means for how universities use student ratings.
Koji Education Team
Product
In brief: The most systematic review of the field — Spooren, Brockx and Mortelmans (2013), On the Validity of Student Evaluation of Teaching: The State of the Art (Review of Educational Research) — concludes that student evaluation of teaching (SET) is neither the worthless popularity contest its critics claim nor the precise measuring instrument its defenders assume. SET scores carry genuine signal about some dimensions of teaching, but their validity is conditional, multidimensional, and easily undermined by how institutions collect and use them. The honest answer to "Are course evaluations valid?" is: valid for what, measured how, used for which decision?
What the research says
For decades the SET literature has been a battlefield of duelling single studies — one showing ratings correlate with learning, the next showing they correlate only with grades. Pieter Spooren, Bert Brockx and Dimitri Mortelmans, then at the University of Antwerp, set out to cut through the noise. Their 2013 paper in Review of Educational Research (one of the highest-impact venues in education) is a narrative synthesis of peer-reviewed SET research published since roughly 2000, organised around a validation framework rather than a simple "biased / not biased" verdict.
Their central move is to treat validity not as a property a questionnaire either has or lacks, but as an argument that must be assembled from several kinds of evidence. Borrowing from modern validity theory (Messick; the Standards for Educational and Psychological Testing), they examine:
- Construct validity — does SET measure the latent construct "teaching effectiveness", and is that construct multidimensional? They side with Herbert Marsh''s long programme of work (e.g., the SEEQ instrument) showing that student ratings are reliable and multidimensional, capturing distinguishable factors such as clarity, organisation, enthusiasm, and workload — not a single undifferentiated "good teacher" blob.
- Criterion validity — do ratings predict an external criterion such as student learning? Here the evidence is mixed and shrinking. Older meta-analyses (Cohen, 1981; Feldman) reported moderate positive correlations between ratings and achievement in multisection studies; later reanalyses cast doubt on their robustness.
- Threats to validity — they catalogue the now-familiar list of potential biasing variables (grading leniency, course difficulty, class size, discipline, instructor gender) and conclude the evidence is inconsistent: effects appear in some designs and vanish in others, and few are large enough to overturn a rating on their own.
Crucially, Spooren et al. argue the field has been asking a poorly specified question. "Is SET valid?" has no answer until you say which inference you want to license. Using ratings to give an instructor formative feedback on clarity is a very different validity claim from using a single global mean to deny tenure.
Two later works sharpen the same point from opposite directions. Uttl, White and Gonzalez (2017), re-meta-analysing the multisection literature with proper attention to study size, found that once small-sample studies are weighted correctly the SET–learning correlation is essentially zero — a strong corroboration of Spooren et al.''s caution about criterion validity. By contrast, Benton and Cashin''s (2012) IDEA review reaches a more sanguine verdict, treating ratings as a useful if imperfect source. The disagreement is itself the finding: reasonable syntheses diverge because the underlying studies measure different things in different ways.
Why it matters for course evaluation in practice
The state-of-the-art review hands quality-assurance offices three operational lessons.
1. Match the inference to the evidence. Spooren et al.''s validity-as-argument framing maps directly onto QA practice. Formative use (helping an instructor improve) rests on the strongest part of the evidence base — the multidimensional construct validity Marsh documented. Summative, high-stakes use (promotion, contract renewal, ranking) leans on criterion validity, which is exactly the weakest link. An institution that uses the same number for both is over-claiming on the high-stakes side.
2. Multidimensionality is a feature, not noise. Because ratings reliably distinguish dimensions of teaching, collapsing them into one "overall satisfaction" score throws away the most defensible information. A course can be well-organised but poorly assessed; a single mean hides exactly the diagnostic detail a programme director needs.
3. Bias is real but conditional. The review''s refusal to declare SET uniformly "biased" is not fence-sitting — it reflects genuinely inconsistent effect sizes. The practical implication is that you cannot wave away an individual result by invoking "bias" in the abstract, nor can you assume your instrument is clean. You have to check your own data for differential functioning rather than import a verdict from a US study of a different cohort.
This is why the modern QA consensus — echoed in the ESG (Standards and Guidelines for Quality Assurance in the European Higher Education Area) — treats student feedback as one triangulated source among several (peer observation, learning analytics, self-evaluation), never as a standalone metric.
Limitations and honest caveats
A PhD reader will and should push back on several fronts.
- It is a narrative review, not a formal meta-analysis. Spooren et al. synthesise and interpret; they do not pool effect sizes with pre-registered inclusion criteria. That makes the review broad and readable but leaves more room for author judgement than a PRISMA-style quantitative synthesis would.
- The literature it reviews is dominated by North American, paper-based, large-enrolment settings. Generalising to small European seminars, multilingual cohorts, or online delivery is an extrapolation, not a direct finding. Response styles and consumer norms differ across systems.
- "Teaching effectiveness" remains an underdetermined criterion. Much of the validity debate is circular because the field lacks an agreed gold-standard measure of learning. Multisection studies use end-of-course exams, which themselves capture only part of what a course is meant to achieve (see our note on value-added learning).
- The review is now over a decade old. It predates the explosion of online evaluation, the gender-bias quasi-experiments of the mid-2010s, and AI-assisted analysis. Its framework endures; some of its specific effect-size readings have been overtaken by later reanalyses.
None of these caveats overturns the core conclusion. They reinforce it: validity is local, conditional, and tied to use.
How Koji incorporates this
Koji is built around exactly the distinction Spooren, Brockx and Mortelmans insist on — that a student evaluation is only as valid as the specific inference it is asked to support. Several mechanisms operationalise the review''s conclusions:
- Preserving multidimensionality. Koji structures evaluations as discrete, theory-driven questions (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) so that distinguishable dimensions — clarity, organisation, assessment, workload — are reported separately rather than averaged into one figure. This is designed to retain the diagnostic, construct-valid signal Marsh identified rather than discard it.
- Probing beyond the Likert number. Because criterion validity is the weakest link, Koji''s AI-moderated conversational interview follows up a rating with "why" and "can you give an example?" The aim is to surface the reasoning behind a score so that QA staff are reading evidence, not just a mean — a direct response to the review''s warning against treating a single number as self-explanatory.
- Bias-aware reporting, locally checked. In line with the finding that bias is conditional rather than universal, Koji''s reporting is designed to flag where score distributions diverge across cohorts so an institution can investigate its own data for differential functioning, instead of importing or dismissing bias claims wholesale.
- Separating formative from summative. Koji supports mid-cycle, formative collection distinct from end-of-term summative reporting, helping institutions avoid the over-claim of using one instrument for both purposes. We frame this as designed to mitigate misuse, not to certify any single score as fit for high-stakes personnel decisions.
The same AI-moderated interview engine powers Koji''s core research platform at koji.so for product and customer research — the underlying method (structured questions plus adaptive probing) is identical; only the context changes.
The bottom line for practice
The enduring contribution of Spooren, Brockx and Mortelmans is procedural, not doctrinal: before you act on a course-evaluation number, name the inference it is meant to support and ask whether the evidence base actually licenses that inference. A score that is perfectly fit for a developmental conversation with an instructor is rarely fit, unaided, for a career decision. Validity is something you argue for one use at a time, not a badge the questionnaire wears for all of them.
Related resources
- Dimensionality of student ratings: global vs profile (d'Apollonia & Abrami)
- What student evaluations measure: Marsh and multidimensionality
- Peer observation vs student evaluations: convergent validity
- Student evaluations, teaching and learning: the meta-analytic record
- Value-added learning vs student evaluations (Carrell & West)
- An evaluation of course evaluations (Stark & Freishtat)
References
- Spooren, P., Brockx, B., & Mortelmans, D. (2013). On the validity of student evaluation of teaching: The state of the art. Review of Educational Research, 83(4), 598–642. https://doi.org/10.3102/0034654313496870
- Marsh, H. W. (2007). Students'' evaluations of university teaching: Dimensionality, reliability, validity, potential biases and usefulness. In R. P. Perry & J. C. Smart (Eds.), The Scholarship of Teaching and Learning in Higher Education (pp. 319–383). Springer. https://doi.org/10.1007/1-4020-5742-3_9
- Uttl, B., White, C. A., & Gonzalez, D. W. (2017). Meta-analysis of faculty''s teaching effectiveness: Student evaluation of teaching ratings and student learning are not related. Studies in Educational Evaluation, 54, 22–42. https://doi.org/10.1016/j.stueduc.2016.08.007
- Benton, S. L., & Cashin, W. E. (2012). Student ratings of teaching: A summary of research and literature (IDEA Paper No. 50). The IDEA Center. https://www.ideaedu.org/idea_papers/student-ratings-of-teaching-a-summary-of-research-and-literature/
- Cohen, P. A. (1981). Student ratings of instruction and student achievement: A meta-analysis of multisection validity studies. Review of Educational Research, 51(3), 281–309. https://doi.org/10.3102/00346543051003281
Related articles
An Evaluation of Course Evaluations: What the Berkeley Statisticians Concluded About Reporting Student Ratings
Two Berkeley statisticians argued that averaging student-rating scores to rank instructors for promotion and tenure should be abandoned for both statistical and substantive reasons. What Stark & Freishtat (2014) actually recommend — and how to report ratings responsibly instead.
What Do Student Evaluations Actually Measure? Marsh, the SEEQ, and the Case for Multidimensional Feedback
Herbert Marsh spent decades showing that well-built student evaluations are multidimensional, reliable, and stable — and that a single global score throws away most of what they can tell you. What his work establishes, where critics push back, and how to design feedback that is actually usable.
Does the Professor Students Rate Highest Teach Them Most? Value-Added Evidence vs Student Ratings
The strongest causal evidence — random-assignment studies that track follow-on course performance — shows the instructors students rate highest are often not the ones whose students learn most over time. What Carrell & West (2010) and Braga et al. (2014) found, and what it means for course evaluation.
One Number or Many? The Dimensionality Debate and How to Use Student Ratings for Personnel Decisions
Should a promotion committee use a single global teaching score or a detailed profile of many dimensions? The 1997 d'Apollonia & Abrami vs Marsh & Roche debate set the terms — and the answer depends on whether the purpose is summative judgement or formative improvement.