Do Student Evaluations Measure Learning? The Uttl Meta-Analysis Revisited
A research-grounded reading of the Uttl, White & Gonzalez (2017) meta-analysis on the SET–learning relationship, with implications for European course evaluation and quality-assurance policy.
Koji Education Team
Product
Answer in brief
The best current meta-analytic evidence — Uttl, White, and Gonzalez (2017) — finds that student evaluations of teaching (SET) explain at most about 1% of the variance in objective student learning, once small-sample artefacts and publication bias are properly corrected. SET should therefore not be treated as a measure of teaching effectiveness for high-stakes decisions; it is a measure of student perception of the course experience, which is informative but distinct. European quality-assurance offices should pair SET with evidence of learning outcomes and peer review before drawing conclusions about teaching quality.
What the research says
The anchor paper for this brief is Uttl, B., White, C. A., & Gonzalez, D. W. (2017). "Meta-analysis of faculty's teaching effectiveness: Student evaluation of teaching ratings and student learning are not related." Studies in Educational Evaluation, 54, 22–42. The authors re-examined the classic multisection validity studies — designs in which multiple sections of the same course, taught by different instructors but using a common final exam, are used to test whether sections that rated their instructor higher also learned more.
The traditional reading, established by Cohen (1981) and reinforced by Feldman through the 1990s, was that section-average SET correlates moderately with section-average exam performance (r ≈ 0.43 in Cohen's 1981 meta-analysis), and that this correlation justifies SET as a validity criterion. Uttl et al. (2017) re-coded the multisection studies and added more recent ones, then applied corrections for small-sample bias and publication bias. They found that the apparent SET–learning correlation collapsed: the largest, most carefully designed studies show effectively no relationship, and the weighted estimate is close to zero. Their headline interpretation — that SET ratings explain "at most 1%" of variance in measured learning — has been heavily cited and is, eight years later, the most defensible position from the multisection literature.
Uttl's work is corroborated and extended by several adjacent lines of evidence.
Boring, Ottoboni, and Stark (2016), "Student evaluations of teaching (mostly) do not measure teaching effectiveness," published in ScienceOpen Research, re-analysed two large datasets and showed that the SET–learning relationship is weak and depends on grading leniency and student characteristics. Stark and Freishtat (2014), "An evaluation of course evaluations," an open-access methodological critique, made the now-well-known argument that taking means of ordinal Likert data violates basic measurement assumptions and that comparisons across instructors are unsupported by the data structure.
The contrasting view — that SET does track some component of teaching quality — has not disappeared. Marsh and Roche (1997) and Marsh (2007) argue that SET captures a meaningful "perceived instructional quality" dimension that has theoretical and practical value, even if it is not a learning meter. Esarey and Valdes (2020), "Unbiased, reliable, and valid student evaluations can still be unfair," Assessment & Evaluation in Higher Education (45(8), 1106–1120), offers an important nuance: even granting that SET measures something real, using it for individual-level decisions produces unacceptable error rates because percentile thresholds amplify the noise in the underlying ratings.
The synthesis: there is no robust evidence that SET means correlate with objective learning at a magnitude that would justify high-stakes decisions, and there is robust evidence that SET captures perception (which is a legitimate but different construct).
Why it matters for course evaluation in practice
For a Dutch, German, or Nordic quality-assurance office, the practical implications are sharp.
First, do not call SET a measure of teaching effectiveness. The phrasing matters. SET measures perceived course experience. Treating it as a proxy for teaching effectiveness misrepresents the evidence and creates legal and reputational exposure when individuals are disadvantaged on the basis of that label. Internal documentation, accreditation submissions, and HR processes should describe SET accurately.
Second, separate formative from summative use. SET is empirically useful as input to a teacher's own improvement cycle (this is the use case where Cohen 1980 and Penny & Coe 2004 find measurable instructional gains; see Mid-Semester Feedback and the Power of Consultation). It is empirically problematic as a primary signal in tenure, promotion, or contract renewal.
Third, triangulate with learning evidence. Where SET is part of a teaching-quality argument, pair it with at least one independent learning indicator: course-level achievement of intended learning outcomes (ILOs), exit-test or progression data, peer observation, and external examiner reports. The ENQA Standards and Guidelines for Quality Assurance in the European Higher Education Area (ESG 2015) makes this triangulation an institutional expectation under standards 1.2 (design and approval of programmes) and 1.5 (teaching staff).
Fourth, beware threshold rules. Many institutional policies trigger investigation if a course or instructor falls below a percentile or absolute SET cut-off. Esarey and Valdes (2020) show that even when SET is unbiased and reliable, percentile thresholds produce false-positive rates above 25% for individuals — meaning a large share of instructors flagged as "underperforming" by SET alone are not actually underperforming. Policies that escalate cases on SET alone are likely to mis-identify problems and miss real ones.
Limitations and honest caveats
The Uttl et al. (2017) meta-analysis has critics, and it is important to surface them rather than nod past them.
The authors used a re-coding of the multisection database that some methodologists argue is too aggressive in its publication-bias corrections; if you apply different corrections, the SET–learning correlation does not vanish entirely but settles at a small positive value (r in the range 0.1–0.2). The multisection design itself has limitations: it requires a common final exam, which only certain disciplines (introductory STEM, economics, languages) typically have, and the result may not generalise to seminar or design-studio teaching where there is no common exam. Multisection studies also generally compare instructors of the same course in the same term — a strong design for ruling out cohort effects but a narrow slice of the teaching that institutions evaluate.
It is also worth being honest that "learning" in these studies usually means short-term recall measured by an end-of-course exam. The relationship between SET and longer-term, transfer-rich learning is much harder to measure and remains an open question.
None of these caveats restore SET as a reliable proxy for teaching effectiveness. The defensible reading is the one above: SET is informative about experience, not about learning, and should not be the primary evidence in high-stakes decisions about individuals.
How Koji incorporates this
Koji for Education (edu.koji.so) is built on the position that what students can describe well is their experience — and that the right instrument for that is a conversation, not a Likert scale. We translate the SET–learning literature into product decisions in four ways.
Concept separation in reporting. Koji's reports never label aggregated student ratings as "teaching effectiveness." We surface three distinct constructs: perceived course experience (what students rated and described), behaviour-level signal (what specific instructor and design behaviours students named), and learning outcomes (where the institution supplies grade, ILO, or exit-test data). The data model treats these as separate quantities that can be triangulated rather than collapsed.
Conversational depth over Likert collapse. Koji's AI-moderated interviews probe the why behind a rating, so a 4/5 means something concrete (clear lectures, useful problem sets) rather than a stereotype-loaded global judgment. This is designed to mitigate exactly the construct-validity problem that Uttl et al. (2017) and Stark & Freishtat (2014) raise about Likert means.
ILO-aligned questions. Evaluation design in Koji supports structured items (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) that can be aligned to specific intended learning outcomes of the module. Programme committees can then read student responses against the ILOs rather than against an abstract notion of "good teaching." This is consistent with ESG 1.2 and 1.5 expectations.
Quality scoring and bias-aware reporting. Koji scores response quality (length, depth, specificity) and flags concentrations of generic or stereotype-loaded language. This is not a magic bias-removal tool — it is a way to surface to reviewers which parts of the dataset can carry inferential weight and which cannot.
The central design choice: Koji is not trying to build a "better SET scalar." It is trying to give QA teams the qualitative and structured evidence they actually need to support a defensible judgement, alongside the institution's own learning-outcome data. Teams that also run user, learner, or product research outside the course-evaluation context can use the same engine at koji.so.
Related resources
- Gender Bias in Student Evaluations of Teaching: Evidence and Mitigation
- Selection Bias in Course Evaluations: What Goos and Salomons Found
- Mid-Semester Feedback and the Power of Consultation
- Closing the Feedback Loop in Programme Quality Assurance
References
- Uttl, B., White, C. A., & Gonzalez, D. W. (2017). Meta-analysis of faculty's teaching effectiveness: Student evaluation of teaching ratings and student learning are not related. Studies in Educational Evaluation, 54, 22–42. https://doi.org/10.1016/j.stueduc.2016.08.007
- Cohen, P. A. (1981). Student ratings of instruction and student achievement: A meta-analysis of multisection validity studies. Review of Educational Research, 51(3), 281–309. https://doi.org/10.3102/00346543051003281
- Stark, P. B., & Freishtat, R. (2014). An evaluation of course evaluations. ScienceOpen Research. https://doi.org/10.14293/S2199-1006.1.SOR-EDU.AOFRQA.v1
- Esarey, J., & Valdes, N. (2020). Unbiased, reliable, and valid student evaluations can still be unfair. Assessment & Evaluation in Higher Education, 45(8), 1106–1120. https://doi.org/10.1080/02602938.2020.1724875
- Spooren, P., Brockx, B., & Mortelmans, D. (2013). On the validity of student evaluation of teaching: The state of the art. Review of Educational Research, 83(4), 598–642. https://doi.org/10.3102/0034654313496870
Related articles
Gender Bias in Student Evaluations of Teaching: Evidence and Mitigation
A research-grounded synthesis of gender bias in student evaluations of teaching (SET), anchored to MacNell, Driscoll & Hunt (2015) and corroborated by Boring (2017) and Mengel, Sauermann & Zölitz (2019), with practical mitigations for European quality-assurance teams.
Mid-Semester Feedback and the Power of Consultation: What the Meta-Analyses Show
Penny & Coe (2004) and Cohen (1980) on the formative effectiveness of mid-semester student feedback when combined with consultation, and the implications for European course-evaluation design.
Selection Bias in Course Evaluations: What Goos and Salomons Found
A research-grounded reading of Goos & Salomons (2017) on selection bias in online course evaluations, with practical implications for response-rate policy and reporting at European universities.
Racial and Ethnic Bias in Student Evaluations of Teaching: What the Evidence Shows
Peer-reviewed evidence shows students rate instructors of colour systematically lower than White peers on the same teaching. We review Reid (2010), Chávez & Mitchell (2020) and Bavishi et al. (2010), the limitations, and how to design evaluation so race-linked bias does not contaminate quality decisions.