New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods12 min read

Do Student Evaluations Measure Learning? The Uttl Meta-Analysis Revisited

A research-grounded reading of the Uttl, White & Gonzalez (2017) meta-analysis on the SET–learning relationship, with implications for European course evaluation and quality-assurance policy.

Koji Education Team

Product

Answer in brief

The best current meta-analytic evidence — Uttl, White, and Gonzalez (2017) — finds that student evaluations of teaching (SET) explain at most about 1% of the variance in objective student learning, once small-sample artefacts and publication bias are properly corrected. SET should therefore not be treated as a measure of teaching effectiveness for high-stakes decisions; it is a measure of student perception of the course experience, which is informative but distinct. European quality-assurance offices should pair SET with evidence of learning outcomes and peer review before drawing conclusions about teaching quality.

What the research says

The anchor paper for this brief is Uttl, B., White, C. A., & Gonzalez, D. W. (2017). "Meta-analysis of faculty's teaching effectiveness: Student evaluation of teaching ratings and student learning are not related." Studies in Educational Evaluation, 54, 22–42. The authors re-examined the classic multisection validity studies — designs in which multiple sections of the same course, taught by different instructors but using a common final exam, are used to test whether sections that rated their instructor higher also learned more.

The traditional reading, established by Cohen (1981) and reinforced by Feldman through the 1990s, was that section-average SET correlates moderately with section-average exam performance (r ≈ 0.43 in Cohen's 1981 meta-analysis), and that this correlation justifies SET as a validity criterion. Uttl et al. (2017) re-coded the multisection studies and added more recent ones, then applied corrections for small-sample bias and publication bias. They found that the apparent SET–learning correlation collapsed: the largest, most carefully designed studies show effectively no relationship, and the weighted estimate is close to zero. Their headline interpretation — that SET ratings explain "at most 1%" of variance in measured learning — has been heavily cited and is, eight years later, the most defensible position from the multisection literature.

Uttl's work is corroborated and extended by several adjacent lines of evidence.

Boring, Ottoboni, and Stark (2016), "Student evaluations of teaching (mostly) do not measure teaching effectiveness," published in ScienceOpen Research, re-analysed two large datasets and showed that the SET–learning relationship is weak and depends on grading leniency and student characteristics. Stark and Freishtat (2014), "An evaluation of course evaluations," an open-access methodological critique, made the now-well-known argument that taking means of ordinal Likert data violates basic measurement assumptions and that comparisons across instructors are unsupported by the data structure.

The contrasting view — that SET does track some component of teaching quality — has not disappeared. Marsh and Roche (1997) and Marsh (2007) argue that SET captures a meaningful "perceived instructional quality" dimension that has theoretical and practical value, even if it is not a learning meter. Esarey and Valdes (2020), "Unbiased, reliable, and valid student evaluations can still be unfair," Assessment & Evaluation in Higher Education (45(8), 1106–1120), offers an important nuance: even granting that SET measures something real, using it for individual-level decisions produces unacceptable error rates because percentile thresholds amplify the noise in the underlying ratings.

The synthesis: there is no robust evidence that SET means correlate with objective learning at a magnitude that would justify high-stakes decisions, and there is robust evidence that SET captures perception (which is a legitimate but different construct).

Why it matters for course evaluation in practice

For a Dutch, German, or Nordic quality-assurance office, the practical implications are sharp.

First, do not call SET a measure of teaching effectiveness. The phrasing matters. SET measures perceived course experience. Treating it as a proxy for teaching effectiveness misrepresents the evidence and creates legal and reputational exposure when individuals are disadvantaged on the basis of that label. Internal documentation, accreditation submissions, and HR processes should describe SET accurately.

Second, separate formative from summative use. SET is empirically useful as input to a teacher's own improvement cycle (this is the use case where Cohen 1980 and Penny & Coe 2004 find measurable instructional gains; see Mid-Semester Feedback and the Power of Consultation). It is empirically problematic as a primary signal in tenure, promotion, or contract renewal.

Third, triangulate with learning evidence. Where SET is part of a teaching-quality argument, pair it with at least one independent learning indicator: course-level achievement of intended learning outcomes (ILOs), exit-test or progression data, peer observation, and external examiner reports. The ENQA Standards and Guidelines for Quality Assurance in the European Higher Education Area (ESG 2015) makes this triangulation an institutional expectation under standards 1.2 (design and approval of programmes) and 1.5 (teaching staff).

Fourth, beware threshold rules. Many institutional policies trigger investigation if a course or instructor falls below a percentile or absolute SET cut-off. Esarey and Valdes (2020) show that even when SET is unbiased and reliable, percentile thresholds produce false-positive rates above 25% for individuals — meaning a large share of instructors flagged as "underperforming" by SET alone are not actually underperforming. Policies that escalate cases on SET alone are likely to mis-identify problems and miss real ones.

Limitations and honest caveats

The Uttl et al. (2017) meta-analysis has critics, and it is important to surface them rather than nod past them.

The authors used a re-coding of the multisection database that some methodologists argue is too aggressive in its publication-bias corrections; if you apply different corrections, the SET–learning correlation does not vanish entirely but settles at a small positive value (r in the range 0.1–0.2). The multisection design itself has limitations: it requires a common final exam, which only certain disciplines (introductory STEM, economics, languages) typically have, and the result may not generalise to seminar or design-studio teaching where there is no common exam. Multisection studies also generally compare instructors of the same course in the same term — a strong design for ruling out cohort effects but a narrow slice of the teaching that institutions evaluate.

It is also worth being honest that "learning" in these studies usually means short-term recall measured by an end-of-course exam. The relationship between SET and longer-term, transfer-rich learning is much harder to measure and remains an open question.

None of these caveats restore SET as a reliable proxy for teaching effectiveness. The defensible reading is the one above: SET is informative about experience, not about learning, and should not be the primary evidence in high-stakes decisions about individuals.

How Koji incorporates this

Koji for Education (edu.koji.so) is built on the position that what students can describe well is their experience — and that the right instrument for that is a conversation, not a Likert scale. We translate the SET–learning literature into product decisions in four ways.

Concept separation in reporting. Koji's reports never label aggregated student ratings as "teaching effectiveness." We surface three distinct constructs: perceived course experience (what students rated and described), behaviour-level signal (what specific instructor and design behaviours students named), and learning outcomes (where the institution supplies grade, ILO, or exit-test data). The data model treats these as separate quantities that can be triangulated rather than collapsed.

Conversational depth over Likert collapse. Koji's AI-moderated interviews probe the why behind a rating, so a 4/5 means something concrete (clear lectures, useful problem sets) rather than a stereotype-loaded global judgment. This is designed to mitigate exactly the construct-validity problem that Uttl et al. (2017) and Stark & Freishtat (2014) raise about Likert means.

ILO-aligned questions. Evaluation design in Koji supports structured items (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) that can be aligned to specific intended learning outcomes of the module. Programme committees can then read student responses against the ILOs rather than against an abstract notion of "good teaching." This is consistent with ESG 1.2 and 1.5 expectations.

Quality scoring and bias-aware reporting. Koji scores response quality (length, depth, specificity) and flags concentrations of generic or stereotype-loaded language. This is not a magic bias-removal tool — it is a way to surface to reviewers which parts of the dataset can carry inferential weight and which cannot.

The central design choice: Koji is not trying to build a "better SET scalar." It is trying to give QA teams the qualitative and structured evidence they actually need to support a defensible judgement, alongside the institution's own learning-outcome data. Teams that also run user, learner, or product research outside the course-evaluation context can use the same engine at koji.so.

Related resources

References