Do Student Ratings Track Real Learning? What Cohen's 1981 Multisection Meta-Analysis Found — and Why the Answer Changed
Cohen (1981) found a moderate positive correlation between student ratings and student achievement in multisection courses — the foundational evidence that ratings have validity. Forty years of reanalysis has since shrunk that correlation toward zero. Here is what the multisection paradigm actually proves, and how to use ratings responsibly given both findings.
Koji Education Team
Product
The short answer
In a 1981 meta-analysis of multisection validity studies — the cleanest natural experiment available for testing whether student ratings reflect learning — Peter Cohen found a moderate positive correlation between how students rated an instructor and how much that instructor's sections actually learned, measured by a common final exam. The average correlation between an overall instructor rating and achievement was about r = .43, and for an overall course rating about r = .47. For two decades this was the single strongest piece of evidence that student evaluations of teaching (SET) measure something real about teaching effectiveness.
That conclusion has not survived replication intact. When Uttl, White and Gonzalez (2017) re-ran the multisection meta-analysis controlling for section size — and when Uttl, Cnudde and White (2019) sorted the studies by the authors' conflicts of interest — the correlation collapsed toward zero, especially in studies published after 1981. The honest position today is that the multisection design remains the best test we have, that Cohen's synthesis was methodologically careful for its era, and that the effect is much smaller and more fragile than the .43 headline suggests. That nuance — not a slogan in either direction — is what a course-evaluation programme should be built on.
What the research says
The multisection design, and why it matters
Most attempts to validate student ratings against learning are hopelessly confounded: different courses, different students, different exams, different content. The multisection validity study removes most of those confounds. A single large course is split into many sections that share the same syllabus, the same textbook, and crucially the same final examination, but are taught by different instructors. Students are assigned to sections in a way that approximates randomness (timetable, registration order). You then ask: do the sections whose instructor got higher student ratings also score higher on the common exam?
Because the exam is identical and the curriculum is fixed, a positive rating–achievement correlation is hard to dismiss as an artifact of easier content or a friendlier test. This is why Cohen — and his critics — treated multisection studies as the gold standard.
Cohen's 1981 synthesis
Cohen, P. A. (1981), Student Ratings of Instruction and Student Achievement: A Meta-Analysis of Multisection Validity Studies, published in Review of Educational Research, pooled 41 independent validity studies covering 68 separate multisection courses. Using the then-new techniques of meta-analysis, he reported:
- Overall instructor rating × achievement: mean r ≈ .43
- Overall course rating × achievement: mean r ≈ .47
- Specific dimensions varied: ratings of Skill and Structure/Organization correlated most strongly with achievement; ratings of dimensions like rapport or interaction correlated more weakly.
- Moderators mattered: correlations were larger for full-time faculty, when students knew their final grades before rating the instructor, and when an external evaluator (not the instructor) scored the achievement test.
Cohen's reading was cautiously positive: student ratings, particularly global ones and those tapping clarity and organization, carry genuine information about how much students learn. Feldman (1989), extending the synthesis to the dimension level in Research in Higher Education, reached a compatible conclusion — specific instructional behaviours such as clarity, preparation and stimulation of interest were the ones most associated with achievement.
What replication did to the number
The story does not end in 1981. Two findings reframe it:
-
Section size and small-sample noise. Uttl, White and Gonzalez (2017), Meta-analysis of faculty's teaching effectiveness (Studies in Educational Evaluation), pointed out that many multisection studies have very few sections — sometimes a handful — so the per-section correlations are estimated on tiny samples and are extraordinarily noisy. When they weighted properly and accounted for small-section artifacts, the SET–learning correlation shrank toward zero. Their blunt summary: students do not learn more from professors with higher ratings.
-
Publication era and conflict of interest. Uttl, Cnudde and White (2019), Conflict of interest explains the size of student evaluation of teaching and learning correlations in multisection studies (PeerJ, DOI 10.7717/peerj.7225), found that studies published before 1981 reported r ≈ .31, while those published in 1981 and after reported r ≈ .06 — essentially nothing. They also found larger correlations in studies authored by people with a stake in SET (administrators, evaluation-unit staff, vendors).
The takeaway is not that Cohen was wrong to compute what he computed; it is that the multisection literature is heterogeneous, small-sampled, and time-dependent, and that the most defensible modern estimate of the rating–learning link is small and positive at best.
Why it matters for course evaluation in practice
For a quality-assurance office, the Cohen-to-Uttl arc carries three practical lessons.
1. Ratings are not a learning thermometer. Whatever your students' satisfaction scores say, they are at most a weak proxy for how much was learned. A 4.6/5 instructor is not demonstrably teaching more than a 4.1/5 instructor. Treating ratings as a direct measure of learning outcomes — the kind of inference an accreditation panel might wish you could make — is not supported by the strongest available design.
2. The content of the rating matters more than the overall number. Both Cohen and Feldman found that the dimensions most tied to achievement were clarity, organization/structure, and skill — not global likability or rapport. If you must use ratings as one input, weight the items that ask about concrete instructional behaviour, and discount global "overall, this was a great course" items, which behave more like satisfaction.
3. Use ratings for the inferences they support, and triangulate for the rest. Ratings are reasonable evidence of the student experience — clarity, workload, accessibility, perceived support. They are weak evidence of learning gain. For learning, you need other instruments: aligned assessment data, peer observation, and direct outcome measures. This is exactly the logic accreditation frameworks such as the ESG and ENQA standards encode when they ask for multiple lines of evidence rather than a single survey mean.
Limitations and honest caveats
A PhD reader will (rightly) push back, and the pushback runs in both directions.
- The multisection design is internally strong but externally narrow. Multisection courses are overwhelmingly large introductory courses (intro psychology, calculus, economics) with standardized exams. Whether the rating–learning relationship in a 600-student intro economics course generalizes to a 12-student final-year seminar with a dissertation deliverable is genuinely unknown.
- A common final exam is itself construct-narrow. It captures the knowledge the exam samples, usually lower-order recall and procedure, on a single occasion. Deep, durable, or transferable learning — the thing universities claim to produce — is not what most of these exams measure. An instructor who builds long-run understanding at the expense of exam-week performance would look worse on this criterion, a possibility the later causal studies (Carrell & West 2010; Braga et al. 2014) make concrete.
- Heterogeneity is large. Pooling 68 courses hides real variation; the average r is not a law of nature.
- The 2019 conflict-of-interest finding is a correlation about authors, not a randomized audit. It is suggestive of bias in the literature, not definitive proof that every pre-1981 estimate is inflated.
None of this means ratings are worthless. It means the validity of a rating depends entirely on the inference you draw from it — a point we develop in our companion piece on argument-based validity.
How Koji incorporates this
Koji is an AI-native course-evaluation platform, and the Cohen literature shapes how it is designed to be used — not as a learning meter, but as a structured instrument for the inferences ratings can actually support.
- Behaviour-specific, low-inference items over global likability. Because the dimensions Cohen and Feldman linked to achievement were clarity, structure and skill, Koji supports structured question types (
scale,single_choice,multiple_choice,ranking,yes_no,open_ended) that let programmes ask about concrete teaching behaviours ("Were the learning objectives stated for each session?") rather than relying on a single overall mean. - AI-moderated conversational interviews that probe the number. A Likert score is silent about why. Koji's AI moderator follows up on a rating in the respondent's own words — turning "I gave clarity a 3" into a specific, codeable account of where the course lost them. This is designed to recover the construct-relevant signal that a bare global rating obscures.
- Triangulation across cohorts and sources. Koji is built to combine course-level feedback with mid-cycle formative collection and cross-cohort comparison, so a programme is not resting a personnel or curriculum decision on one noisy mean — exactly the small-sample fragility the multisection reanalyses exposed.
- Bias-aware, uncertainty-honest reporting. Rather than ranking instructors on raw averages to two decimal places — which Cohen's own moderator analysis and the later misclassification literature warn against — Koji's reporting is designed to surface distributions, response counts and confidence, not false precision.
Koji is careful to frame these as mitigations, not cures: no instrument turns a satisfaction survey into a measure of learning. For teams whose evaluation needs extend beyond teaching — product, customer and user research — Koji's core research platform at koji.so applies the same AI-moderated interview engine to those questions.
Related Resources
- Do Student Evaluations Measure Learning? The Uttl Meta-Analysis Revisited
- Does the Professor Students Rate Highest Teach Them Most? Value-Added Evidence vs Student Ratings
- Students Rate the Classes Where They Learn the Most Lower: The Feeling-of-Learning Gap
- What Do Student Evaluations Actually Measure? Marsh, the SEEQ, and the Case for Multidimensional Feedback
- Stop Asking "Was the Lecturer Clear?": The Case for Low-Inference Teaching-Behaviour Items
- Is Student Evaluation of Teaching Valid? What the State-of-the-Art Review Concludes
References
- Cohen, P. A. (1981). Student ratings of instruction and student achievement: A meta-analysis of multisection validity studies. Review of Educational Research, 51(3), 281–309. https://doi.org/10.3102/00346543051003281
- Feldman, K. A. (1989). The association between student ratings of specific instructional dimensions and student achievement: Refining and extending the synthesis of data from multisection validity studies. Research in Higher Education, 30(6), 583–645. https://doi.org/10.1007/BF00992392
- Uttl, B., White, C. A., & Gonzalez, D. W. (2017). Meta-analysis of faculty's teaching effectiveness: Student evaluation of teaching ratings and student learning are not related. Studies in Educational Evaluation, 54, 22–42. https://doi.org/10.1016/j.stueduc.2016.08.007
- Uttl, B., Cnudde, K., & White, C. A. (2019). Conflict of interest explains the size of student evaluation of teaching and learning correlations in multisection studies: A meta-analysis. PeerJ, 7, e7225. https://doi.org/10.7717/peerj.7225
- Carrell, S. E., & West, J. E. (2010). Does professor quality matter? Evidence from random assignment of students to professors. Journal of Political Economy, 118(3), 409–432. https://doi.org/10.1086/653808
Related articles
What Do Student Evaluations Actually Measure? Marsh, the SEEQ, and the Case for Multidimensional Feedback
Herbert Marsh spent decades showing that well-built student evaluations are multidimensional, reliable, and stable — and that a single global score throws away most of what they can tell you. What his work establishes, where critics push back, and how to design feedback that is actually usable.
Stop Asking "Was the Lecturer Clear?": The Case for Low-Inference Teaching-Behaviour Items
High-inference items like "Is the instructor clear?" tell teachers nothing about what to change. Murray's low-inference behaviour research shows how concrete, observable items make evaluations diagnostic.
Does the Professor Students Rate Highest Teach Them Most? Value-Added Evidence vs Student Ratings
The strongest causal evidence — random-assignment studies that track follow-on course performance — shows the instructors students rate highest are often not the ones whose students learn most over time. What Carrell & West (2010) and Braga et al. (2014) found, and what it means for course evaluation.
Is Student Evaluation of Teaching Valid? What the State-of-the-Art Review Concludes
A close reading of Spooren, Brockx and Mortelmans'' 2013 Review of Educational Research synthesis of SET validity — what it actually concludes, the evidence behind it, and what it means for how universities use student ratings.