The Teachers Who Help You Most Get Rated Worst: What the Bocconi Natural Experiment Proved
Braga, Paccagnella and Pellizzari (2014) used near-random assignment of students to professors at Bocconi University to show that the teachers who most improved students' performance in later courses received the lowest student evaluations. Here is the design, the effect sizes, the caveats, and what it means for European quality assurance.
Koji Education Team
Product
The short answer
Using administrative data from Bocconi University in Milan, where first-year students are assigned to teaching sections in a way that is effectively random, Braga, Paccagnella and Pellizzari (2014) estimated each professor's true effectiveness by tracking how that professor's students later performed in follow-on courses taught by other instructors. They then compared this objective effectiveness measure to the students' end-of-course evaluations. The result is one of the most cited and most unsettling findings in the field: teacher effectiveness is negatively correlated with student evaluations. The professors whose students did best in subsequent, more advanced courses tended to receive the worst ratings; the professors who earned the highest ratings were associated with weaker downstream performance.
The most plausible explanation is that students evaluate on realized utility now — how easy, pleasant, and well-graded the immediate course felt — rather than on learning that pays off later. Because the design is quasi-experimental and set in a European institution with fixed syllabuses and common exams, the study is unusually hard to wave away, and it is essential reading for any European QA office tempted to treat evaluation scores as a proxy for teaching quality.
What the research says
A natural experiment that fixes the confounds
The credibility of this study rests on its design. At Bocconi, students in a given degree follow a fixed sequence of compulsory courses with common syllabuses and common exams across sections, and they are allocated to sections (and thus professors) in a way that approximates random assignment. This setup neutralizes the usual confounds:
- Students cannot self-select into easier professors.
- The curriculum and assessment are held constant across sections.
- Because students later take follow-on courses taught by different professors, you can measure how well a first-year professor prepared them — independent of that professor's own grading or popularity.
The authors define a professor's effectiveness as the (residualized) performance of their students in the subsequent related course, after adjusting for student ability and other factors. This is a value-added-style criterion, but measured on a downstream exam the original professor neither sets nor grades — closing the loophole that students might just be rewarding lenient marking.
The headline results
Published in Economics of Education Review (2014, vol. 41, pp. 71–88), the study reports two findings that belong together:
- Professors matter. Even with fixed syllabuses, the variation in downstream student performance attributable to the first-year professor is substantial — teaching quality is real and measurable.
- Evaluations point the wrong way. The correlation between a professor's measured effectiveness and their students' evaluations is negative. Professors who raised follow-on performance received lower evaluations; those who were rated highly were associated with lower subsequent achievement. The negative relationship was especially pronounced in the more demanding, quantitative courses and among the most academically able students, who appear to penalize the very rigour that helps them later.
The authors interpret this through the lens of immediate versus deferred payoff: a professor who makes students work hard, struggle productively, and confront difficulty depresses the in-the-moment experience that drives satisfaction ratings, while building competence that only shows up in the next course.
Convergent evidence
Braga et al. did not produce an isolated curiosity. The pattern echoes across independent designs:
- Carrell and West (2010), using random assignment at the U.S. Air Force Academy (Journal of Political Economy, 118(3)), found that professors who boosted contemporaneous exam performance — and earned higher student evaluations — were associated with worse performance in mandatory follow-on courses. Same shape, different continent and institution.
- Deslauriers et al. and related work on active learning show students often feel they learn less in the classes where they objectively learn more, depressing satisfaction with the more effective pedagogy.
- The broader meta-analytic re-estimation by Uttl, White and Gonzalez (2017) finds essentially no positive rating–learning correlation once methodological artifacts are controlled.
Three different methods — quasi-random assignment in Italy, random assignment in the U.S., and meta-analysis — converge on the same uncomfortable conclusion.
Why it matters for course evaluation in practice
For a European quality-assurance office, the Bocconi study is not an argument to abandon student feedback. It is an argument to stop using satisfaction scores as a proxy for teaching effectiveness, and to be especially wary in exactly the cases where the bias is largest.
1. High scores can be a warning, not a reward. If a demanding, quantitatively rigorous module earns lower satisfaction than a gentle one, the Bocconi evidence says that may reflect better preparation for what comes next — not worse teaching. Ranking instructors or modules by raw evaluation means can systematically penalize the rigour your programme exists to deliver.
2. The bias is patterned, so target it. Because the negative relationship concentrates in hard, quantitative courses and among strong students, those are precisely the contexts where you should refuse to read satisfaction as quality. STEM and quantitative-social-science programmes are most exposed.
3. Separate the two questions you are actually asking. "Did students have a good experience?" and "Did students learn what they will need next?" are different questions with different right instruments. Student evaluations answer the first reasonably well and the second badly. For the second, you need downstream and outcome evidence — progression data, performance in follow-on modules, aligned assessment.
Limitations and honest caveats
A careful reader should not over-read a single study, even a strong one.
- "Subsequent-course performance" is a proxy for learning, not learning itself. It captures what the next exam rewards. If both courses over-weight narrow, examinable skills, the criterion inherits that narrowness.
- External validity is bounded. Bocconi is a selective, quantitatively oriented Italian business-and-economics university with fixed syllabuses and common exams. The near-random assignment that makes the design work is precisely what makes it unrepresentative of, say, a small humanities seminar with elective enrolment. The effect's sign is corroborated elsewhere, but its size may not transfer.
- Negative on average is not negative everywhere. The relationship is a tendency across professors, not a guarantee that every well-liked instructor is secretly ineffective. Plenty of teaching is both effective and enjoyable.
- It does not show evaluations are useless. It shows they are a poor proxy for deferred learning. They remain informative about clarity, organization, accessibility and the student experience — the things students are well positioned to observe.
How Koji incorporates this
Koji is an AI-native course-evaluation platform, and the Bocconi finding shapes a core design stance: measure the things students can validly report, surface the rigour–satisfaction tension instead of hiding it, and never let a single satisfaction mean masquerade as a quality verdict.
- Disentangling experience from difficulty. Koji's AI-moderated conversational interviews probe why a student rated a course as they did, and its automatic thematic analysis can separate "the course was hard" from "the course was badly taught" — the exact conflation the Bocconi study warns about. A low score driven by productive struggle reads very differently from one driven by genuine disorganization, and Koji is designed to make that difference legible.
- Structured items that capture preparation, not just enjoyment. Beyond a global satisfaction
scale, Koji supportsopen_ended,single_choice,multiple_choice,rankingandyes_noitems so programmes can ask whether a module built skills students expect to need later — moving the instrument toward the deferred-value question evaluations usually miss. - Triangulation with downstream evidence and closing the loop. Because student evaluations are weak on deferred learning, Koji is built to sit alongside progression and follow-on outcome data rather than replace it, and to support action tracking so a rigorous module flagged as "hard" is investigated, not penalized.
- Uncertainty-honest, bias-aware reporting. Rather than league-tabling instructors on raw means — the practice most vulnerable to the Bocconi bias — Koji's reporting favours distributions, response counts and context, designed to mitigate the misclassification of demanding teachers as poor ones.
These are mitigations, not cures: no feedback instrument can fully recover deferred learning from an end-of-course survey. For programmes and teams that also conduct customer, product or user research, Koji's core research platform at koji.so applies the same AI-moderated interview engine to those domains.
Related Resources
- Does the Professor Students Rate Highest Teach Them Most? Value-Added Evidence vs Student Ratings
- Do Student Evaluations Measure Learning? The Uttl Meta-Analysis Revisited
- Students Rate the Classes Where They Learn the Most Lower: The Feeling-of-Learning Gap
- Does Grading Leniency Inflate Student Evaluations? What the Evidence Actually Shows
- Do Student Evaluations Encourage Grade Inflation? The Incentive Problem
- The Student-as-Consumer Effect: What a Consumer Mindset Does to Course Evaluations
References
- Braga, M., Paccagnella, M., & Pellizzari, M. (2014). Evaluating students' evaluations of professors. Economics of Education Review, 41, 71–88. https://doi.org/10.1016/j.econedurev.2014.04.002
- Carrell, S. E., & West, J. E. (2010). Does professor quality matter? Evidence from random assignment of students to professors. Journal of Political Economy, 118(3), 409–432. https://doi.org/10.1086/653808
- Uttl, B., White, C. A., & Gonzalez, D. W. (2017). Meta-analysis of faculty's teaching effectiveness: Student evaluation of teaching ratings and student learning are not related. Studies in Educational Evaluation, 54, 22–42. https://doi.org/10.1016/j.stueduc.2016.08.007
- Cohen, P. A. (1981). Student ratings of instruction and student achievement: A meta-analysis of multisection validity studies. Review of Educational Research, 51(3), 281–309. https://doi.org/10.3102/00346543051003281
Related articles
Does the Professor Students Rate Highest Teach Them Most? Value-Added Evidence vs Student Ratings
The strongest causal evidence — random-assignment studies that track follow-on course performance — shows the instructors students rate highest are often not the ones whose students learn most over time. What Carrell & West (2010) and Braga et al. (2014) found, and what it means for course evaluation.
The Student-as-Consumer Effect: What a Consumer Mindset Does to Course Evaluations
When students see themselves as paying customers, what happens to how they rate courses — and how they learn? A look at Bunce, Baird & Jones (2017) and why the consumer frame quietly distorts the meaning of satisfaction scores.
Students Rate the Classes Where They Learn the Most Lower: The Feeling-of-Learning Gap
A randomised Harvard experiment found students learned more in active classrooms but rated their own learning lower. What the feeling-of-learning gap means for interpreting course-evaluation items that ask how much students learned.
Do Student Evaluations Measure Learning? The Uttl Meta-Analysis Revisited
A research-grounded reading of the Uttl, White & Gonzalez (2017) meta-analysis on the SET–learning relationship, with implications for European course evaluation and quality-assurance policy.