Class Size and Student Evaluations: What Bedard and Kuhn Found
Does class size bias student evaluations of teaching? Bedard and Kuhn (2008) found a large, non-linear negative effect of enrolment on instructor ratings even after controlling for instructor and course. This article synthesises the evidence and explains how to stop class size from contaminating cross-instructor comparisons.
Koji Education Team
Product
In brief: Larger classes tend to receive lower instructor evaluation scores, and the best-identified study on the question — Bedard & Kuhn (2008) — found this negative effect is large, highly significant, and non-linear even after holding the instructor and the course constant. That means comparing a 15-student seminar to a 300-student lecture on raw scores is comparing unlike things. Class size should be treated as a known confounder and reported alongside scores, not silently averaged away.
The question, stated precisely
When a quality-assurance office ranks instructors or programmes on evaluation scores, it implicitly assumes the scores are comparable across very different teaching contexts. One of the most consequential threats to that assumption is class size. If students systematically rate large classes lower for reasons that have nothing to do with the instructor's competence, then enrolment becomes a hidden tax on anyone who teaches the big required lecture — and a hidden subsidy for anyone who teaches small electives.
The empirical question is therefore not "are large classes worse?" but "does class size move evaluation scores independently of the instructor and the course content?" Answering it cleanly requires data that can separate the size effect from the confounds that travel with it — the fact that different instructors teach different-sized classes, and that some subjects are taught only in large formats.
What the research says
The strongest single piece of evidence is Bedard & Kuhn (2008), "Where class size really matters: Class size and student ratings of instructor effectiveness" (Economics of Education Review, 27(3), 253–265). Using data on every economics class offered at the University of California, Santa Barbara from Fall 1997 to Spring 2004, the authors exploited a research-design strength that most evaluation studies lack: they could control for both instructor fixed effects and course fixed effects simultaneously. That means they compared the same instructor teaching the same course at different enrolments, stripping out the obvious confound that better or worse teachers might be assigned systematically to large or small sections.
Their headline finding: a large, highly significant, and non-linear negative impact of class size on student ratings of instructor effectiveness, robust to those fixed effects. The non-linearity matters — the rating penalty is steepest as classes grow from small to moderate size, then flattens, so the marginal effect of one more student is not constant. Crucially, because instructor identity was held fixed, the effect cannot be explained away as "worse teachers happen to teach bigger classes."
This finding sits in tension with the more sanguine validity tradition. Marsh & Roche (2000) — and the broader body of Marsh's multidimensional SEEQ research — characterised evaluations as "relatively unaffected" by background variables including class size, treating size effects as small enough to ignore under appropriate conditions. The contrast is instructive: a careful fixed-effects econometric design (Bedard & Kuhn) surfaces a sizeable size effect that instrument-validation studies (Marsh & Roche) had tended to downplay. Part of the discrepancy is methodological — without within-instructor variation, a raw correlation between size and ratings can be muted by assignment patterns, masking a real causal effect.
More recent work continues to find size-related effects. Fisher, Vu & Lai (2024), "Faculty course evaluations and class size" (Active Learning in Higher Education, advance online), report that larger enrolments are associated with less favourable course evaluations, echoing the direction Bedard and Kuhn identified and underscoring that the pattern is not an artefact of one campus or one cohort.
The convergent reading: class size is a genuine, non-trivial driver of evaluation scores, and the cleaner the causal identification, the clearer that signal becomes.
Why it matters for course evaluation in practice
The practical stakes for a QA office are direct:
- Cross-instructor league tables are confounded by enrolment. An instructor scoring 4.0 in a 250-seat lecture may be teaching more effectively than a colleague scoring 4.3 in a 12-person seminar. Ranking them on raw scores embeds a structural penalty for teaching at scale.
- It distorts workload and staffing incentives. If large-class teaching reliably lowers scores and scores feed promotion, rational faculty will avoid the very high-enrolment courses that institutions most need staffed well.
- Programme-level trends can be misread. A department that grows its cohort may see average evaluation scores fall for reasons that have nothing to do with declining teaching quality — a classic case of mistaking a composition shift for a quality shift.
The defensible response is to treat class size as a reported covariate, compare like with like (size band against size band), and resist collapsing heterogeneous teaching contexts into a single rank order.
Limitations and honest caveats
A careful reader should temper the conclusion in several ways:
- Single-discipline, single-institution identification. Bedard and Kuhn's clean design comes from economics classes at one US university. The existence of a size effect is well supported, but its precise magnitude and curvature may differ across disciplines, countries, and assessment cultures.
- "Class size" bundles several mechanisms. A large class differs from a small one in interaction, anonymity, assessment type, and room dynamics. The studies estimate the net effect of size; they do not isolate which mechanism does the work, so "control for size" is a pragmatic correction, not a causal explanation.
- Effect size is not destiny. Even a statistically large size effect may be small relative to the genuine teaching-quality variance an institution cares about. Over-correcting for size could obscure real differences in effectiveness.
- Fixed-effects designs assume stable teaching. Holding instructor identity constant assumes the same person teaches similarly across enrolments; if instructors adapt their methods substantially at scale, part of the "size effect" is really a teaching-response effect.
Acknowledging these limits is what separates a defensible adjustment from a mechanical one.
How Koji incorporates this
Koji for Education is designed to keep class size from silently distorting evaluation evidence:
- Size-aware, like-for-like reporting. Rather than collapsing every course into one league table, Koji's reporting keeps enrolment context attached to results so a 300-student lecture is benchmarked against comparable large classes, not against small seminars — directly addressing the confound Bedard and Kuhn quantified.
- Conversational depth that scales without diluting. The penalty large classes pay is partly about anonymity and thin feedback. Koji's AI-moderated interview gives every student in a 300-person cohort the same adaptive, probing conversation a seminar student would get, so qualitative richness no longer collapses as enrolment rises.
- Automatic thematic analysis across high-volume cohorts. With
open_endedresponses analysed thematically at scale, evaluators can see whether low scores in a big class reflect genuine teaching issues or structural factors (room, pacing, assessment load) that size introduces — separating signal from the size artefact. - Triangulation across cohorts and time. By comparing the same instructor across differently sized offerings, Koji supports the within-instructor comparison that makes Bedard and Kuhn's design so credible, helping institutions distinguish a teaching problem from an enrolment effect.
Koji is built to mitigate the class-size confound, not to claim it disappears: enrolment will always shape the student experience, but size should never be allowed to masquerade as teaching quality in a personnel file. The same AI-moderated interview engine underpins Koji's core research platform at koji.so, where comparing feedback fairly across very different sample sizes is the same methodological challenge.
Related resources
- Mid-Semester Feedback and the Power of Consultation
- Do Student Evaluations Measure Learning? The Uttl Meta-Analysis Revisited
- Does Grading Leniency Inflate Student Evaluations? What the Evidence Actually Shows
- Gender Bias in Student Evaluations of Teaching: Evidence and Mitigation
- Selection Bias in Course Evaluations: What Goos and Salomons Found
Practical guidance for evaluation committees
For committees that benchmark instructors or programmes, the class-size evidence supports a handful of concrete practices. Begin by attaching enrolment to every reported score and comparing within size bands — large lectures against other large lectures, seminars against seminars — so that no one is penalised simply for teaching at scale. Where possible, exploit within-instructor comparisons in the spirit of Bedard and Kuhn: when the same person teaches the same course at different enrolments, the difference in scores is a far cleaner signal of a size effect than any cross-instructor ranking. Treat a falling programme-level average with suspicion if the cohort has grown, because a composition shift can masquerade as a quality decline. Avoid mechanical size corrections that assume a fixed penalty per student; the effect is non-linear, steepest in the move from small to moderate classes, so a flat adjustment will misprice most courses. Finally, protect high-enrolment teaching from a structural scoring disadvantage in promotion criteria, or the institution will quietly disincentivise exactly the large-class teaching it most needs staffed well. None of this denies that large classes genuinely change the student experience — the aim is to stop size from being mistaken for teaching quality.
References
- Bedard, K., & Kuhn, P. (2008). Where class size really matters: Class size and student ratings of instructor effectiveness. Economics of Education Review, 27(3), 253–265. https://doi.org/10.1016/j.econedurev.2006.08.007
- Marsh, H. W., & Roche, L. A. (2000). Effects of grading leniency and low workload on students' evaluations of teaching: Popular myth, bias, validity, or innocent bystanders? Journal of Educational Psychology, 92(1), 202–228. https://doi.org/10.1037/0022-0663.92.1.202
- Fisher, C., Vu, P., & Lai, P. (2024). Faculty course evaluations and class size. Active Learning in Higher Education. https://doi.org/10.1177/14697874221126739
Related articles
Gender Bias in Student Evaluations of Teaching: Evidence and Mitigation
A research-grounded synthesis of gender bias in student evaluations of teaching (SET), anchored to MacNell, Driscoll & Hunt (2015) and corroborated by Boring (2017) and Mengel, Sauermann & Zölitz (2019), with practical mitigations for European quality-assurance teams.
Do Student Evaluations Measure Learning? The Uttl Meta-Analysis Revisited
A research-grounded reading of the Uttl, White & Gonzalez (2017) meta-analysis on the SET–learning relationship, with implications for European course evaluation and quality-assurance policy.
Mid-Semester Feedback and the Power of Consultation: What the Meta-Analyses Show
Penny & Coe (2004) and Cohen (1980) on the formative effectiveness of mid-semester student feedback when combined with consultation, and the implications for European course-evaluation design.
Does Grading Leniency Inflate Student Evaluations? What the Evidence Actually Shows
A research-grounded look at the grading-leniency hypothesis in student evaluations of teaching: what Marsh and Roche, Greenwald and Gillmore, and Stroebe found, why the grades-ratings correlation is genuinely ambiguous, and how to design evaluation so leniency cannot quietly buy better scores.