New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods9 min read

Class Size and Student Evaluations: What Bedard and Kuhn Found

Does class size bias student evaluations of teaching? Bedard and Kuhn (2008) found a large, non-linear negative effect of enrolment on instructor ratings even after controlling for instructor and course. This article synthesises the evidence and explains how to stop class size from contaminating cross-instructor comparisons.

Koji Education Team

Product

In brief: Larger classes tend to receive lower instructor evaluation scores, and the best-identified study on the question — Bedard & Kuhn (2008) — found this negative effect is large, highly significant, and non-linear even after holding the instructor and the course constant. That means comparing a 15-student seminar to a 300-student lecture on raw scores is comparing unlike things. Class size should be treated as a known confounder and reported alongside scores, not silently averaged away.

The question, stated precisely

When a quality-assurance office ranks instructors or programmes on evaluation scores, it implicitly assumes the scores are comparable across very different teaching contexts. One of the most consequential threats to that assumption is class size. If students systematically rate large classes lower for reasons that have nothing to do with the instructor's competence, then enrolment becomes a hidden tax on anyone who teaches the big required lecture — and a hidden subsidy for anyone who teaches small electives.

The empirical question is therefore not "are large classes worse?" but "does class size move evaluation scores independently of the instructor and the course content?" Answering it cleanly requires data that can separate the size effect from the confounds that travel with it — the fact that different instructors teach different-sized classes, and that some subjects are taught only in large formats.

What the research says

The strongest single piece of evidence is Bedard & Kuhn (2008), "Where class size really matters: Class size and student ratings of instructor effectiveness" (Economics of Education Review, 27(3), 253–265). Using data on every economics class offered at the University of California, Santa Barbara from Fall 1997 to Spring 2004, the authors exploited a research-design strength that most evaluation studies lack: they could control for both instructor fixed effects and course fixed effects simultaneously. That means they compared the same instructor teaching the same course at different enrolments, stripping out the obvious confound that better or worse teachers might be assigned systematically to large or small sections.

Their headline finding: a large, highly significant, and non-linear negative impact of class size on student ratings of instructor effectiveness, robust to those fixed effects. The non-linearity matters — the rating penalty is steepest as classes grow from small to moderate size, then flattens, so the marginal effect of one more student is not constant. Crucially, because instructor identity was held fixed, the effect cannot be explained away as "worse teachers happen to teach bigger classes."

This finding sits in tension with the more sanguine validity tradition. Marsh & Roche (2000) — and the broader body of Marsh's multidimensional SEEQ research — characterised evaluations as "relatively unaffected" by background variables including class size, treating size effects as small enough to ignore under appropriate conditions. The contrast is instructive: a careful fixed-effects econometric design (Bedard & Kuhn) surfaces a sizeable size effect that instrument-validation studies (Marsh & Roche) had tended to downplay. Part of the discrepancy is methodological — without within-instructor variation, a raw correlation between size and ratings can be muted by assignment patterns, masking a real causal effect.

More recent work continues to find size-related effects. Fisher, Vu & Lai (2024), "Faculty course evaluations and class size" (Active Learning in Higher Education, advance online), report that larger enrolments are associated with less favourable course evaluations, echoing the direction Bedard and Kuhn identified and underscoring that the pattern is not an artefact of one campus or one cohort.

The convergent reading: class size is a genuine, non-trivial driver of evaluation scores, and the cleaner the causal identification, the clearer that signal becomes.

Why it matters for course evaluation in practice

The practical stakes for a QA office are direct:

  • Cross-instructor league tables are confounded by enrolment. An instructor scoring 4.0 in a 250-seat lecture may be teaching more effectively than a colleague scoring 4.3 in a 12-person seminar. Ranking them on raw scores embeds a structural penalty for teaching at scale.
  • It distorts workload and staffing incentives. If large-class teaching reliably lowers scores and scores feed promotion, rational faculty will avoid the very high-enrolment courses that institutions most need staffed well.
  • Programme-level trends can be misread. A department that grows its cohort may see average evaluation scores fall for reasons that have nothing to do with declining teaching quality — a classic case of mistaking a composition shift for a quality shift.

The defensible response is to treat class size as a reported covariate, compare like with like (size band against size band), and resist collapsing heterogeneous teaching contexts into a single rank order.

Limitations and honest caveats

A careful reader should temper the conclusion in several ways:

  • Single-discipline, single-institution identification. Bedard and Kuhn's clean design comes from economics classes at one US university. The existence of a size effect is well supported, but its precise magnitude and curvature may differ across disciplines, countries, and assessment cultures.
  • "Class size" bundles several mechanisms. A large class differs from a small one in interaction, anonymity, assessment type, and room dynamics. The studies estimate the net effect of size; they do not isolate which mechanism does the work, so "control for size" is a pragmatic correction, not a causal explanation.
  • Effect size is not destiny. Even a statistically large size effect may be small relative to the genuine teaching-quality variance an institution cares about. Over-correcting for size could obscure real differences in effectiveness.
  • Fixed-effects designs assume stable teaching. Holding instructor identity constant assumes the same person teaches similarly across enrolments; if instructors adapt their methods substantially at scale, part of the "size effect" is really a teaching-response effect.

Acknowledging these limits is what separates a defensible adjustment from a mechanical one.

How Koji incorporates this

Koji for Education is designed to keep class size from silently distorting evaluation evidence:

  • Size-aware, like-for-like reporting. Rather than collapsing every course into one league table, Koji's reporting keeps enrolment context attached to results so a 300-student lecture is benchmarked against comparable large classes, not against small seminars — directly addressing the confound Bedard and Kuhn quantified.
  • Conversational depth that scales without diluting. The penalty large classes pay is partly about anonymity and thin feedback. Koji's AI-moderated interview gives every student in a 300-person cohort the same adaptive, probing conversation a seminar student would get, so qualitative richness no longer collapses as enrolment rises.
  • Automatic thematic analysis across high-volume cohorts. With open_ended responses analysed thematically at scale, evaluators can see whether low scores in a big class reflect genuine teaching issues or structural factors (room, pacing, assessment load) that size introduces — separating signal from the size artefact.
  • Triangulation across cohorts and time. By comparing the same instructor across differently sized offerings, Koji supports the within-instructor comparison that makes Bedard and Kuhn's design so credible, helping institutions distinguish a teaching problem from an enrolment effect.

Koji is built to mitigate the class-size confound, not to claim it disappears: enrolment will always shape the student experience, but size should never be allowed to masquerade as teaching quality in a personnel file. The same AI-moderated interview engine underpins Koji's core research platform at koji.so, where comparing feedback fairly across very different sample sizes is the same methodological challenge.

Related resources

Practical guidance for evaluation committees

For committees that benchmark instructors or programmes, the class-size evidence supports a handful of concrete practices. Begin by attaching enrolment to every reported score and comparing within size bands — large lectures against other large lectures, seminars against seminars — so that no one is penalised simply for teaching at scale. Where possible, exploit within-instructor comparisons in the spirit of Bedard and Kuhn: when the same person teaches the same course at different enrolments, the difference in scores is a far cleaner signal of a size effect than any cross-instructor ranking. Treat a falling programme-level average with suspicion if the cohort has grown, because a composition shift can masquerade as a quality decline. Avoid mechanical size corrections that assume a fixed penalty per student; the effect is non-linear, steepest in the move from small to moderate classes, so a flat adjustment will misprice most courses. Finally, protect high-enrolment teaching from a structural scoring disadvantage in promotion criteria, or the institution will quietly disincentivise exactly the large-class teaching it most needs staffed well. None of this denies that large classes genuinely change the student experience — the aim is to stop size from being mistaken for teaching quality.

References