New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Evaluation bias8 min read

Big Classes, Lower Scores: Is Class Size Quietly Biasing Your Course Evaluations?

Larger courses tend to receive lower student evaluations, and the effect is largely a property of the format rather than the teaching. Here is what the evidence actually shows, why it matters for fair comparison, and how to stop penalising instructors for teaching at scale.

Koji Education Team

Product · July 12, 2026

Bottom line up front: Across decades of research, bigger classes tend to earn lower student evaluation scores — and much of that gap reflects the constraints of teaching at scale, not weaker teaching. If your quality-assurance process compares a 15-student seminar leader against a 250-student lecture convenor on the same raw mean, you are measuring class size as much as competence. The fix is not to abandon student feedback; it is to interpret it in context and to collect the kind of feedback that a headcount cannot distort.

The pattern is old, consistent, and widely replicated

The negative association between class size and student ratings is one of the more stable findings in the student-evaluation literature. Kenneth Feldman's careful review, Class size and college students' evaluations of teachers and courses: A closer look (1984, Research in Higher Education), concluded that the relationship is generally negative but modest, and — crucially — uneven across what is being rated. The dip is largest on items tied to interaction, rapport, and the instructor's availability and responsiveness; it is smaller or negligible on items about the instructor's knowledge or organisation. In other words, class size mostly depresses the dimensions that large-format teaching structurally constrains.

A more rigorous causal estimate comes from Kelly Bedard and Peter Kuhn, Where class size really matters: Class size and student ratings of instructor effectiveness (2008, Economics of Education Review, 27, 253–265). Using instructor fixed effects — comparing the same instructor teaching sections of different sizes, which strips out the "good teachers get assigned small classes" confound — they found a large, statistically robust, and non-linear negative effect of class size on ratings. The relationship is steepest at the small end: going from a seminar to a mid-sized class hurts ratings far more than going from a large class to a very large one, where scores have already flattened out.

That non-linearity matters for interpretation. It means the raw evaluation gap between a 12-person and a 40-person class can be substantial and driven almost entirely by format, while the gap between a 120-person and a 200-person class may be trivial. A naïve league table treats all of those differences as signal about the teacher. They are not.

Why size depresses scores (and why that is not the instructor's fault)

Several mechanisms plausibly combine, and they are worth separating because they have different implications:

  • Reduced interaction. Large lectures mechanically limit questions, individual attention, discussion, and timely feedback — precisely the experiences students reward. This is the Feldman finding: the interaction items take the hit.
  • Pedagogical constraint. Large enrolments push instructors toward lecturing and automated or rubric-based assessment, away from the seminar-style discussion many students prefer. The format, not the person, narrows the toolkit.
  • Compositional effects. Big classes are disproportionately compulsory, introductory, and service courses — which independently attract lower ratings (a distinct bias we cover in Required Courses Get Worse Evaluations). Class size and course type are entangled, and failing to separate them double-counts the penalty.
  • Statistical artefacts of aggregation. Larger samples produce tighter, more reliable means; smaller classes swing more. That is a separate issue from the size bias on the score level, but it compounds the unfairness of head-to-head comparison — see What Is a Fair Score for a Class of 12?.

None of these is a claim that a large-class instructor is doing worse teaching. They are structural properties of teaching hundreds of students in a room.

But doesn't a low score in a big class still tell you something real?

This is the strongest counterargument, and it deserves a straight answer: yes, sometimes it does — and that is exactly why crude adjustment is the wrong response.

If large classes systematically frustrate students, the frustration is real and worth acting on. A student who cannot get a question answered in a 300-person lecture has a legitimate grievance. The problem is not that the feedback is false; it is that attributing it to the instructor's competence is a category error. The signal is often about resourcing, room design, tutorial provision, and assessment support — decisions made by heads of department and timetablers, not by the lecturer standing at the front.

So the goal is not to "correct away" the class-size effect until every course looks identical. That would erase genuine information about which formats are failing students. The goal is attribution: to know how much of a score reflects the person and how much reflects the conditions they were handed. Blanket size adjustments, applied mechanically, can over-correct and hide a real problem in a badly run large module. The honest approach is to model the effect, flag it, and then look at the qualitative detail — not to launder it out of the numbers.

What defensible practice looks like

  1. Never compare raw means across radically different class sizes. Benchmark like with like, or use models that account for size. Multilevel and empirical-Bayes approaches — see Stop Ranking Instructors on Raw Means and Your Instructor Averages Are Mostly Noise — let you separate instructor effects from course-level conditions instead of pretending a mean is a clean measure of a person.
  2. Disaggregate size from course type. A compulsory 200-person first-year module carries three overlapping penalties (size, compulsion, introductory level). Report them as such.
  3. Read the dimensions, not just the overall score. If the dip is concentrated on interaction items and knowledge/organisation hold up, that is the class-size signature — and it points at provision, not competence. Averaging everything into one number destroys exactly this diagnostic detail (see Why Averaging Likert Scores Misleads).
  4. Ask better questions than a headcount can distort. A number invites the size penalty; a specific account of what a student experienced does not.

What a fair large-class benchmark looks like in practice

Concretely, a defensible quality process does four things before it lets a class-size-affected score influence a decision. It groups courses into comparable size bands rather than ranking across the whole range, so a 250-seat lecture is judged against other large lectures, not against tutorials. It records the confound explicitly — flagging where a low score coincides with compulsory status and introductory level, so a reader sees three overlapping penalties rather than one damning number. It reads the item-level pattern, because a dip concentrated on interaction and responsiveness with organisation and knowledge intact is the class-size signature, not a competence signal. And it routes the finding to the right owner: if the problem is tutorial ratios or assessment turnaround at scale, that is a resourcing decision for a head of department, not a mark against the lecturer.

This is not exotic. It is simply taking seriously the European Standards and Guidelines expectation that quality judgments rest on evidence used fairly, rather than on a raw mean lifted out of context. The instructor who volunteers to teach the 300-person first-year gateway module should not be quietly penalised in the next promotion round for doing the job the department needed doing.

Where Koji fits

The class-size problem is, at root, a problem of thin data and unfair aggregation. A single Likert mean gives you nowhere to go: it is depressed by format, and you cannot tell why. Koji for Education is built to attack that thinness directly.

Instead of a static form that returns a number, Koji runs AI-moderated conversational interviews that probe beyond the rating — when a student in a large module says the course was "hard to follow," the interviewer asks what specifically got lost, whether they could get help, and what would have changed it. That turns a size-depressed score into an actionable account of whether the issue was the lecturer, the tutorial ratio, or the assessment design. Koji's automatic thematic analysis then aggregates those accounts across hundreds of respondents, so a 250-person lecture yields structured themes rather than one blunt average — you see what scaled badly, not merely that the number is low. Because the moderation is standardised and bias-aware, every student in every class is probed consistently, which is exactly what fair cross-format comparison requires. And programme- and institution-level reporting lets you view instructor signal against the backdrop of format and provision, rather than dropping a seminar leader and a mega-lecture convenor into the same league table.

Koji does not eliminate the class-size effect — nothing does, because part of it is a real experience worth hearing. What Koji does is surface and contextualise it, so your quality process acts on the right cause.

The same conversational interview engine powers the main Koji platform for general user and customer research, where the identical problem — a satisfaction number that hides its own causes — shows up in product feedback.

If your evaluation reports still rank a lecture hall against a seminar on a single mean, you are grading the timetable, not the teaching. See how Koji for Education handles evaluation at scale.