What Is a Fair Score for a Class of 12? Small Classes and the Statistics of Uncertainty in Course Evaluation
A 4.2 from twelve students and a 4.2 from two hundred are not the same evidence — but most evaluation reports show them identically. Why small-class means are statistically fragile, and how confidence intervals, reliability thresholds, and shrinkage make them honest.
Koji Education Team
Product ·
The short answer: A mean of 4.2 from a class of twelve and a mean of 4.2 from a class of two hundred look identical on a dashboard, but they are not the same evidence. Small-class evaluation means are statistically fragile: one or two atypical responses can swing them, their confidence intervals are wide, and ranking instructors on them is close to reading tea leaves. The remedy is not to hide small classes or pretend the number is solid. It is to report uncertainty honestly — with confidence intervals, reliability thresholds, and shrinkage estimates — and to lean on qualitative depth where the quantitative signal is thin.
The problem hiding in plain sight
Most institutional evaluation reports present a course mean to one or two decimal places, often with a tidy comparison against a departmental or institutional benchmark. What they rarely show is how much that mean would move if a couple of students had answered differently. For a large lecture, the answer is "barely." For a seminar of twelve, the answer is "a lot."
This is not a niche concern. Seminars, labs, options, postgraduate modules, and specialist courses routinely enrol fewer than twenty students, and at many institutions a large share of all courses are small. Yet the same reporting template — a mean, a benchmark, sometimes a traffic-light colour — is applied regardless of N. The result is that the most statistically uncertain scores are presented with exactly the same false confidence as the most solid ones.
The statistics of uncertainty
Two facts from basic sampling theory drive the whole issue.
First, the precision of a mean depends on sample size. The standard error of a mean shrinks with the square root of N, so estimates from small classes carry much wider uncertainty. A useful way to see this is the response rate literature. Nulty (2008), in a widely cited paper, worked out the response rates needed for evaluation data to be representative at given margins of error. For a relatively forgiving 10% margin, a class of 20 needs a 58% response rate, a class of 50 needs 35%, and a class of 100 needs 21%. Tighten the margin to 3% and the same classes need 97%, 93%, and 87% respectively (Nulty, 2008, Assessment & Evaluation in Higher Education). The smaller the class, the closer to a census you need — and the smaller class is precisely where a near-census is hardest to obtain.
Second, reliability rises with the number of raters. Decades of psychometric work on student ratings show that the reliability of a class-mean rating climbs with class size, and that roughly 25 responding students are needed before a global teaching rating reaches the conventional 0.90 reliability threshold (a benchmark consistent with Marsh and Roche's long-standing guidance). Below that, the proportion of a score that reflects genuine teaching quality — rather than the idiosyncratic perceptions of a few students — drops. Feistauer and Richter (2017), analysing 4,224 evaluations with cross-classified multilevel models, found that the interaction between individual students and teachers was the single largest source of variance — concluding that aggregated evaluation scores should be used with caution (Feistauer & Richter, 2017, Assessment & Evaluation in Higher Education). In a class of twelve, that student-specific noise is not averaged away; it dominates.
Put bluntly: for a small class, the difference between a 3.9 and a 4.3 is very often noise.
Three ways to make small-class scores honest
You do not have to abandon quantitative evaluation for small classes. You have to report it truthfully.
1. Show confidence intervals, not just point estimates. A mean of 4.2 with a 95% interval of [3.6, 4.8] is a different statement than a bare 4.2. Intervals make the uncertainty visible and stop readers from over-interpreting decimal-place differences — the over-interpretation problem we examine in whether evaluations should decide tenure and promotion.
2. Set and enforce reliability thresholds. Many institutions adopt a minimum-N rule — for example, not reporting or formally using a mean below a set number of responses, and never using small-class means in isolation for personnel decisions. This is consistent with the broader case that evaluations should be interpreted with explicit reliability evidence.
3. Use shrinkage (empirical-Bayes) estimates. This is the statistically sophisticated move and the least familiar to most administrators. A shrinkage estimator pulls each small-class mean partway toward the overall grand mean, by an amount that depends on how little data supports it. A class of 200 barely moves; a class of 8 moves a lot. The effect is to stop tiny, noisy samples from producing the extreme highs and lows that fill the top and bottom of naive league tables. Esarey and Valdes (2020) make a closely related argument: even under generous assumptions, ranking instructors on raw means produces serious misclassification, and statistical correction is needed before such scores carry weight.
"But we just need one number to compare courses"
This is the strongest and most honest counterargument, and it deserves a straight answer rather than a dismissal.
Administrators are not being unreasonable when they ask for comparability. Quality processes, annual reviews, and accreditation evidence all benefit from consistent metrics, and confidence intervals can feel like an invitation to do nothing. There are three legitimate worries:
- Shrinkage can mask real differences. Pulling a genuinely excellent small seminar toward the mean is a real cost, not just a benefit. A brilliant tutor of ten students will look more average than they are.
- Intervals can be weaponised. "It is not statistically significant" can become an excuse to ignore a clear, consistent pattern of student concern.
- Operational simplicity matters. A 4.2 is easy to put in a spreadsheet; a posterior distribution is not.
The resolution is not to pretend small samples are precise. It is to change what the number is for. For small classes, the quantitative mean should be treated as a weak prior — useful for spotting outliers worth investigating, never as a precise ranking. The decisive evidence for a class of twelve should be qualitative: what twelve students actually said, read carefully, will almost always be more informative than where their average landed to one decimal place. And a consistent qualitative theme ("three of twelve independently said the assessment brief was unclear") is real signal even when the mean is statistically mushy.
How Koji approaches small classes
This is where conversational evaluation has a structural advantage. Koji for Education does not rest the weight of a small class on a fragile mean. Its AI-moderated conversational interviews produce rich, attributable qualitative feedback from each respondent, and its automatic thematic analysis surfaces recurring themes even when N is small — turning "twelve scattered comments" into "the two issues most students raised." Because the analysis identifies what students consistently say rather than only how high they rated, it extracts usable signal from samples too small for a stable Likert mean.
Where quantitative items are used, Koji's reporting is designed to present them in context — alongside the qualitative themes and quality scoring that tell you whether the responses are substantive — rather than as decimal-place rankings. And because the conversational format lifts engagement, small classes are more likely to reach the near-census that Nulty's numbers show they need. The same interview engine underpins general research on the main Koji platform, where the small-sample problem is just as acute in early-stage user research.
To be precise: Koji does not make a class of twelve statistically equivalent to a class of two hundred — nothing can. It reduces the institution's dependence on a fragile number, and surfaces the qualitative evidence that is genuinely informative at small N.
The bottom line
A score is only as trustworthy as the sample behind it. Reporting a class of twelve and a class of two hundred with the same confident decimal place is not neutral — it actively misleads the people making decisions. Show the uncertainty, set reliability thresholds, consider shrinkage, and let careful qualitative reading carry the weight where the numbers cannot.
See how Koji surfaces themes from small cohorts at Koji for Education, and read our companions on reliability vs validity and why averaging Likert scores misleads.