Stop Comparing Raw Averages: Empirical-Bayes Shrinkage for Course-Evaluation Scores
Why raw course-evaluation means mislead — especially for small classes — and how empirical-Bayes shrinkage (partial pooling), the method Kane and Staiger applied to teacher value-added, produces fairer, more accurate instructor estimates.
Koji Education Team
Product
In brief
A course that returns ten responses with a 4.8 average is not reliably better than one that returns eighty responses with a 4.5 average — the 4.8 is mostly noise, and the right statistical correction is shrinkage. Empirical-Bayes shrinkage (also called partial pooling or the best-linear-unbiased-predictor approach) pulls each course's raw mean toward the overall average by an amount that depends on how noisy that mean is: small, high-variance samples get pulled hard, large stable samples barely move. Kane and Staiger (2008) made this the standard method for estimating teacher value-added, and the same logic applies directly to course-evaluation means. Reporting raw averages, then ranking instructors by them, systematically rewards small classes and lucky draws; shrinkage is the established fix, with important caveats about when it helps.
What the research says
The statistical foundation is older than course evaluation itself. Efron and Morris (1977), in their classic Scientific American exposition of "Stein's paradox," showed something counterintuitive: when you are estimating many quantities at once (many instructors' true ratings), the individual sample averages are not the best estimates. You can do provably better — lower total squared error — by shrinking each average toward the group mean. The noisier an individual estimate, the more it should be shrunk. This is not a heuristic; it is a theorem about estimation under uncertainty.
Kane and Staiger (2008), in an experimental evaluation of teacher value-added (NBER Working Paper 14607), turned this into the dominant applied method in education measurement. Facing the problem that a teacher's measured effect in a single year is heavily contaminated by sampling noise — small classes, idiosyncratic cohorts — they applied an empirical-Bayes procedure that shrinks each teacher's estimate toward the mean in proportion to its unreliability. The shrunken estimates are the best linear unbiased predictors: they minimise expected mean-squared error between the estimate and the teacher's true contribution. Critically, Kane and Staiger validated the approach against an experiment with random assignment and found the shrunken predictions were approximately unbiased forecasts of future performance — exactly the property you want if you are going to act on a score.
The method is now standard but not uncontested. Guarino and colleagues (2015), in the Journal of Educational and Behavioral Statistics, evaluated empirical-Bayes estimators rigorously and found they perform well under random assignment of students to classes, but their advantage erodes — and shrinkage alone does not rescue them — when assignment is non-random and systematically related to the thing being measured. That caveat transfers cleanly to course evaluation: shrinkage fixes noise, not bias.
Why it matters for course evaluation in practice
The single most common error in course-evaluation reporting is treating a raw mean as a precise score and ranking on it. Shrinkage exposes why this is wrong and what to do instead.
Small classes dominate the extremes. Because variance is larger when n is small, the highest and lowest raw averages in any institution are disproportionately small classes — not because small classes are taught better or worse, but because small samples produce extreme means by chance. Rank by raw mean and you build a leaderboard of sample sizes. This is the mechanism behind regression to the mean in course evaluations: an extreme score one year is partly luck and tends to move toward the average the next, even if nothing changed.
Shrinkage produces fairer comparisons. A shrunken estimate answers the decision-relevant question — "what is our best prediction of this instructor's true rating?" — rather than the misleading one, "what number did this particular sample produce?" An instructor with a 4.9 from eight students and one with a 4.6 from a hundred may have nearly identical shrunken estimates, because the 4.9 carries far less information. Presenting the shrunken estimate, with an interval, prevents committees from over-reading a fragile number.
It complements, not replaces, reliability thinking. Shrinkage is the estimation counterpart to the reliability arguments in can you fairly rank instructors by their scores and the misclassification analysis in why course-evaluation scores misclassify instructors. Where those articles show that low-reliability scores cannot support fine rankings, shrinkage gives the constructive method for reporting the best defensible estimate you can. It pairs naturally with the variance-partitioning logic of generalizability theory and the rater-and-item modelling of many-facet Rasch measurement.
A worked example
Suppose a faculty's courses average 4.3 with a between-course standard deviation of 0.3, and the typical within-course sampling noise for a small seminar is large. A seminar of 8 students returns a raw mean of 4.9. Empirical Bayes asks: how much of that 0.6 gap above the faculty mean is real signal versus sampling noise? Because n is tiny, the estimated reliability of the 4.9 is low — say 0.3 — so the shrunken estimate is roughly the faculty mean plus 0.3 times the gap: about 4.3 + 0.3 x 0.6 = 4.48. The same calculation for a 120-student lecture scoring 4.6 with reliability near 0.9 barely moves it: about 4.57. The raw ranking put the seminar far ahead; the shrunken estimates put them essentially level. The shrinkage has not punished the seminar — it has declined to over-interpret eight data points, which is exactly the discipline a personnel committee needs.
Limitations and honest caveats
Shrinkage is powerful but not a cure-all, and a careful reader should hold these objections.
Shrinkage corrects noise, not bias. This is the central caveat from Guarino et al. (2015). If course-evaluation scores are systematically biased — by instructor gender, class difficulty, grading leniency, or discipline — shrinking them pulls biased estimates toward a biased mean. You get a more precise version of a biased measure, which can be more dangerous than an obviously noisy one because it looks authoritative. Shrinkage must sit downstream of bias adjustment, not in place of it.
It assumes a sensible prior. Empirical Bayes estimates the "shrink toward what, and how hard" from the data, which requires a defensible population and distributional assumptions. Pooling dissimilar courses — a 400-student lecture and an 8-student seminar, or courses across faculties with different rating cultures — toward a single grand mean can distort rather than improve. Shrinkage should usually be done within comparable strata, which connects to measurement-invariance concerns about whether groups are comparable at all.
Shrunken estimates are conservative by design. Because they pull toward the mean, shrinkage compresses the spread and will understate genuine outliers, especially the genuinely excellent small-class teacher. The estimate is optimised for aggregate predictive accuracy, not for being fair to any single individual — an ethical tension when the output feeds a personnel decision about one person.
Communication is hard. A committee that wanted a clean number now receives a shrunken estimate with an interval and a caveat. Without statistical literacy support, the temptation is to revert to the raw mean. Reporting design matters as much as the method.
How Koji incorporates this
Koji is built so that the reporting layer reflects the uncertainty in the data rather than hiding it behind a single decimal.
Koji's bias-aware reporting is designed to present course-evaluation results with their precision attached — uncertainty intervals and response-count context rather than a bare mean — so that a 4.8 from ten responses is never displayed as though it outranked a 4.5 from eighty. This operationalises the core warning of the shrinkage literature at the point of decision: the platform is designed to discourage the raw-mean leaderboard that small samples corrupt. It complements, rather than substitutes for, an institution's own statistical modelling.
Because shrinkage corrects noise but not bias, Koji attacks the bias problem at the measurement stage. Its AI-moderated conversational interviews and structured question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) gather richer, more behaviourally specific evidence than a lone global rating, and its automatic thematic analysis surfaces why a course scored as it did — so a low mean driven by a workload-design problem is not confused with one driven by teaching quality. Koji frames every estimate as designed to mitigate over-interpretation, never as an objective ranking, and supports triangulation across cohorts so that an instructor is judged on accumulated evidence across courses and terms rather than one small, noisy sample — the most practical real-world remedy for the small-n problem shrinkage addresses.
For research teams applying the same precision-aware reporting beyond education, Koji's core platform at koji.so uses the same conversational engine for customer and product research, where small-sample over-interpretation is an equally common trap.
Related resources
- Can You Fairly Rank Instructors by Their Course-Evaluation Scores? What Esarey & Valdes (2020) Proved
- Why a Low-Scoring Course Usually 'Improves' Next Year Even If Nothing Changed: Regression to the Mean
- The 4.2 vs 4.4 Trap: Why Small Differences in Course-Evaluation Means Are Usually Noise
- Even Fair Evaluations Can Rank Unfairly: Why Course-Evaluation Scores Misclassify Instructors
- Generalizability Theory and the Reliability of Student Ratings: Why One Class Is Not Enough
- Some Raters Are Harsh, Some Items Are Hard: Why Rasch and Many-Facet Models Beat Averaging
References
- Kane, T. J., & Staiger, D. O. (2008). Estimating Teacher Impacts on Student Achievement: An Experimental Evaluation. NBER Working Paper No. 14607. National Bureau of Economic Research. https://doi.org/10.3386/w14607
- Guarino, C. M., Maxfield, M., Reckase, M. D., Thompson, P. N., & Wooldridge, J. M. (2015). An Evaluation of Empirical Bayes' Estimation of Value-Added Teacher Performance Measures. Journal of Educational and Behavioral Statistics, 40(2), 190–222. https://doi.org/10.3102/1076998615574771
- Efron, B., & Morris, C. (1977). Stein's Paradox in Statistics. Scientific American, 236(5), 119–127. https://doi.org/10.1038/scientificamerican0577-119
Related articles
Can You Fairly Rank Instructors by Their Course-Evaluation Scores? What Esarey & Valdes (2020) Proved
Even if course evaluations were unbiased, reliable, and valid, ranking instructors by their scores would still misclassify many good teachers. A walk through the Esarey & Valdes (2020) simulation and what it means for how you report and use evaluation data.
Generalizability Theory and the Reliability of Student Ratings: Why One Class Is Not Enough
Reliability is not one number. Generalizability theory (Gillmore, Kane & Naccarato 1978; Marsh 1984) decomposes the variance in student ratings into student, teacher, course and occasion components — and shows that a single class can be reliable for the course yet a poor estimate of the teacher. What that means for fair evaluation.
The 4.2 vs 4.4 Trap: Why Small Differences in Course-Evaluation Means Are Usually Noise
Faculty and administrators routinely read meaning into tiny gaps between course-evaluation means. Boysen (2015) and Boysen et al. (2014) show this happens even when confidence intervals say the difference is nothing — and that warnings barely help. How to report uncertainty honestly.
Even Fair Evaluations Can Rank Unfairly: Why Course-Evaluation Scores Misclassify Instructors
A simulation-based finding that even unbiased, reliable, and valid course-evaluation scores produce high misclassification rates when used to rank or compare instructors — and what it means for tenure, promotion, and merit decisions.