New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Methodology9 min read

Stop Ranking Instructors on Raw Means: Empirical-Bayes Shrinkage for Fair Comparison

A 4.8 from an 8-student seminar and a 4.5 from a 200-student lecture are not comparable, and ranking them as if they were is a statistical error with a known fix. Empirical-Bayes shrinkage has been the standard answer in every other field for fifty years.

Koji Education Team

Product ·

Bottom line up front: When an institution ranks instructors or departments on their raw mean course evaluation score, it systematically rewards small classes for being noisy and punishes large ones for being precise. A 4.8 from an eight-student seminar carries far less information than a 4.5 from a 200-student lecture, yet a raw-mean ranking places the seminar above the lecture. The fix is not new and not controversial: empirical-Bayes shrinkage (equivalently, partial pooling via a multilevel model) pulls unreliable small-sample means toward the overall average in proportion to how little we trust them. It has been the standard estimator for exactly this problem — comparing many units measured with unequal precision — since the 1970s. Course evaluation is one of the last high-stakes domains still using raw means.

The problem: unequal precision, equal treatment

Every course produces a mean from a different number of responses. A seminar of eight, a workshop of fifteen, a lecture of three hundred. The mean from eight students is a wild thing — one unusually generous or unusually bitter respondent moves it by half a point. The mean from three hundred is stable; no single respondent matters much.

A raw-mean ranking ignores this entirely. It lines up 4.8, 4.7, 4.6, 4.5 as if each were measured to the same precision, and then someone draws a line and calls everything below it "needs improvement". The instructors near the top of such a list are disproportionately those teaching small classes — not because they teach better, but because small samples produce more extreme means in both directions. The same mechanism that lets a small class hit 4.9 also lets a small class crater to 3.1 on one bad cohort. Extremity is a property of the sample size, not the teaching. We made the year-over-year version of this argument in regression to the mean in course evaluation scores; shrinkage is the same insight applied across units at a single point in time.

Why small samples are untrustworthy: the reliability evidence

This is not an abstract worry. Using generalizability theory, Li and colleagues (2018) found that roughly 20 student raters are needed before a course evaluation mean reaches an acceptable dependability of 0.80, and that below that threshold the number of respondents has a large influence on how stable the score is. In other words, a substantial share of real courses — seminars, options, lab groups, dissertations — are being ranked on means that the measurement theory says are not yet reliable. We discuss the small-class case directly in the statistics of uncertainty in small-class evaluations.

If a score is not reliable, treating its exact value as meaningful — and ranking on it — is a category error. The question is what to do instead.

The fix: borrow strength, shrink the unreliable

The answer comes from one of the most celebrated results in twentieth-century statistics. In 1977, Bradley Efron and Carl Morris published Stein's Paradox in Statistics in Scientific American, popularising a finding that had unsettled statisticians since Charles Stein's 1956 proof: when you are estimating many quantities at once, the obvious estimator — use each unit's own observed average — is inadmissible. You can always do better, in total squared error, by shrinking each individual average toward the global average.

Their famous worked example used the 1970 batting averages of 18 Major League players after their first 45 at-bats. Taking each player's early average at face value is a worse predictor of their season-end ability than an estimator that pulls every early average toward the league mean. The player who hit .400 in 45 at-bats almost certainly is not a .400 hitter; the right estimate shrinks him toward the pack. The amount of shrinkage is principled: the less data behind an average, and the further it sits from the group, the harder it is pulled back.

This is empirical-Bayes shrinkage, and its modern, routine implementation is the multilevel (hierarchical) model. Borrowing strength across units to stabilise sparse estimates is the same variance-reduction logic that now governs how sports analytics rate players, how epidemiologists map disease rates across small areas, and how value-added models estimate teacher effects in school systems. Applied to course evaluation, partial pooling does three honest things at once:

  1. It pulls small-sample means toward the institutional average, in proportion to their unreliability. An 8-student 4.8 gets pulled toward the mean substantially; a 300-student 4.5 barely moves.
  2. It produces an interval, not just a point. Each adjusted estimate comes with an honest measure of uncertainty, so a committee can see that two instructors who differ by 0.2 have overlapping intervals and are, statistically, indistinguishable.
  3. It separates real variation between instructors from sampling noise. This is exactly the variance-partitioning logic we describe in multilevel models and variance partitioning for course evaluation.

But doesn't shrinkage just punish good teachers in small classes?

This is the most common and most reasonable objection, and it deserves a direct answer. If an instructor really is excellent, why should we drag their 4.8 down toward the mean just because the class was small?

Three responses. First, shrinkage does not claim the instructor is mediocre — it states honestly that eight ratings are not enough evidence to support a 4.8 claim, and reports the most probable true value given that limited evidence. The instructor is not being punished; the score is being made truthful about its own uncertainty. Second, shrinkage is symmetric: it protects small-class instructors from being unfairly tanked by one harsh cohort exactly as much as it tempers an inflated high. The teacher who got a 3.1 from a single difficult group of nine is the same beneficiary as the one capped from 4.8. A raw-mean system that "rewards" the lucky high is the same system that destroys the unlucky low. Third, if a small-class instructor genuinely is excellent, the signal will persist across cohorts and semesters, and a hierarchical model that accumulates evidence over time will converge on the high estimate as the data justify it. Shrinkage is not a permanent ceiling; it is a demand for evidence proportional to the claim.

The deeper point: the alternative to shrinkage is not "fairness to small classes". The alternative is a system that quietly makes class size, not teaching, the strongest predictor of where you land in the ranking. That is the unfair option, and it is the one most institutions currently use.

What this means in practice

  • Never rank on raw means across units of different sizes. It is comparing measurements of different precision as though they were equivalent. If you must produce a comparison, shrink first.
  • Report uncertainty intervals alongside every adjusted score. A point estimate without an interval invites exactly the false-precision comparisons we warn about in whether a 0.3 difference is a real effect.
  • Resist the league table. Even shrunken estimates, sorted into a rank order and used for tenure and promotion decisions, invite over-interpretation of differences that are within noise. Shrinkage makes the comparison honest; it does not make a noisy comparison decisive. The same caution applies to benchmarking scores across departments.
  • Collect richer evidence so you depend less on the fragile mean. The fundamental fragility of a small-class mean is that it compresses eight people's experience into one number. The more your evaluation captures why — the more it gives you something other than a brittle average to reason about — the less a single noisy statistic has to carry.

Where Koji fits

Shrinkage is the right statistical treatment for the data you already have. But the deeper fix is to stop forcing small cohorts into a single brittle number in the first place. Koji for Education changes the unit of evidence. Instead of eight Likert responses averaged into a 4.8 that one student could swing, its AI-moderated conversational interviews generate structured, comparable qualitative depth from every respondent — so a small class produces rich evidence rather than a fragile mean. Where you do need numbers, Koji's programme- and institution-level reporting is built to aggregate appropriately rather than naively rank course-level averages of wildly different sizes, and its automatic thematic analysis surfaces what is consistent across a small cohort — a far more stable signal than the cohort's mean score.

Because the AI moderation is standardized and bias-aware, the qualitative evidence is genuinely comparable across courses, which is precisely the property that makes cross-unit comparison defensible in the first place. The same interview engine underpins the main Koji platform for customer and user research, where teams face the identical trap of over-reading a small sample's average.

To be precise about the claim: Koji does not make a small cohort's data behave like a large one — no method can. It reduces an institution's reliance on the single most noise-prone statistic in the building, and surfaces the qualitative signal that small samples can actually support.

The takeaway

Ranking instructors and departments on raw means is a fifty-year-old statistical mistake with a fifty-year-old fix. Means built on few responses are unreliable — the measurement theory puts the threshold around 20 raters — and treating them as equivalent to means built on hundreds rewards noise and punishes precision. Empirical-Bayes shrinkage, or its multilevel-model equivalent, pulls unreliable estimates toward the average in proportion to their uncertainty, reports honest intervals, and stops class size from masquerading as teaching quality. Every field that compares many noisily-measured units adopted it decades ago. Course evaluation should too.

Comparing courses and programmes across very different cohort sizes? See how Koji for Education builds comparable, defensible evidence instead of ranking fragile averages.