Did Your Teaching Improve, or Is It Just Regression to the Mean?
A course that scored badly last year usually scores better this year — even if nothing changed. Why regression to the mean quietly corrupts year-over-year course evaluation comparisons, and how to read the numbers honestly.
Koji Education Team
Product · June 17, 2026
Short answer: When a course or instructor records an unusually low evaluation score one year, the next score will very probably be higher — and an unusually high score will very probably be followed by a lower one — even if nothing about the teaching changed at all. This is regression to the mean, a statistical certainty wherever a measure contains random noise. It is one of the most common ways that quality-assurance committees fool themselves into seeing improvement, decline, or the "effect" of an intervention that did nothing. If you compare this year's number to last year's and draw a conclusion, you are probably measuring noise.
The 150-year-old trap your committee keeps falling into
The phenomenon was named by Francis Galton in 1886, who noticed that tall parents tend to have somewhat shorter children and short parents somewhat taller ones — the offspring "regress" toward the population average. The mechanism is general: whenever an observed value is part true signal and part luck, an extreme observation owes some of its extremity to luck, and luck does not repeat. The next measurement therefore drifts back toward the mean.
Daniel Kahneman tells the canonical version in Thinking, Fast and Slow (2011). Addressing Israeli Air Force flight instructors, he argued praise works better than punishment. A veteran instructor objected: every time he praised a cadet for a clean manoeuvre, the next attempt was worse; every time he screamed at a cadet for a bad one, the next was better. The instructor concluded that punishment works and praise backfires. He was wrong, and the reason is pure regression to the mean: an exceptionally good manoeuvre is, by definition, near the top of that cadet's range and will usually be followed by something more average; an exceptionally bad one will usually be followed by something better — regardless of what the instructor said. Kahneman called the realisation "one of the most satisfying eureka experiences of my career."
Now replace "manoeuvre" with "course evaluation score" and "instructor's reaction" with "your improvement plan." A module scores 3.4 — alarmingly low. The department acts: a teaching consultation, a redesign, a stern conversation. Next year it scores 3.9. Everyone congratulates the intervention. But a module that scored 3.4 was, in part, unlucky — a difficult cohort, a clash with an exam, a couple of disengaged respondents in a small class. Remove the bad luck and it would have scored higher anyway. You may have changed nothing and seen "improvement," or done genuine good work and been unable to tell how much.
This is not a fringe worry — it has been measured
The effect is well documented in the closely related field of teacher value-added measurement. In a study of the Los Angeles Times's public release of teacher ratings (Pope, 2019, Journal of Public Economics), researchers had to explicitly model regression to the mean to separate real responses to the ratings from the statistical artifact: low-rated teachers' student scores tended to rise the following year, exactly the pattern regression predicts before any behaviour change is invoked. The methodological lesson transfers directly to course evaluations: low scores tend to bounce up and high scores tend to settle down, and any naive before/after comparison silently bakes that in.
The smaller your class and the noisier your instrument, the larger the regression. A class of twelve can swing half a point on the mood of two students. A five-point Likert scale already compresses almost everyone into the 3.5–4.5 band, so the random component is a large fraction of the year-to-year movement you are trying to interpret. This compounds the deeper problem we describe in why averaging Likert scores misleads: the number you are tracking is a fragile summary to begin with.
But doesn't this mean evaluation data is useless?
No — and this is the counterargument worth taking seriously, because the lazy reading of "regression to the mean" is nihilism. Three honest responses:
-
Real change exists; regression just sets the baseline you must beat. Teaching genuinely improves and declines. The point is that the expected movement from an extreme score is non-zero even with no real change, so improvement should be judged against that expected rebound, not against zero. A low course that rebounds exactly as much as regression predicts has shown no evidence of improvement; one that rebounds substantially more has.
-
You can design around it. Do not select courses for intervention solely because they hit a one-year low, then judge success by the next single year — that is the maximally regression-prone design. Use multiple years, control or comparison groups, and confidence intervals around differences (we cover the inference side in small mean differences and confidence intervals).
-
Richer evidence is less regressive. Random noise dominates thin, single-number measures. Qualitative themes that recur across cohorts — "the assessment brief was unclear," "the lab was disconnected from the lectures" — are far more stable than a mean, because they describe structural features of the course rather than the luck of one sample.
How to read year-over-year scores without fooling yourself
- Never act on a single year's extreme in isolation. Look at three or more years; a genuine problem persists, a fluke does not.
- Expect the rebound and quantify it. If low scores in your institution typically recover by, say, 0.3 points on their own, an intervention needs to clear that bar before you credit it.
- Put uncertainty on the difference, not just the level. "3.4 to 3.9" with overlapping confidence intervals is not a finding.
- Trust recurring themes over moving averages. Structural feedback is the signal that survives resampling.
- Use a comparison group where you can — similar untreated modules regress too, and the difference in regression is your real effect.
Where Koji fits
Koji does not repeal statistics — regression to the mean applies to any measure, including ours. What Koji changes is the ratio of signal to noise in what you collect, which is the only lever that actually shrinks the problem. Instead of a single averaged Likert number per course, Koji's AI-moderated conversational interviews surface why a course scored as it did, and its automatic thematic analysis identifies issues that recur across cohorts — the stable, structural signal that does not evaporate when you resample students. When a low score is driven by a concrete, repeatable cause ("the group project marking felt arbitrary"), you can see it is real rather than guessing from a number that may simply rebound next year. Quality scoring and programme-level reporting let you watch themes across multiple cycles rather than chasing one year's mean, and formative mid-cycle collection means you can act on a problem while the cohort is still in the room — not gamble a personnel decision on a noisy annual figure. The same conversational engine powers the main Koji platform for general research, where the same logic — themes are more stable than averages — applies.
If your improvement reports rest on comparing one year's average to the next, regression to the mean is quietly writing some of your conclusions. See how Koji for Education helps you measure the structural signal instead.
A worked example you can show a committee
Picture twenty modules that each scored 3.3 last year — the bottom of your distribution. You intervene in all of them. Next year their average is 3.7. Triumph? Not necessarily. Take twenty other modules that scored 3.3 and did nothing: regression to the mean predicts they will also rise, perhaps to 3.6, purely because last year's 3.3 was partly bad luck that does not recur. Your intervention's real effect is the difference between the treated rebound (3.7) and the untreated rebound (3.6) — a tenth of a point, not the four-tenths the naive before/after comparison claimed. Without that comparison group you would have credited the programme with roughly four times the effect it produced. This is why the most confident-sounding improvement stories — "we acted on the feedback and scores jumped" — are often the least trustworthy: they are exactly the design in which regression masquerades as success.
The bottom line
A low score that rises and a high score that falls are the default expectation of any noisy measure, not evidence of anything. Before you credit an intervention or panic over a dip, ask whether you are looking at teaching — or at Galton's ghost. The defence is more years, more uncertainty, and richer, theme-level evidence that does not regress.