New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
analysis-reporting10 min read

Why a Low-Scoring Course Usually 'Improves' Next Year Even If Nothing Changed: Regression to the Mean in Course Evaluations

Year-over-year movements in course-evaluation scores are mostly statistical noise regressing toward an instructor's true average. Here is why regression to the mean fools quality-assurance committees, and how to read evaluation trends responsibly.

Koji Education Team

Product

Bottom line: Course-evaluation scores swing from year to year largely because a single term's mean is measured with substantial error. When you flag a course because its score was unusually low, statistics alone predict it will rise next term — back toward the instructor's true average — whether or not anyone intervened. This is regression to the mean (RTM), and mistaking it for the effect of a quality-improvement action is one of the most common errors in how universities read evaluation trends. Treat a single dip (or spike) as noise until repeated measurement, or a controlled comparison, says otherwise.

What regression to the mean actually is

Regression to the mean was first described by Francis Galton (1886), who noticed that tall parents tend to have children shorter than themselves, and short parents children taller than themselves: heights "regress" toward the population average. The mechanism is not biological — it is measurement. Any observed score mixes a stable true component with transient noise: a brutal exam week, one unusually vocal cohort, a heatwave during the survey window, a timetable clash that soured the room. When you select a measurement because it is extreme, you have disproportionately picked cases where the noise happened to push the score far from the truth. On the next independent measurement that particular noise is gone, and the score drifts back toward the true value.

Daniel Kahneman (2011) gives the canonical applied example in Thinking, Fast and Slow. Israeli flight instructors were convinced that praising a cadet for a good manoeuvre produced a worse next attempt, while criticising a bad manoeuvre produced improvement — so criticism "worked" and praise "backfired." They had inverted cause and effect. An exceptionally good flight is partly luck; the next flight regresses downward whether or not you praise. An exceptionally bad flight regresses upward whether or not you scold. The instructors had bolted a causal story onto pure RTM. The parallel to course evaluations is exact: a department that launches an improvement plan on a low-rated course and sees the score rise the following year may simply be watching the cadet's next flight.

What the research says about course-evaluation stability

For RTM to bite, two conditions must hold: scores must contain meaningful noise, and courses must be selected for attention based on extreme scores. Both are demonstrably true of Student Evaluation of Teaching (SET).

The reliability literature shows that a single course mean is a noisy estimate of underlying teaching. Feistauer and Richter (2017), analysing 4,224 evaluations from 480 students across three years with cross-classified multilevel models, found that although instructor and course differences explained meaningful variance, "a similar proportion of variance was due to students, and the interaction of students and teachers was the strongest source of variance." In plain terms, a large share of any score reflects which students happened to be in the room and how they personally clicked with the instructor — precisely the term-to-term noise that drives regression.

Longitudinal work tells the same story from the stability side. Marsh (2007), synthesising decades of SET research, reports that evaluations are reasonably stable when aggregated across many sections, but that any single administration is comparatively unreliable; he repeatedly urges averaging over multiple courses before drawing inferences about an instructor. A 2025 analysis in Assessment & Evaluation in Higher Education of more than thirty years of evaluation data from a large public university found the highest reliability when the same instructor taught the same course within the same semester (r ≈ 0.50), with lower reliability across different semesters and a small decline over time (https://doi.org/10.1080/02602938.2025.2504618). A test-retest correlation near 0.5 is exactly the condition under which RTM is large: roughly half of any deviation from an instructor's mean is expected to evaporate on the next measurement.

Combine the two facts — noisy single-term means, plus the near-universal habit of flagging the lowest-rated courses for remediation — and you get a near-guarantee that flagged courses will appear to improve next year, and celebrated courses will appear to slip, with or without any intervention.

Why it matters for course evaluation in practice

The practical danger is false attribution. A quality cycle that selects the bottom decile of courses, applies an intervention, and measures again next year is structurally designed to manufacture the appearance of success. Because the selected courses were chosen for being extreme, most will rise toward their true mean regardless. The committee congratulates the intervention; the intervention may have done nothing. Worse, the same logic runs in reverse for praised courses: a department that rewards this year's top scorers and watches several "decline" may wrongly conclude the recognition bred complacency.

Three concrete consequences follow:

  1. Uncontrolled before-after comparisons cannot evaluate QA actions. If you want to know whether peer mentoring, a curriculum redesign, or a new assessment format actually improved teaching, you need a comparison group that was not selected on an extreme score — otherwise RTM and the intervention are hopelessly confounded.
  2. "Improvement required" thresholds penalise noise. Triggering a formal review whenever a course falls below a fixed cut-off will repeatedly snare instructors who had one statistically unlucky term, then "vindicate" them next year. That is administrative churn, not quality assurance, and it erodes faculty trust in the whole system.
  3. Trend lines need baselines, not anecdotes. A two-point time series (last year vs this year) is the worst possible evidence, because it maximises exposure to RTM. Multi-year baselines, control-chart style limits, and explicit uncertainty are the antidote.

Limitations and honest caveats

RTM is not a claim that interventions never work. It is a claim that uncontrolled before-after comparison cannot tell a real effect apart from statistical drift. Three caveats keep the argument honest:

  • Magnitude depends on reliability. The more reliable the score (more respondents, more aggregation across sections), the smaller the regression. A programme that already pools several cohorts before acting has less RTM to worry about than one reacting to a single 18-student seminar.
  • Real trends exist. Genuine deterioration — a course that has drifted out of date, an instructor under strain — produces sustained, directional movement, not a one-term blip. RTM is the reason you should require persistence before believing a trend, not a reason to ignore all change.
  • Selection is the trigger, not measurement alone. If you look at every course's change with no selection on extremes, the regression effects across the cohort roughly cancel. RTM becomes a problem specifically when you act on, or report, the tails. Most QA processes do exactly that, which is why it matters here.

A PhD reader will also note that RTM interacts with other SET threats — nonresponse swings, cohort composition, discipline and class-size effects — so the cleanest defence is never a single statistic but triangulation across cycles and methods.

How Koji incorporates this

Koji is built so that an evaluation does not have to be read as a single fragile number, which is what makes RTM so easy to mis-read.

  • Multi-cycle aggregation and confidence-aware reporting. Koji reports an instructor's or course's results across cohorts and cycles, not just the latest term, and surfaces the number of respondents and the uncertainty around a mean. A dashboard that shows "4.1 this term, 95% interval 3.7-4.5, against a three-year mean of 4.2" makes it obvious that a dip is within noise — the visual equivalent of refusing to over-read one flight.
  • Small-sample flagging. Because RTM is largest when reliability is low, Koji marks low-response or small-cohort results so committees do not trigger a formal review on a statistically unlucky seminar. This is designed to mitigate the "improvement required threshold" trap, not to hide weak scores.
  • Mechanism, not just movement. Koji's AI-moderated conversational interviews probe why a score moved — following up on a low scale rating with open_ended questions, so you can distinguish a genuine, describable change ("the new group project had no support") from undifferentiated dissatisfaction. A real cause that recurs across students is the signal that something more than regression is happening.
  • Closing-the-loop action tracking with honest baselines. When a department acts on a course, Koji's action tracking records the intervention against the multi-year baseline and, where the design allows, an unselected comparison cohort — so the credit (or blame) is attached to evidence rather than to the next term's inevitable bounce. Koji frames this as designed to reduce false attribution; it cannot eliminate confounding that only a controlled study could resolve.

The same discipline applies beyond the classroom: Koji's core research platform at koji.so applies the same AI-moderated interview engine and confidence-aware reporting to product and customer research, where regression to the mean fools teams reading week-over-week NPS just as readily as it fools QA committees reading SET.

Related Resources

References

  • Galton, F. (1886). Regression towards mediocrity in hereditary stature. Journal of the Anthropological Institute of Great Britain and Ireland, 15, 246-263. https://doi.org/10.2307/2841583
  • Kahneman, D. (2011). Thinking, Fast and Slow. New York: Farrar, Straus and Giroux. (Regression to the mean and the flight-instructor example, Ch. 17.)
  • Feistauer, D., & Richter, T. (2017). How reliable are students' evaluations of teaching quality? A variance components approach. Assessment & Evaluation in Higher Education, 42(8), 1263-1279. https://doi.org/10.1080/02602938.2016.1261083
  • Marsh, H. W. (2007). Students' evaluations of university teaching: Dimensionality, reliability, validity, potential biases and usefulness. In R. P. Perry & J. C. Smart (Eds.), The Scholarship of Teaching and Learning in Higher Education: An Evidence-Based Perspective (pp. 319-383). Springer. https://doi.org/10.1007/1-4020-5742-3_9
  • The reliability of student evaluations of teaching (2025). Assessment & Evaluation in Higher Education, 50(7). https://doi.org/10.1080/02602938.2025.2504618

Related articles

analysis-reporting

How Many Responses Do You Need for a Reliable Course Evaluation?

Reliability of a class-average rating rises steeply with the number of students who respond. Generalizability theory — from Marsh''s syntheses to Dzakadzie (2026) — shows raters matter more than items, and a class of 12 cannot be evaluated like a class of 120.

analysis-reporting

Generalizability Theory and the Reliability of Student Ratings: Why One Class Is Not Enough

Reliability is not one number. Generalizability theory (Gillmore, Kane & Naccarato 1978; Marsh 1984) decomposes the variance in student ratings into student, teacher, course and occasion components — and shows that a single class can be reliable for the course yet a poor estimate of the teacher. What that means for fair evaluation.

analysis-reporting

The 4.2 vs 4.4 Trap: Why Small Differences in Course-Evaluation Means Are Usually Noise

Faculty and administrators routinely read meaning into tiny gaps between course-evaluation means. Boysen (2015) and Boysen et al. (2014) show this happens even when confidence intervals say the difference is nothing — and that warnings barely help. How to report uncertainty honestly.

analysis-reporting

Even Fair Evaluations Can Rank Unfairly: Why Course-Evaluation Scores Misclassify Instructors

A simulation-based finding that even unbiased, reliable, and valid course-evaluation scores produce high misclassification rates when used to rank or compare instructors — and what it means for tenure, promotion, and merit decisions.