New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Methodology9 min read

Is That Dip Real or Just Noise? Reading Course-Evaluation Trends With Statistical Process Control

A module drops from 4.3 to 4.0 and a review is triggered. But was that a real change or ordinary variation? Statistical process control gives quality offices a disciplined answer — and stops two expensive mistakes at once.

Koji for Education

Research & Editorial Team · July 11, 2026

Answer first: When a module''s evaluation score moves from one term to the next, most quality offices react to the movement itself — a dip triggers a review, a rise earns praise. This is a mistake. Nearly every term-to-term wobble is ordinary noise, and treating noise as signal ("tampering," in W. Edwards Deming''s language) makes systems worse, not better. Statistical process control (SPC) — specifically run charts and control charts — gives you a disciplined rule for telling a real change from random variation. It is one of the most useful, least-used tools in course-evaluation analysis.

Two ways to be wrong about a number that moved

Every time a course score changes, you can make one of two errors:

  • React to noise. You investigate a "decline" that is just the normal scatter you would see even if nothing changed. You waste a programme director''s afternoon, and — worse — you may "fix" something that was never broken, injecting real instability. Deming called this tampering.
  • Miss a signal. You dismiss a genuine deterioration as "probably just variation," and a course quietly gets worse for three years before anyone acts.

You cannot avoid both errors by staring harder at the raw numbers. You need a rule that separates the two kinds of variation. That is exactly what SPC was built to do.

Common cause vs special cause

The foundational insight, from Walter Shewhart in the 1920s and popularised by Deming, is that all processes contain two kinds of variation (ASQ, Control Chart):

  • Common-cause variation is the inherent, stable scatter of a process operating as designed. In course evaluation, it is the term-to-term wobble produced by which students happened to respond, the cohort''s mood, the assessment timing, the weather in exam week. A process showing only common-cause variation is stable — its future behaviour is predictable within limits, and any single point tells you nothing special.
  • Special-cause variation is assignable: something genuinely changed. A new instructor, a redesigned assessment, a room move, a doubled class size. This is the variation worth investigating, because it has a findable cause.

Deming relabelled Shewhart''s "chance" and "assignable" variation as common and special cause precisely to make the management lesson unavoidable: you should investigate special causes and leave common causes alone (LifeQI, common vs special cause). Reacting to common-cause variation as if it were special is the single most common analytical error in quality management — and course-evaluation dashboards invite it on every screen.

What a control chart does

A control chart plots a metric over time with three reference lines: a centre line (the process mean) and upper and lower control limits, typically set three standard deviations out. As long as points fall inside the limits and show no non-random patterns, the process is "in control" — the variation is common-cause, and no point deserves a special explanation. When a point breaches a limit, or a run of points trends or clusters in a way chance would rarely produce, the process is signalling a special cause worth investigating.

Applied to a module evaluated every term, this reframes the entire conversation. Instead of "the score fell 0.3, why?" the question becomes "is this point outside what this course''s own history would predict?" Most of the time the answer is no — and the correct action is to do nothing, which is itself a finding. Occasionally the answer is yes, and now you are investigating a change that is real rather than chasing ghosts. This is the same logic behind warnings about regression to the mean and the multiple-comparisons trap, given an operational rule.

"But a university is not a factory" — the honest objections

SPC came from manufacturing, and a PhD audience will immediately push back. The objections are fair and worth stating plainly.

"Teaching outcomes are not widgets; three-sigma limits assume a stable, high-volume process." True. A module run once a year to a class of 15 does not generate the data volume SPC was designed for, and control limits built on a handful of points are themselves noisy. This is a real limitation, not a quibble. The response is proportionate: use SPC as a disciplining heuristic — a rule that stops you over-reacting — not as a mechanical trigger. For small classes, the wide, honest uncertainty a control chart shows is a feature: it tells you that most apparent movement is not interpretable, which is the correct conclusion.

"Evaluation scores are bounded, skewed Likert averages, not normal measurements." Correct — ceiling effects and restricted range violate the tidy assumptions. Run charts (which use non-parametric rules about runs and trends rather than standard-deviation limits) are more defensible here than classic Shewhart charts, and modern variants exist for bounded proportions. The point is not to import a factory formula uncritically; it is to adopt the habit of asking whether a movement exceeds the process''s natural variation before you act.

"This just adds statistical theatre to a number that is already the wrong measure." The strongest objection of all. If the underlying metric — an averaged Likert score — is a poor proxy for teaching quality, then charting it precisely does not make it valid. SPC controls for noise over time; it does nothing about construct validity. A course can sit perfectly in control and still be measured by the wrong instrument. SPC is necessary discipline, not sufficient insight.

Where Koji fits

SPC tells you when a change is real. It cannot tell you what changed or why — for that you need evidence with more resolution than a single trended number. This is where Koji for Education complements the discipline. Because Koji collects formative, mid-cycle feedback rather than only an end-of-term score, a special-cause signal can be caught and diagnosed while the course is still running, not a term later in a post-mortem. When a control chart flags a real shift, Koji''s thematic analysis of open-text and its AI-moderated conversational interviews supply the assignable cause — the redesigned assessment, the pacing change, the new tutor — instead of leaving a committee to guess. And Koji''s closing-the-loop action tracking and programme-level reporting let you check whether the intervention actually moved the process, rather than declaring victory on the next lucky data point (which is just regression to the mean wearing a medal).

Used together, the logic is clean: SPC decides whether a number deserves attention; Koji supplies the reasoning and confirms whether your fix held. The same conversational interview engine powers general research on the main Koji platform for teams who track other metrics over time.

The next time a module''s score dips, resist the reflex to convene a review. Ask first whether the movement exceeds what this course''s own history would predict. Most of the time it does not — and knowing that, reliably, is worth more than another dashboard.

The run-chart rules worth knowing

You do not need the full apparatus of three-sigma control limits to get most of SPC''s protective value. A simple run chart — the metric plotted over time against its median — supports a handful of non-parametric rules that flag non-random behaviour without assuming a normal distribution, which makes them a better fit for bounded Likert data:

  • A shift: roughly seven or eight consecutive points all on the same side of the median. A single low term is noise; eight low terms in a row is not.
  • A trend: six or more points in a row all rising or all falling. Steady drift is the pattern most likely to be a real, findable change — and the one raw dashboards miss because each individual step looks small.
  • Too few or too many runs: if the line crosses the median far less often than chance predicts, the process is not behaving randomly.

The discipline these rules impose is deliberately conservative: they tell you to wait for a pattern before acting, which is exactly the brake a reactive quality culture needs. A module that dropped 0.3 this term and rose 0.2 last term has produced no pattern at all, and the correct response is to keep watching, not to convene. Applied consistently across a programme, run-chart rules turn a wall of twitchy numbers into a short, defensible list of courses whose behaviour has actually changed — the only ones worth a director''s time.

Turn evaluation trends into evidence you can act on — see Koji for Education.