New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
analysis-reporting10 min read

Statistical Process Control for Course Evaluations: Using Control Charts to Separate Signal from Noise

How Shewhart control charts and statistical process control distinguish real changes in teaching quality from ordinary random variation in course evaluation scores - so QA teams stop reacting to noise. Grounded in Sivena and Nikolaidis (2019, 2022).

Koji Education Team

Product

In brief: Most year-to-year wobble in a course-evaluation mean is ordinary random variation, not a real change in teaching. Statistical process control (SPC) - specifically Shewhart-style control charts - gives quality-assurance teams a principled rule for telling common-cause noise from a special-cause signal, so they act on genuine shifts and ignore the rest. Sivena and Nikolaidis (2019, 2022) built and validated exactly this framework for student evaluations in higher education.

What the research says

Statistical process control originates with Walter Shewhart in the 1920s and was popularised by W. Edwards Deming: any repeatedly measured process shows variation, and the central task is to distinguish common-cause variation (the inherent, stable noise of the process) from special-cause variation (a real, assignable shift). A control chart plots the metric over time with a centre line (the process mean) and upper and lower control limits, conventionally set around three standard deviations from the centre. Points inside the limits, with no non-random patterns, indicate a stable process that should be left alone; points outside the limits, or runs and trends, flag a special cause worth investigating.

Sofia Sivena and Yiannis Nikolaidis brought this machinery directly to course evaluations. In Communications in Statistics - Simulation and Computation (2019; volume 51, pages 1289-1312) they proposed a framework using average and dispersion charts (X-bar and S charts) to monitor faculty teaching-evaluation ratings, and used simulation to compare candidate chart types and identify which perform best for the skewed, bounded, small-sample data that SET produces. Their stated aim is to give decision-makers a reliable tool to monitor the teaching process and to distinguish effective from ineffective performance without over-reacting to routine fluctuation. In a 2022 follow-up (same journal; volume 53, pages 2291-2315) they revisited and expanded the framework, adding tools to compare faculty members in pairs and refining chart selection.

The reason this matters is that naive SET reporting routinely violates the common-cause/special-cause distinction. A department reads a drop from 4.3 to 4.0 as a teaching problem and a rise back to 4.2 as a recovery, when all three numbers may sit comfortably within the normal noise band of a small class. This is closely related to regression to the mean: extreme scores in one year tend to move back toward the average the next, purely by chance, and SPC makes that expectation explicit rather than surprising.

Why it matters for course evaluation in practice

Control charts change three things about how a QA office operates:

  1. Fewer false alarms, fewer missed signals. A three-sigma rule calibrates the trade-off between chasing noise (treating random dips as failures) and missing a real decline. Reacting to every fluctuation wastes staff time, demoralises faculty, and erodes trust in the evaluation system; SPC replaces gut-feel thresholds with a defensible statistical rule.
  2. Fair, noise-aware comparison over time. Because control limits widen automatically for small classes (where the standard error is larger), SPC bakes in the uncertainty that raw mean comparisons ignore. A small seminar bouncing between 3.8 and 4.6 may be perfectly in control; a large course drifting by 0.3 over four terms may be a genuine special cause. This is the same insight behind confidence intervals and empirical-Bayes shrinkage, expressed as a monitoring tool.
  3. A quality-cycle narrative accreditors respect. ESG/ENQA-aligned quality assurance asks institutions to show they monitor and act on evidence systematically. "We flag a course when its evaluation trend breaches control limits, then investigate the assignable cause" is a far stronger process claim than "we look at the numbers each year." SPC operationalises continuous improvement rather than reactive fire-fighting.

Limitations and honest caveats

A careful reader should not treat control charts as a cure-all:

  • Garbage in, garbage out. SPC monitors the stability of a metric; it says nothing about whether the metric is valid. If SET only weakly reflects learning (Uttl, White and Gonzalez, 2017, put the shared variance at roughly 1%), a beautifully in-control chart may be faithfully tracking a poor proxy for teaching quality. SPC disciplines interpretation of the number; it does not fix the number.
  • Distributional assumptions. Classic Shewhart limits assume roughly independent, identically distributed data. SET scores are bounded, ceiling-prone and skewed, and successive cohorts are not identical. This is precisely why Sivena and Nikolaidis used simulation to choose robust chart types - but any off-the-shelf X-bar chart applied naively can mislead.
  • Small numbers. Many courses have few respondents, so control limits are wide and only large shifts trip them. SPC will rightly tell you that most small-class movements are uninformative - which some stakeholders find unsatisfying.
  • Special cause is not diagnosis. A breach flags that something changed; it does not say what. The change could be teaching, but it could equally be a new room, a timetable move, a syllabus change, a marking policy, or a response-rate shift. SPC starts the investigation; it does not end it.
  • Gaming risk. Once a control limit becomes a stakes threshold, Campbell and Goodhart pressures apply and the indicator can be managed rather than improved.

How Koji incorporates this

Koji is an AI-native course-evaluation platform, and its reporting layer is built around the same signal-versus-noise discipline that SPC formalises.

Report uncertainty, not just means. Koji surfaces distributions, dispersion and respondent counts alongside the average, so a small-class swing is presented with its inherent uncertainty rather than as a false precision. This is the reporting analogue of a control chart widening its limits for small samples - it discourages readers from treating a noisy dip as a verdict.

Trend monitoring across cohorts and terms. Because Koji retains structured evaluation data over time, teams can watch a course longitudinally and ask whether a movement is a stable-process wobble or a sustained shift - the practical question a control chart answers. Triangulating the numeric trend against the thematic trend in open-text feedback is a strong special-cause test: if the mean drops but the teaching-related themes are unchanged, the cause is likely construct-irrelevant (a room, a slot, a cohort), not teaching.

Explaining the assignable cause. SPC flags that something changed but not what; this is where Koji's AI-moderated conversational interviews earn their keep. When a course trend breaches expectation, the open-text and probe data give the why - separating a genuine teaching change from a syllabus, logistics or grading shift - so the investigation a control chart triggers actually reaches a diagnosis. Koji frames these as candidate explanations for human judgement, never as automatic scoring of an instructor.

Used together, control-chart thinking plus Koji's conversational depth turn evaluation reporting from reactive number-watching into a defensible quality cycle. Koji's core research platform at koji.so applies the same interview engine to product and customer research, where separating signal from noise in tracking metrics is equally central.

Reading a control chart in practice

A worked example makes the discipline concrete. Suppose a module has been evaluated for eight consecutive terms with mean overall-satisfaction scores of 4.2, 4.4, 4.1, 4.3, 4.2, 3.9, 4.3 and 4.2 on a five-point scale, with 25-35 respondents each term. Plotting these with a centre line at roughly 4.2 and three-sigma control limits of, say, 3.7 and 4.7, every point falls inside the limits. The term at 3.9 that triggered a worried email is, statistically, ordinary common-cause variation - the process is stable and no action is warranted. Now suppose the next three terms read 3.6, 3.5 and 3.4: two points below the lower limit and a downward run of three. That is a special-cause signal worth investigating.

Beyond the three-sigma rule, practitioners use supplementary run rules (often called the Western Electric or Nelson rules) to catch shifts the limits alone miss - for example, several consecutive points on one side of the centre line, or a sustained upward or downward trend. These increase sensitivity to gradual drift at the cost of a slightly higher false-alarm rate, which is exactly the trade-off Sivena and Nikolaidis (2019, 2022) examined by simulation when choosing chart types for the skewed, small-sample reality of SET.

Two cautions keep the tool honest. First, control limits should be estimated from a stable baseline period, not recomputed every term, or a slow decline can silently move the limits down with it and hide itself. Second, a chart for a small seminar will have very wide limits, so only large shifts register - which correctly signals that small-class scores carry little information term to term, and should not be over-interpreted. The practical payoff is cultural as much as statistical: a QA office that adopts control-chart thinking stops treating every wobble as a verdict and reserves its limited attention, and its faculty's goodwill, for changes that are real.

Related resources

References

  • Sivena, S., & Nikolaidis, Y. (2019). Improving the quality of Higher Education teaching through the exploitation of student evaluations and the use of control charts. Communications in Statistics - Simulation and Computation, 51(3), 1289-1312. https://doi.org/10.1080/03610918.2019.1667390
  • Sivena, S., & Nikolaidis, Y. (2022). Higher education teaching and exploitation of student evaluations through the use of control charts: Revisited and expanded. Communications in Statistics - Simulation and Computation, 53(5), 2291-2315. https://doi.org/10.1080/03610918.2022.2074457
  • Uttl, B., White, C. A., & Gonzalez, D. W. (2017). Meta-analysis of faculty's teaching effectiveness: Student evaluation of teaching ratings and student learning are not related. Studies in Educational Evaluation, 54, 22-42. https://doi.org/10.1016/j.stueduc.2016.08.007
  • Deming, W. E. (1986). Out of the Crisis. MIT Press. (Foundational treatment of common-cause vs special-cause variation.)

Related articles

analysis-reporting

Can You Fairly Rank Instructors by Their Course-Evaluation Scores? What Esarey & Valdes (2020) Proved

Even if course evaluations were unbiased, reliable, and valid, ranking instructors by their scores would still misclassify many good teachers. A walk through the Esarey & Valdes (2020) simulation and what it means for how you report and use evaluation data.

analysis-reporting

Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees

The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.

analysis-reporting

The 4.2 vs 4.4 Trap: Why Small Differences in Course-Evaluation Means Are Usually Noise

Faculty and administrators routinely read meaning into tiny gaps between course-evaluation means. Boysen (2015) and Boysen et al. (2014) show this happens even when confidence intervals say the difference is nothing — and that warnings barely help. How to report uncertainty honestly.

analysis-reporting

Why a Low-Scoring Course Usually 'Improves' Next Year Even If Nothing Changed: Regression to the Mean in Course Evaluations

Year-over-year movements in course-evaluation scores are mostly statistical noise regressing toward an instructor's true average. Here is why regression to the mean fools quality-assurance committees, and how to read evaluation trends responsibly.