Which Instructor Scores Are Genuinely Unusual? Funnel Plots for Fair Course-Evaluation Comparison
Ranking instructors by mean evaluation score turns sampling noise into a league table. Funnel plots — Spiegelhalter's method for institutional comparison — show which scores are genuinely unusual and which are just small-sample wobble.
Koji Education Team
Product
In brief
If you sort instructors or courses by their mean evaluation score and read the order as a ranking of quality, you are almost certainly ranking noise. A course with 12 respondents and a course with 300 respondents do not carry the same precision, yet a raw league table treats them identically. The funnel plot, introduced for institutional comparison by David Spiegelhalter (2005), plots each unit''s score against a measure of its precision (typically the number of responses) and draws control limits that fan out — a funnel — for smaller samples. Points inside the funnel are consistent with common-cause variation; only points outside it are genuinely unusual. Funnel plots are "flexible, attractively simple, and avoid spurious ranking of institutions into league tables" (Spiegelhalter, 2005). For course evaluation, they answer the right question: not "who is top?" but "who is meaningfully different from the benchmark, given how much data we actually have?"
What the research says
Spiegelhalter''s "Funnel plots for comparing institutional performance" (Statistics in Medicine, 2005) was written to fix a specific pathology in how organisations are compared. When hospitals, schools, or surgeons are ranked on an outcome and placed in a league table, most of the apparent spread between the middle ranks is sampling variation, not real difference. Ranking forces a total order onto data that cannot support one, and small units — with the widest uncertainty — bounce to the extremes of the table purely by chance.
The funnel plot is his proposed remedy. The construction is deliberately simple:
- Plot each unit''s outcome (here, mean evaluation score, or percent-favourable) on the vertical axis.
- Plot a measure of that estimate''s precision on the horizontal axis — for course evaluation, the number of respondents. More responses means more precision.
- Draw a horizontal target line at the benchmark (for example the departmental or institutional mean).
- Draw control limits (commonly 95% and 99.8%, echoing Shewhart control charts) around the target. Because the standard error of a mean shrinks as the sample grows, these limits are wide at the left (few responses) and narrow at the right (many responses) — producing the characteristic funnel shape.
Units falling inside the funnel are statistically indistinguishable from the benchmark: their departure from the target is within what sampling alone would produce. Units falling outside are the genuine signals worth investigating. Critically, the funnel plot shows this without imposing a rank order — it separates "different from the benchmark" from "ranked 4th versus 5th," which is exactly the distinction a fair review needs.
Spiegelhalter also addressed over-dispersion: when there is more variation between units than pure sampling predicts (because of unmeasured real differences), naïve limits flag too many units as outliers. He described how to widen the limits to account for over-dispersion so that only truly exceptional units are highlighted.
Education research reached the same conclusion by a different route. Leckie and Goldstein (2009), analysing English secondary-school league tables in the Journal of the Royal Statistical Society, Series A, found that once the uncertainty in performance estimates is properly incorporated, only a handful of schools can be statistically separated from the average or from one another — the vast majority of confidence intervals overlap. Their verdict on ranking is blunt: league tables "have very little to offer as guides to school choice." The identical logic applies to ranking instructors on evaluation means.
Why it matters for course evaluation in practice
The dominant way institutions consume evaluation data is comparison: this instructor versus the department average, this course versus last year, this tutor versus their peers. Almost every one of those comparisons is made on point estimates with wildly different sample sizes and no uncertainty attached. The consequences are concrete and unfair:
- Small classes dominate the extremes. A seminar with nine responses can easily average 4.9 or 3.1 by chance alone. In a raw ranking it appears either as a star or a problem case, when in fact its score carries almost no information. Funnel plots put it on the wide left of the funnel, where large deviations are expected and therefore not alarming.
- Large courses look "average" and get ignored. A 250-response lecture sitting at 4.1 is a precise, reliable signal, but in a league table it is buried mid-pack. On a funnel plot, if 4.1 sits outside the narrow right-hand limits it is a real, actionable finding.
- Year-on-year "movement" is mostly regression to the mean. An instructor who was "bottom of the table" last year and "middle" this year has often not changed at all. Funnel plots (and their time-series cousin, control charts) reframe the change as within-limits noise unless it crosses a control limit.
- Personnel decisions inherit the noise. When ranked means feed tenure, renewal, or teaching-award shortlists, the funnel plot is the difference between acting on signal and acting on which small class happened to have a good week.
The practical payoff is a discipline of interpretation: do not compare two units inside the funnel as if one is better than the other. Inside the funnel, the honest statement is "indistinguishable from the benchmark." That single rule prevents most of the misuse that ranking invites.
Limitations and honest caveats
Funnel plots are a tool, not an oracle, and a rigorous reader should hold several caveats.
First, the funnel controls for sampling variation, not for confounding. A course can sit outside the funnel because it is genuinely different or because of situational factors — difficulty, class composition, required status — that the plot does not adjust for. Funnel plots tell you a score is unusual; they do not tell you why. They should be paired with contextual interpretation (and are complementary to, not a substitute for, the situational reading discussed in our piece on the fundamental attribution error).
Second, the choice of benchmark is a judgement. Using the grand mean as the target assumes all courses should score the same on average, which is exactly the assumption the confound literature undermines. A defensible funnel plot often benchmarks within comparable strata (same level, similar size, same discipline) rather than against a single institution-wide line.
Third, over-dispersion is common in evaluation data and, if unaddressed, makes the naïve funnel flag far too many outliers. Analysts must test for and, where present, model over-dispersion — otherwise the plot manufactures false signals.
Fourth, ordinal data assumptions. Mean Likert scores are ordinal-treated-as-interval, and their sampling distribution is not perfectly normal, especially with ceiling effects and small n. Percent-favourable (top-box) proportions can be a more defensible outcome for a funnel plot in skewed data, though they discard information. This connects to the broader ordinal-versus-interval debate in evaluation reporting.
Fifth, funnel plots discourage ranking by design — which some stakeholders dislike. Administrators who want a single ordered list will find "these six are inside the funnel and indistinguishable" unsatisfying. That discomfort is the point: the method refuses to certify distinctions the data cannot support.
How Koji incorporates this
Koji is built to report evaluation results as estimates with uncertainty, not as a bare ordered list, so the comparisons an institution actually makes are honest about sample size.
- Precision-aware comparison, not naïve ranking. Koji ties every reported score to its response count and is designed to present comparisons that account for precision — surfacing whether a course''s deviation from a benchmark is within the range explained by its sample size, in the spirit of Spiegelhalter''s control limits. The design intent is to stop small classes from being read as stars or problem cases on the strength of a handful of responses.
- Benchmarking within comparable strata. Rather than comparing every course to one institution-wide mean, Koji supports benchmarking against comparable cohorts (level, size band, discipline), addressing the "wrong target line" limitation and reducing spurious over-dispersion from mixing unlike courses.
- Response-count and reliability signalling. Koji reports how many responses underlie each result and flags when a course has too few responses to support confident comparison — the same information the horizontal axis of a funnel plot encodes — so reviewers do not over-read thinly sampled units.
- Distinguishing signal from movement. For repeated cycles, Koji tracks results over time so that year-on-year change can be read against expected variation rather than as evidence of improvement or decline, complementing statistical-process-control monitoring and guarding against regression-to-the-mean misreadings.
- Richer evidence for the genuine outliers. When a course is outside the expected range, a number alone does not explain it. Koji''s AI-moderated interviews and automatic thematic analysis of open text provide the why behind an outlier — turning a flagged point into an investigable, actionable finding rather than a rank.
Koji does not claim to remove every statistical hazard from comparison; over-dispersion, confounding, and benchmark choice still require analyst judgement. What it is designed to do is make sample size and uncertainty first-class citizens of the report, so the institution stops ranking noise. The same estimate-with-uncertainty philosophy runs through Koji''s core research platform at koji.so, where comparing segments or cohorts on thin samples is an equally common trap.
Related resources
- Empirical Bayes shrinkage for instructor scores
- Statistical process control charts for monitoring evaluations
- Regression to the mean in year-over-year changes
- Small mean differences and confidence intervals
- Is it fair to rank instructors by evaluation scores?
- How many responses make a course evaluation reliable?
References
- Spiegelhalter, D. J. (2005). Funnel plots for comparing institutional performance. Statistics in Medicine, 24(8), 1185–1202. https://doi.org/10.1002/sim.1970
- Leckie, G., & Goldstein, H. (2009). The limitations of using school league tables to inform school choice. Journal of the Royal Statistical Society: Series A (Statistics in Society), 172(4), 835–851. https://doi.org/10.1111/j.1467-985X.2009.00597.x
- Spiegelhalter, D. J. (2005). Handling over-dispersion of performance indicators. Quality and Safety in Health Care, 14(5), 347–351. https://doi.org/10.1136/qshc.2005.013755
Related articles
How Many Responses Do You Need for a Reliable Course Evaluation?
Reliability of a class-average rating rises steeply with the number of students who respond. Generalizability theory — from Marsh''s syntheses to Dzakadzie (2026) — shows raters matter more than items, and a class of 12 cannot be evaluated like a class of 120.
Can You Fairly Rank Instructors by Their Course-Evaluation Scores? What Esarey & Valdes (2020) Proved
Even if course evaluations were unbiased, reliable, and valid, ranking instructors by their scores would still misclassify many good teachers. A walk through the Esarey & Valdes (2020) simulation and what it means for how you report and use evaluation data.
The 4.2 vs 4.4 Trap: Why Small Differences in Course-Evaluation Means Are Usually Noise
Faculty and administrators routinely read meaning into tiny gaps between course-evaluation means. Boysen (2015) and Boysen et al. (2014) show this happens even when confidence intervals say the difference is nothing — and that warnings barely help. How to report uncertainty honestly.
Why a Low-Scoring Course Usually 'Improves' Next Year Even If Nothing Changed: Regression to the Mean in Course Evaluations
Year-over-year movements in course-evaluation scores are mostly statistical noise regressing toward an instructor's true average. Here is why regression to the mean fools quality-assurance committees, and how to read evaluation trends responsibly.