New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Methodology10 min read

Flag Enough Instructors and Some Will Look Bad by Chance: The Multiple-Comparisons Trap in Evaluation Dashboards

When an evaluation dashboard scans hundreds of instructors and dozens of items for "concerning" scores, the mathematics of multiple comparisons guarantees false alarms — even if every instructor is identical. Why your red flags are noisier than they look, and what to do instead.

Koji Education Team

Product ·

Bottom line up front: Modern course-evaluation dashboards encourage a dangerous habit: scanning every instructor against a benchmark and flagging whoever falls "below." The problem is statistical, not philosophical. When you run hundreds of comparisons at once — every instructor, every item, every subgroup — some will look alarming purely by chance, even in a world where every instructor is genuinely identical. This is the multiple-comparisons problem, it is one of the best-understood failure modes in all of statistics, and almost no institutional evaluation process controls for it. The result is a steady stream of false flags that waste managers'' time, frighten good teachers, and quietly corrode trust in the whole system.

The arithmetic nobody runs

Start with a thought experiment. Suppose every instructor in a faculty of 100 is exactly equally effective, and the only thing that varies between their evaluation scores is random noise — different students, different days, different moods. Now you flag anyone whose score is "significantly" below the mean at the conventional 5% threshold.

How many false flags do you expect? The maths is unforgiving. If you run 100 independent comparisons at a 5% error rate, you expect five false positives on average — five instructors flagged as problems who are, by construction, no different from anyone else. Worse, the probability of getting at least one false flag is 1 − 0.95^100 ≈ 99.4%. You are essentially guaranteed to manufacture a problem out of pure noise. Scale that to a university with thousands of teaching staff evaluated on a dozen items each, and you are running tens of thousands of comparisons every semester with no error-rate control whatsoever.

This is exactly the scenario that statisticians Benjamini and Hochberg confronted in their landmark 1995 paper introducing false discovery rate (FDR) control (Journal of the Royal Statistical Society, Series B). Their insight: when you test many hypotheses at once, you must adjust your threshold to control the proportion of your "discoveries" that are false. Genomics, neuroimaging, and clinical research adopted this discipline decades ago because they routinely test thousands of hypotheses. Course-evaluation analytics, which does the same thing, almost never has.

"But we don''t run significance tests — we just flag below-average scores"

This is the most common objection, and it is worse than the problem it dodges. Ranking instructors against a mean and flagging the bottom guarantees that roughly half of all instructors are "below average" every single cycle, by definition, no matter how good they are. You have not avoided the multiple-comparisons problem; you have replaced an explicit, controllable error rate with an implicit one that flags people for the crime of not being above the median.

It gets worse once you add time. Scores that are low this semester are disproportionately likely to be low because of bad luck this semester — a small class, an unlucky draw of disgruntled respondents, a scheduling clash. By regression to the mean, those same instructors will tend to rebound next cycle with no intervention at all. If you flag them, send them to a teaching-development conversation, and they improve, you will credit the intervention for an effect that was going to happen anyway — and you will have spent your scarce development resources on people who were never the problem.

How unreliable a single mean really is

The multiple-comparisons trap is amplified by how noisy each individual score is to begin with. In a frequently cited analysis, An Evaluation of Course Evaluations by Philip Stark and Richard Freishtat (2014), the authors simulated a generous scenario in which evaluation scores were moderately correlated with actual learning (r = 0.4) — and found that under those conditions, comparing two instructors by their average ratings identified the wrong instructor as the better teacher about 37% of the time. Their blunt methodological conclusion was that averages of rating scores "should not even be calculated, much less compared across instructors, courses, or departments."

If a single head-to-head comparison is wrong more than a third of the time, a dashboard performing hundreds of them simultaneously is not surfacing signal. It is industrialising noise. And the noise is largest exactly where institutions most want certainty: small classes, where a couple of extreme responses swing the mean, and cross-department benchmarking, where disciplinary norms differ enough that "below benchmark" means something different in engineering than in history.

A worked example

Picture a faculty dashboard with twelve survey items, displayed for forty instructors, each broken down by three student subgroups — first-year, returning, and international. That is 12 x 40 x 3 = 1,440 cells, every one of them implicitly compared against a benchmark. At a 5% flag rate, you would expect roughly seventy "concerning" cells to light up even if every instructor taught every group identically well. A manager scanning that screen sees seventy problems; the statistics says they are looking at the expected output of pure noise. The more granular and well-intentioned the dashboard becomes — more items, more cuts, more subgroups — the more false alarms it manufactures, and the more genuine signal drowns among them. Granularity feels like rigour. Statistically, it is the opposite.

What good practice looks like

You cannot make this problem vanish, but you can stop amplifying it.

  1. Stop auto-flagging on point estimates. A score is an estimate with uncertainty around it. A flag should require a difference that survives an honest look at effect size and confidence intervals, not merely a rank below a line.
  2. Use shrinkage, not raw means. Multilevel models pull noisy individual averages toward the group mean in proportion to how little data supports them — the principled antidote to over-reading a class of twelve.
  3. Treat flags as hypotheses, never verdicts. A low score is a reason to look, not a finding. The only legitimate output of a screen is a question to investigate with richer evidence — which is also why student evaluations should not, on their own, drive tenure and promotion.
  4. Reduce the number of comparisons you make. Every additional item, subgroup cut, and league-table column multiplies your false-positive count. Ask fewer, better questions.

Where Koji changes the equation

The multiple-comparisons trap is a symptom of a deeper design choice: building evaluation around scanning thousands of means for outliers. Koji for Education is built around understanding why a cohort responded the way it did, which sidesteps much of the multiplicity problem rather than patching it.

  • Evidence, not just a number to rank. Koji''s AI-moderated conversational interviews produce explanations — what specifically worked, what got in the way — so a low signal comes attached to a reason you can act on, instead of an unexplained dot below a benchmark line.
  • Quality scoring that flags weak data, not weak teachers. Koji scores the informativeness of responses, helping you down-weight low-effort or contradictory feedback before it pollutes an average — a far better use of "flagging" than hunting for outlier instructors.
  • Automatic thematic analysis turns hundreds of open-text comments into a handful of recurring themes, so insight comes from convergence across many students rather than from significance-testing one mean against another.
  • Programme- and institution-level reporting designed to read patterns across cohorts, where idiosyncratic noise washes out, instead of inviting managers to over-interpret a single instructor''s mean.
  • Formative, mid-cycle collection surfaces issues while a course is still running, so you respond to a developing situation rather than retrospectively flagging a finished one.

The same conversational AI engine powers the general-purpose koji.so platform that research and product teams use to make sense of open-ended feedback at scale — the education edition simply adds the question design, GDPR/AVG-aligned data handling, and quality-assurance reporting that universities need.

The takeaway

A red flag on a dashboard feels like information. Often it is arithmetic. Before you act on a "below-benchmark" instructor, ask the only question that matters: if every instructor here were equally good, how many would this process have flagged anyway? If the answer is "quite a few" — and with hundreds of simultaneous comparisons it always is — then a flag is a prompt to listen, not a verdict to deliver.

Want evaluation that explains rather than merely ranks? See how Koji for Education works.