When You Flag a "Course of Concern", How Often Are You Wrong? The Base-Rate Trap in Evaluation Dashboards
A flagging rule tells you how often a genuinely poor course trips the threshold. It does not tell you the thing that matters: if a course is flagged, how likely is it actually poor? When bad teaching is rare, most flags are false alarms — the same maths that makes a positive rare-disease screen usually wrong.
Koji Education Team
Product · July 29, 2026
Bottom line up front: The rule "flag every course below 3.5" answers one question — how often does a genuinely poor course trip the threshold? — but committees act on a different one: if a course is flagged, how likely is it actually poor? Those two probabilities are not the same, and when genuinely poor teaching is rare, the gap between them is enormous. A dashboard that flags courses without accounting for the base rate will mostly generate false alarms, wasted review effort, and quiet injustice to the people flagged. This is the same base-rate mathematics that makes a positive result on a rare-disease screen far more likely to be wrong than right.
Two probabilities everyone confuses
Every flagging rule has a sensitivity — the chance it flags a course that really is poor — and a specificity — the chance it clears a course that really is fine. Those are properties of the rule, measured looking forward from the truth to the flag: P(flag | poor).
But a head of department reading a dashboard reasons backward: here is a flag, how worried should I be? That is P(poor | flag) — the positive predictive value. And the positive predictive value depends on something the rule itself does not contain: how common genuinely poor teaching is in the first place. Confusing P(flag | poor) with P(poor | flag) is the base-rate fallacy, one of the most robust findings in the judgment literature since Kahneman and Tversky, and it is wired into almost every evaluation dashboard in the sector.
Watch it happen with real numbers
The fix that decades of research keeps confirming is to stop reasoning in percentages and use natural frequencies instead — Gerd Gigerenzer's work shows the proportion of people who reason correctly roughly triples when the same facts are framed as counts of people (or here, courses) rather than conditional probabilities.
So take 1,000 courses. Suppose — generously to the dashboard — that the flag is a decent test: it correctly flags 80% of genuinely poor courses (sensitivity) and correctly clears 85% of adequate ones (specificity). And suppose that genuinely poor teaching is, as most quality data suggest, uncommon: 5% of courses.
- Of the 1,000 courses, 50 are genuinely poor. The rule flags 80% of them: 40 true flags.
- The other 950 are adequate. The rule wrongly flags 15% of them: about 143 false flags.
- Total flagged: roughly 183. True flags among them: 40.
The positive predictive value is 40 / 183 ≈ 22%. Roughly four out of every five flagged courses are, in fact, perfectly adequate. The dashboard has not identified eighty-odd problem courses; it has generated a to-do list that is mostly noise, wrapped in the authority of a red cell.
The classic medical parallel makes the point stick: Gigerenzer describes physicians told a screening test had a 0.3% prevalence, 50% sensitivity and a 3% false-positive rate, then asked the probability of disease given a positive result. Estimates ranged from 1% to 99%; the correct answer, by Bayes' rule, is about 5%. Experts get this wrong constantly. So will your teaching committee, unless the base rate is put in front of them.
This is not the multiple-comparisons problem
It is worth being precise, because the corpus already has a piece on why flagging enough instructors guarantees some look bad by chance. That is the multiple-comparisons trap: run enough tests against a null and some cross the line by luck. The base-rate trap is different and compounds it. Here the flag can be a genuinely informative test — sensitivity and specificity well above chance — and it still produces mostly false positives, purely because the condition it screens for is rare. Fix your multiple comparisons and this problem remains.
Why the base rate is low — and why that is good news misread as bad
Most courses at most universities are adequate to good. Genuinely poor teaching — the kind that warrants intervention — is the tail, not the body, of the distribution. That is a healthy state of affairs. But it is precisely what makes low predictive value inevitable: the rarer the thing you are screening for, the more the false positives from the large "fine" majority swamp the true positives from the small "poor" minority. A sector that has largely good teaching will, with a naive threshold, drown in false concern flags. The better the underlying teaching, the worse a fixed cut-off performs.
"Surely it is safer to over-flag than to miss a bad course?"
This is the strongest objection and it deserves a direct answer. The intuition is that false negatives (missing a genuinely poor course) are costlier than false positives (reviewing a fine one), so err toward flagging. But false positives are not free. Each one consumes scarce review capacity, erodes trust in the system, and — landing disproportionately on contingent and early-career staff whose contracts hang on these numbers — does real career harm. Worse, a screen with 22% predictive value trains reviewers to ignore flags: alarm fatigue is a well-documented failure mode of low-specificity warning systems, and an ignored flag misses the genuinely poor course anyway.
The resolution is not to abandon flagging but to stage it. Use a cheap, sensitive first screen to cast a wide net, then a specific, evidence-rich confirmatory review before anything is called a "concern". Two stages recover both goals; a single crude threshold sacrifices both.
What to actually do
- Estimate your base rate. How many courses per year genuinely warrant intervention? Even a rough figure lets you compute the predictive value of your threshold and stop over-interpreting flags.
- Report predictive value to committees, in natural frequencies. "Of the 183 flags this term, historically about 40 reflect a real problem" is a sentence that changes how a committee behaves.
- Treat a flag as triage, not a verdict. A flag should open an inquiry, never close one. Pair every quantitative flag with qualitative corroboration before any consequence attaches. This is also the case for standard-setting rather than arbitrary cut-offs.
- Raise the bar for consequences, not for curiosity. It is fine to look widely; it is not fine to sanction on a signal with one-in-five predictive value.
Recompute it for your own numbers
The illustrative 22% is not a universal constant; it is a function of three quantities you can estimate for your own institution. Push the base rate up — say a genuinely troubled department where one course in five warrants intervention — and predictive value climbs. Push it down toward the 2–3% typical of a healthy faculty and it collapses further still. The lesson is not a fixed number but a habit: before a committee treats a flag as a finding, someone should plug this term's sensitivity, specificity and base-rate estimate into the same natural-frequency table and state, out loud, roughly how many of the flags are expected to be real. That single sentence reframes the entire meeting from "here are the problem courses" to "here are the courses worth a closer look" — and it is the difference between a process that informs judgment and one that manufactures unwarranted certainty.
Where Koji fits
A flag is only as useful as what sits behind it. Koji is built so that a flag opens straight into evidence rather than a spreadsheet cell. Its AI-moderated conversational interviews probe why students rated a course as they did, and its automatic thematic analysis surfaces the specific, recurring issues — or their absence — behind a low number. That corroborating qualitative signal is exactly what raises predictive value: a course flagged low and carrying a consistent, substantiated theme is a genuine concern; a course flagged low with no coherent problem behind it is usually noise. Koji does not eliminate false positives — no instrument can — but it lets a committee separate the 40 from the 143 with evidence, instead of treating all 183 the same. The same interview engine powers koji.so for teams doing general user research, where acting on an unreliable signal is just as expensive.
Base rates are not a statistical nicety. They are the difference between a warning system your colleagues trust and a red-cell generator they have learned to scroll past.