Where Should You Set the Review Trigger? Signal Detection Theory for Course-Evaluation Thresholds
Deciding which courses to flag for review is a signal-detection problem. What Swets, Dawes & Monahan and the ROC literature say about separating how well a score discriminates from where you set the cut — and how to choose a threshold that reflects the real cost of misses versus false alarms.
Koji Education Team
Product
In short: Choosing a cut-off for flagging a course "in need of review" is a signal-detection decision, and signal detection theory separates two things institutions routinely confuse: how well an evaluation score can tell a genuinely struggling course from a fine one (its discriminability), and where you put the threshold (a policy choice that trades misses against false alarms). Swets, Dawes and Monahan (2000) show that the ROC curve summarises the first, while the best cut-off for the second depends on the base rate of real problems and the relative cost of a missed problem versus a false alarm. A single fixed number like "flag anything below 3.5/5" quietly buries both decisions and is almost never optimal.
The problem: a threshold is a decision, not a measurement
Most quality-assurance systems reduce a rich evaluation to a single binary act: this course is flagged for review, or it is not. Somewhere a line is drawn — below 3.5 out of 5, or bottom decile, or "two standard deviations below the mean" — and courses fall on one side or the other. That line feels like a technical detail. It is in fact the most consequential decision in the whole process, because it fixes the balance between two kinds of error, and it usually gets made without anyone naming the trade-off.
There are four possible outcomes each time you apply a threshold. A genuinely struggling course that you flag is a hit. A struggling course you miss is a miss (a false negative). A fine course you wrongly flag is a false alarm (a false positive). A fine course you correctly leave alone is a correct rejection. Every threshold buys more of one error by accepting more of the other: lower the bar and you catch more real problems but haul in more innocent courses; raise it and you spare the innocent but let real problems slide. Signal detection theory (SDT) is the framework built precisely to reason about this trade-off, and it has been doing so in radiology, weather forecasting, and forensic decision-making for decades.
What the research says
SDT originated in psychophysics and engineering — Green and Swets (1966), Signal Detection Theory and Psychophysics, is the canonical text — and its core move is to split a detection decision into two mathematically independent parts:
- Discriminability (d′) — how far apart the "signal" distribution (struggling courses) and the "noise" distribution (fine courses) are on your evidence scale. This is a property of the evidence: a more valid, more reliable evaluation instrument separates the two groups better and has a higher d′. No choice of threshold can improve it; only a better measure can.
- Criterion (the threshold) — where you decide to draw the line. This is a policy choice, entirely separate from d′. You can slide the criterion anywhere along the evidence scale, changing the mix of hits, misses, and false alarms, without changing how well the evidence itself discriminates.
The Receiver Operating Characteristic (ROC) curve plots the hit rate against the false-alarm rate as you slide the criterion from lenient to strict. Its shape captures discriminability independent of threshold, and the area under the ROC curve (AUC) is a single, threshold-free summary of how well the score distinguishes struggling from fine courses — 0.5 is useless (a coin flip), 1.0 is perfect.
The article that carried this into applied decision-making is Swets, Dawes and Monahan (2000), "Psychological science can improve diagnostic decisions," which occupied the entire inaugural issue of Psychological Science in the Public Interest. Their central, practical arguments transfer directly to course evaluation. First, improving the evidence and choosing the threshold are different projects — buying a better instrument raises d′; arguing about where to flag only moves the criterion. Second, the optimal criterion depends on two things outside the data: the base rate of genuine problems (if only 5% of courses are truly struggling, even a good test produces mostly false alarms at a lenient cut) and the relative cost of a miss versus a false alarm. Third, mechanical, explicitly-set thresholds tend to outperform holistic clinical judgement — a finding with a long pedigree (Dawes, Faust and Meehl, 1989) that argues for a transparent rule, but the right transparent rule.
Two supporting tools are worth naming. Youden''s J (Youden, 1950), defined as sensitivity + specificity − 1, is a common way to pick a "balanced" operating point on an ROC curve — but note that it implicitly assumes misses and false alarms are equally costly, which they rarely are in QA. Macmillan and Creelman (2005), Detection Theory: A User''s Guide, is the standard modern reference for computing d′ and criterion measures and for the pitfalls (extreme rates, small samples) that bite when you estimate them from thin data.
Why it matters for course evaluation in practice
Reframing the review trigger as a signal-detection decision changes several habits:
- Stop treating a low d′ instrument as if a better threshold could save it. If your evaluation barely separates good courses from bad (AUC near 0.6), no cut-off will flag well — you will trade misses for false alarms along a nearly diagonal ROC. The fix is better evidence (richer questions, triangulation), not a cleverer line.
- Make the base rate explicit. Genuinely failing courses are usually rare. At a 5% base rate, a test with 80% sensitivity and 80% specificity flags a set of courses of which the majority are false alarms — a direct consequence of the same base-rate arithmetic covered in Koji''s guidance on base-rate neglect. A threshold chosen without the base rate in view will drown reviewers in noise.
- Name the cost asymmetry, then set the criterion to match it. Is a missed struggling course (students harmed, problem festers) worse than a false alarm (an instructor unfairly scrutinised, staff time wasted)? Usually the two costs are not equal, and the criterion should lean toward the cheaper error. That is a governance judgement — but SDT forces you to make it rather than let a round number make it for you.
- Report the operating point, not just the flag. "This threshold has ~85% sensitivity and ~90% specificity at our base rate" is an accountable statement a committee can debate. "We flag below 3.5" hides every assumption.
Limitations and honest caveats
- "Struggling" is not a clean binary. SDT models two underlying distributions — signal and noise. Real course quality is continuous and multidimensional, so the neat two-distribution picture is an idealisation. It is a useful lens for the decision, not a literal model of teaching.
- You often lack a gold standard. Estimating d′ or an ROC requires knowing which courses were truly struggling — but if evaluations were a perfect ground truth you would not need the threshold. In practice you approximate ground truth with follow-up review outcomes, peer observation, or learning data, all imperfect. This makes AUC estimates uncertain, especially with few known cases.
- Estimates are unstable in small samples. d′ and criterion measures behave badly at extreme hit/false-alarm rates and with small n — exactly the regime of a small department''s annual review. Macmillan and Creelman''s corrections help but cannot manufacture information that is not there.
- Costs and base rates are contestable. The "optimal" criterion is only optimal given a base rate and a cost ratio that reasonable people will dispute. SDT organises the argument; it does not settle it.
- A good operating point cannot fix a biased score. If the underlying ratings are contaminated by response-rate or demographic confounds, tuning the threshold just relocates unfairness. Discriminability assumes the evidence is measuring the right thing.
How Koji incorporates this
Koji for Education is designed to raise the quality of the evidence (d′) and to make the threshold decision explicit rather than hidden inside a dashboard rule:
- Richer evidence to raise discriminability. A single Likert average is a low-d′ signal. Koji''s AI-moderated conversational interviews probe beyond the number, its automatic thematic analysis surfaces why a course is struggling, and its quality scoring flags low-information responses — all of which sharpen the separation between courses that genuinely need help and those that merely scored a little low by chance. Better separation is the only thing that improves every possible operating point at once.
- Threshold framing that exposes the trade-off. Rather than presenting a bare cut-off, Koji''s reporting is built to show where a course sits relative to expected variation — echoing its funnel-plot and empirical-Bayes shrinkage guidance — so a review committee is choosing an operating point on an informed curve, not applying a number nobody can justify.
- Base-rate and uncertainty awareness. Because most courses are fine, Koji''s bias-aware reporting is designed to keep the base rate and the resulting false-alarm burden in view, consistent with the base-rate-neglect and misclassification guidance, so a lenient threshold does not silently flood reviewers with correct-rejection candidates.
- Mid-cycle collection to change the cost of a miss. SDT''s cost asymmetry is not fixed by nature — you can lower the cost of a miss by catching problems earlier. Koji''s formative, mid-cycle evaluation lets an institution act before a struggling course reaches the end of term, which shifts the whole calculus toward less punitive thresholds.
Koji is designed to help institutions make the threshold decision well, not to hand them a universally correct cut-off — SDT is explicit that no such number exists independent of base rates and costs. The same discriminability-versus-criterion logic governs the AI-moderated studies on Koji''s core research platform at koji.so, where deciding which signals warrant follow-up is the same kind of decision.
Frequently asked questions
What does signal detection theory add beyond "pick a cut-off"? It separates two decisions that a bare cut-off fuses: how well the evidence discriminates struggling from fine courses (discriminability, d′, summarised by the ROC/AUC) and where you draw the line (the criterion). Recognising they are independent stops institutions from arguing about the threshold when the real problem is a weak instrument, and vice versa.
What is an ROC curve in a course-evaluation context? It plots the hit rate (struggling courses correctly flagged) against the false-alarm rate (fine courses wrongly flagged) as you slide the review threshold from lenient to strict. The area under it (AUC) is a single, threshold-free measure of how well your evaluation score distinguishes courses that truly need review from those that do not.
How should the base rate of struggling courses affect my threshold? Strongly. When genuine problems are rare — say 5% of courses — even an accurate test produces mostly false alarms at a lenient cut, because there are so many more fine courses to misclassify. Ignoring the base rate leads to thresholds that overwhelm reviewers with false positives, the same arithmetic behind base-rate neglect.
Is there a statistically correct place to set the threshold? No. The optimal criterion depends on the base rate of real problems and the relative cost of a miss versus a false alarm — both of which are governance judgements, not computations. Tools like Youden''s J offer a "balanced" point, but only under the assumption that both errors cost the same, which is rarely true in quality assurance.
How do I improve flagging without just moving the line? Raise discriminability by improving the evidence: add probing open-text and conversational follow-up, triangulate with peer observation or learning data, and screen out careless responses. A higher-d′ instrument improves the hit rate at every false-alarm rate — the whole ROC curve lifts — which no repositioning of the threshold can achieve.
Can I compute d′ from a small department''s data? Cautiously. d′ and criterion estimates are unstable at extreme rates and small samples, exactly the situation in a small unit''s annual cycle. You also need an approximate ground truth (which courses truly struggled) to estimate an ROC at all. Treat the numbers as rough and lean on shrinkage and multi-year aggregation.
Related resources
- One Vivid Comment Is Not a Pattern: Base-Rate Neglect in Reading Course Evaluations
- Which Instructor Scores Are Genuinely Unusual? Funnel Plots for Fair Course-Evaluation Comparison
- Statistical Process Control for Course Evaluations: Using Control Charts to Separate Signal from Noise
- Even Fair Evaluations Can Rank Unfairly: Why Course-Evaluation Scores Misclassify Instructors
- Stop Comparing Raw Averages: Empirical-Bayes Shrinkage for Course-Evaluation Scores
- The 4.2 vs 4.4 Trap: Why Small Differences in Course-Evaluation Means Are Usually Noise
References
- Swets, J. A., Dawes, R. M., & Monahan, J. (2000). Psychological science can improve diagnostic decisions. Psychological Science in the Public Interest, 1(1), 1–26. https://doi.org/10.1111/1529-1006.001
- Green, D. M., & Swets, J. A. (1966). Signal Detection Theory and Psychophysics. New York: Wiley.
- Macmillan, N. A., & Creelman, C. D. (2005). Detection Theory: A User''s Guide (2nd ed.). Mahwah, NJ: Lawrence Erlbaum Associates.
- Youden, W. J. (1950). Index for rating diagnostic tests. Cancer, 3(1), 32–35. https://doi.org/10.1002/1097-0142(1950)3:1<32::AID-CNCR2820030106>3.0.CO;2-3
- Dawes, R. M., Faust, D., & Meehl, P. E. (1989). Clinical versus actuarial judgment. Science, 243(4899), 1668–1674. https://doi.org/10.1126/science.2648573
Related articles
The 4.2 vs 4.4 Trap: Why Small Differences in Course-Evaluation Means Are Usually Noise
Faculty and administrators routinely read meaning into tiny gaps between course-evaluation means. Boysen (2015) and Boysen et al. (2014) show this happens even when confidence intervals say the difference is nothing — and that warnings barely help. How to report uncertainty honestly.
Even Fair Evaluations Can Rank Unfairly: Why Course-Evaluation Scores Misclassify Instructors
A simulation-based finding that even unbiased, reliable, and valid course-evaluation scores produce high misclassification rates when used to rank or compare instructors — and what it means for tenure, promotion, and merit decisions.
Stop Comparing Raw Averages: Empirical-Bayes Shrinkage for Course-Evaluation Scores
Why raw course-evaluation means mislead — especially for small classes — and how empirical-Bayes shrinkage (partial pooling), the method Kane and Staiger applied to teacher value-added, produces fairer, more accurate instructor estimates.
Statistical Process Control for Course Evaluations: Using Control Charts to Separate Signal from Noise
How Shewhart control charts and statistical process control distinguish real changes in teaching quality from ordinary random variation in course evaluation scores - so QA teams stop reacting to noise. Grounded in Sivena and Nikolaidis (2019, 2022).