Your 'Significant' Instructor Difference Might Point the Wrong Way: Type S and Type M Errors in Small-Cohort Evaluation
Statistical significance does not protect a small-cohort evaluation from pointing the wrong way or exaggerating the gap. Type S and Type M error analysis shows why, and what to report instead.
Koji Education Team
Product
In brief: In a small, noisy course-evaluation comparison, a statistically significant difference between two instructors or two cohorts can still have a meaningful probability of being in the wrong direction (a Type S, or sign, error) and is almost guaranteed to overstate the true gap (a Type M, or magnitude, error). Gelman and Carlin (2014) show that when statistical power is low, "significance" selects for exaggerated and occasionally sign-flipped estimates. For evaluation offices comparing tens of responses, the practical takeaway is to stop treating a significant p-value as proof of a real, correctly-signed difference and to report effect sizes with honest uncertainty instead.
Course-evaluation data is a textbook low-power setting: a seminar with 18 respondents, a two-tenths-of-a-point difference on a 1–5 scale, and a lot of measurement noise. Committees routinely act on comparisons like these — ranking instructors, flagging a "declining" course, or crediting a redesign. This article explains why a significant result in that regime is far weaker evidence than it looks, using the framework of design analysis and Type S / Type M errors.
What the research says
The foundational reference is Andrew Gelman and John Carlin's "Beyond Power Calculations: Assessing Type S (Sign) and Type M (Magnitude) Errors" (Perspectives on Psychological Science, 2014, 9(6), 641–651). They introduce two quantities that a conventional power analysis ignores:
- Type S (sign) error — the probability that a statistically significant estimate has the opposite sign to the true effect. If Instructor A is genuinely rated slightly higher than Instructor B, a Type S error is concluding, "significantly," that B is higher.
- Type M (magnitude) error, also called the exaggeration ratio — the factor by which the absolute size of a statistically significant estimate overstates the true effect. An exaggeration ratio of 3 means significant estimates are, on average, three times too big.
Their central result is uncomfortable: these errors are driven not by sample size alone but by the ratio of the true effect to the standard error. When a study is underpowered — small samples, noisy measures, small true effects — the only estimates that clear the significance threshold are the ones that happened to land far from zero. Conditioning on significance therefore selects for exaggeration, and when power falls low enough, for sign reversal. Gelman and Carlin demonstrate cases where a study with, say, 6% power to detect a plausible effect yields significant estimates that are exaggerated roughly ninefold and carry a non-trivial (double-digit percentage) chance of the wrong sign.
This is not an isolated argument. Button, Ioannidis, Mokrysz and colleagues (2013), in "Power failure: why small sample size undermines the reliability of neuroscience" (Nature Reviews Neuroscience, 14, 365–376), show empirically that chronically low-powered fields produce literatures dominated by inflated effect sizes and poor replicability — the "winner's curse." John Ioannidis (2005), "Why most published research findings are false" (PLoS Medicine, 2(8), e124), makes the same point from the direction of positive predictive value: when power is low and many comparisons are run, a large share of "significant" findings are false positives. And in the specific context of teaching evaluation, the meta-analysis by Uttl, White and Gonzalez (2017) — "Meta-analysis of faculty's teaching effectiveness" (Studies in Educational Evaluation, 54, 22–42) — found that once small-sample studies were properly weighted, the apparent relationship between student ratings and learning largely dissolved, exactly the pattern you expect when earlier significant results were small-study artefacts.
The mechanics of a design analysis are simple. You supply (1) a plausible true effect size drawn from prior literature or theory, not from your own noisy data, and (2) the standard error implied by your sample size and measurement variability. From these you compute the power, the Type S probability, and the exaggeration ratio. Crucially, the "plausible effect" must come from outside the current dataset; using your observed estimate defeats the purpose, because that estimate is the very thing suspected of being inflated.
Why it matters for course evaluation in practice
Course-evaluation comparisons live squarely in the low-power regime that manufactures Type S and Type M errors:
- Cohorts are small. Electives, capstones, and postgraduate seminars often return 10–30 usable responses. The standard error of a mean on a 1–5 scale with a standard deviation near 1.0 is roughly 0.2–0.3 at those sizes — the same order of magnitude as the differences committees try to interpret.
- True differences between competent instructors are small. Most teaching in a mature programme is broadly effective. Real gaps are typically a fraction of a scale point, precisely the range where power is weak.
- Many comparisons are run. An institution comparing every instructor against a benchmark, or every course year-on-year, performs hundreds of implicit tests, so the handful that reach "significance" are disproportionately the noisy extremes.
The consequences are concrete. A significant "drop" in a course's rating that triggers a review may, under design analysis, carry a 15–25% chance of pointing the wrong way and overstate the true change several-fold. Ranking instructors by whose scores are "significantly" above the mean rewards small, noisy cohorts that happened to land high — a mechanism that also interacts with the ceiling effects and skew endemic to rating scales. Personnel decisions built on such comparisons are not merely uncertain; they are systematically biased toward exaggeration, which is worse than random error because it is confidently wrong.
Design analysis reframes the reporting task. Instead of "Course X fell significantly (p = 0.03)," the honest statement is "Course X's mean fell by 0.2 points; given our sample this estimate is compatible with anything from a trivial dip to a moderate decline, and comparable small studies typically overstate such changes." That is less satisfying and far more defensible.
Limitations and honest caveats
Type S / Type M analysis is a tool, not a verdict, and it has real limitations a critical reader should weigh:
- It requires a prior on the true effect. The output depends on the plausible effect size you assume. Reasonable analysts can disagree about that number, and a poorly chosen prior yields misleading error rates. The remedy is to report results across a range of plausible effects, not a single guess.
- It assumes a broadly correct model. The framework typically assumes approximately unbiased estimation and roughly normal sampling error. Course-evaluation data is ordinal and ceiling-bound, so the normal approximation is imperfect; the qualitative conclusion (significance selects for exaggeration in low power) survives, but exact numbers should be treated as indicative.
- It does not fix non-response or construct bias. Design analysis addresses sampling noise, not whether the ratings measure teaching quality at all. A confidently-signed, correctly-sized estimate of a biased quantity is still biased. These issues are handled elsewhere in this knowledge base.
- Low power is not a reason to ignore data. The correct response is not to discard small cohorts but to widen the uncertainty you attach to them, pool information across cohorts where appropriate, and avoid threshold-based decisions. Under-powered evidence is still evidence when reported honestly.
None of this makes small-cohort evaluation worthless. It makes the dichotomous reading of it — significant versus not — untrustworthy, and pushes the analysis toward estimation with uncertainty.
How Koji incorporates this
Koji for Education is designed to mitigate the conditions that create Type S and Type M errors, though it cannot repeal the mathematics of small samples.
- Estimation over thresholds. Koji's analysis-and-reporting layer foregrounds the effect size and an uncertainty interval rather than a bare significance verdict, so a 0.2-point difference on 18 responses is presented with its wide interval attached. This aligns with the design-analysis recommendation to report magnitude and uncertainty, and it complements our guidance on bootstrap confidence intervals for small samples.
- Richer signal per respondent, lowering the standard error. The single biggest lever on Type S / Type M error is reducing noise. Koji's AI-moderated conversational interviews probe beyond a Likert number with structured follow-ups (open_ended, scale, single_choice, ranking), extracting more reliable information from each of the few respondents a small cohort provides — which shrinks the effective measurement variance rather than requiring more students.
- Automatic thematic analysis instead of significance-hunting. Rather than ranking instructors on marginally significant score gaps, Koji surfaces what students say drove an experience via thematic coding of open text, giving deans a mechanism-level read that does not hinge on an underpowered comparison.
- Pooling and triangulation across cohorts. Koji is built to aggregate across sections, terms, and cohorts, so a single noisy seminar is interpreted alongside related data rather than in isolation — the practical antidote to conditioning on one small, significant extreme.
- Bias-aware, non-dichotomous reporting. Reports are framed to discourage act-on-the-asterisk decisions, in the same spirit as our warning-label guidance for interpreting scores.
Koji is careful not to overclaim: the platform is designed to mitigate exaggeration by lowering measurement noise and reporting uncertainty honestly; it does not eliminate the fundamental limits of small samples. For teams that also run product or customer research, Koji's core research platform at koji.so applies the same AI-moderated interview engine and estimation-first reporting to those studies.
Frequently asked questions
Isn't a statistically significant result by definition unlikely to be due to chance? A small p-value controls the false-positive rate if the null is true, but it says nothing about the sign or size of a real effect. In low-power settings, the significant estimates are precisely the ones inflated by noise, so significance and exaggeration go together rather than being opposites.
What is the difference between Type S/M errors and ordinary Type I/II errors? Type I and Type II errors are about the decision to reject or retain a null hypothesis. Type S and Type M errors are about the estimate itself — whether a significant estimate has the right sign and a believable magnitude. They are complementary, not substitutes.
How much data do I need to avoid these errors in course evaluation? There is no universal threshold, because it depends on the true effect and measurement noise, not sample size alone. The practical move is to run a design analysis with a plausible effect size before interpreting a comparison, and to widen your uncertainty when cohorts are small. See our companion piece on how many responses make an evaluation reliable.
Does this mean I should ignore small-class evaluations? No. It means you should report them as estimates with honest uncertainty, avoid threshold-triggered decisions, and pool information across cohorts. Discarding the data is an over-correction; over-interpreting a lone asterisk is the actual error.
Can effect-size reporting replace significance testing here? Largely, yes. Reporting the estimated difference, an uncertainty interval, and a plainly-worded interpretation communicates far more than a significant/not-significant label — and is harder to weaponise in a personnel case. Our note on common-language effect sizes shows one accessible format.
Who introduced Type S and Type M errors? Andrew Gelman and John Carlin formalised the terms and the design-analysis procedure in their 2014 Perspectives on Psychological Science paper, building on a long line of work on the "winner's curse" and low power in the empirical sciences.
Related resources
- How Many Responses Do You Need for a Reliable Course Evaluation?
- Bootstrap Confidence Intervals for Small-Sample Course Evaluation
- Ceiling Effects, Skew, and What They Do to Course-Evaluation Statistics
- Regression to the Mean in Year-over-Year Evaluation Changes
- Common-Language Effect Sizes for Reporting Course Evaluation
- Value-Added Evidence vs Student Ratings (Carrell & West)
References
- Gelman, A., & Carlin, J. (2014). Beyond power calculations: Assessing Type S (sign) and Type M (magnitude) errors. Perspectives on Psychological Science, 9(6), 641–651. https://doi.org/10.1177/1745691614551642
- Button, K. S., Ioannidis, J. P. A., Mokrysz, C., Nosek, B. A., Flint, J., Robinson, E. S. J., & Munafò, M. R. (2013). Power failure: Why small sample size undermines the reliability of neuroscience. Nature Reviews Neuroscience, 14, 365–376. https://doi.org/10.1038/nrn3475
- Ioannidis, J. P. A. (2005). Why most published research findings are false. PLoS Medicine, 2(8), e124. https://doi.org/10.1371/journal.pmed.0020124
- Gelman, A., & Stern, H. (2006). The difference between "significant" and "not significant" is not itself statistically significant. The American Statistician, 60(4), 328–331. https://doi.org/10.1198/000313006X152649
- Uttl, B., White, C. A., & Gonzalez, D. W. (2017). Meta-analysis of faculty's teaching effectiveness: Student evaluation of teaching ratings and student learning are not related. Studies in Educational Evaluation, 54, 22–42. https://doi.org/10.1016/j.stueduc.2016.08.007
Related articles
How Many Responses Do You Need for a Reliable Course Evaluation?
Reliability of a class-average rating rises steeply with the number of students who respond. Generalizability theory — from Marsh''s syntheses to Dzakadzie (2026) — shows raters matter more than items, and a class of 12 cannot be evaluated like a class of 120.
Why a Low-Scoring Course Usually 'Improves' Next Year Even If Nothing Changed: Regression to the Mean in Course Evaluations
Year-over-year movements in course-evaluation scores are mostly statistical noise regressing toward an instructor's true average. Here is why regression to the mean fools quality-assurance committees, and how to read evaluation trends responsibly.
When Almost Everyone Scores 4.5: Ceiling Effects, Skew, and What They Do to Course-Evaluation Statistics
Course-evaluation ratings pile up at the top of the scale, producing a strong ceiling effect and negative skew that breaks the statistics most universities still report. Here is what the evidence shows and how to report ratings honestly.
Fair Confidence Intervals for Small Classes: The Bootstrap for Course-Evaluation Reporting
Small classes and skewed rating distributions break the textbook confidence interval. The bootstrap resamples the data you actually have to produce honest uncertainty bounds. Here is the method, its limits, and how to report it.