How Big a Difference Can You Actually Detect? A-Priori Power Analysis for Course Evaluation
Reliability tells you how stable a score is; statistical power tells you whether your sample can detect a real difference at all. A practical guide to a-priori power analysis, minimum detectable effects, and why most single-class course-evaluation comparisons are underpowered before they begin.
Koji Education Team
Product
In brief: Before you compare two instructors, two cohorts, or a course before and after a redesign, there is a question you should answer first: given how many responses you will realistically get, how big a difference could you even detect? That is statistical power — the probability your test flags a real effect of a given size — and computing it in advance is a-priori power analysis. It is a different question from reliability. Reliability (how stable a score is) can be excellent while power (the ability to detect a difference between scores) is hopeless. Most single-class comparisons in higher education are badly underpowered before a single student responds: with 20 students per group you can only reliably detect a difference of roughly three-quarters of a standard deviation. Knowing your minimum detectable effect turns "no significant difference" from an embarrassing surprise into a planned, honest limit.
Power is not reliability, and not "how many responses for a stable average"
A common confusion sinks a lot of evaluation analysis. The corpus already covers how many responses you need for a reliable course evaluation — that question is about the precision of a single course's mean, driven by the response rate and the homogeneity of opinion. Power is a different animal. Power asks: if instructor A is genuinely half a point better than instructor B, what is the probability my comparison will come back statistically significant rather than lost in noise?
You can have one without the other. A course mean of 4.3 estimated from 40 respondents can be highly reliable — a stable, trustworthy summary of that class — and yet a comparison of two such classes can have only a 30% chance of detecting a real, meaningful gap between them. Reliability is about the quality of one number; power is about the sensitivity of a comparison between numbers. Confusing the two is how institutions end up making high-stakes personnel and curriculum decisions on comparisons that never had a chance of working.
What the research says
The framework. Statistical power is governed by four interlocking quantities, and fixing any three determines the fourth: the sample size, the significance level (conventionally 0.05), the true effect size, and power itself (conventionally set to 0.80, an 80% chance of detecting a real effect). Jacob Cohen (1988), in the foundational text on the subject, formalised this and, in his widely cited A Power Primer (1992), provided the conventional effect-size benchmarks — for a difference between two means, a Cohen's d of 0.2 is "small," 0.5 "medium," and 0.8 "large" — together with the sample sizes needed to detect each at 80% power. His headline numbers are sobering: to detect a medium difference between two independent groups at 80% power you need about 64 respondents per group; for a small difference you need roughly 393 per group.
A-priori vs post-hoc. Faul, Erdfelder, Lang, and Buchner (2007), introducing the free G*Power 3 software, distinguish several kinds of power analysis. The one that matters is a-priori: you specify the smallest effect you care about and the power you want, and it returns the required sample size — before you collect data. (Its mirror image, the sensitivity analysis, fixes your realistic sample size and returns the minimum detectable effect — the smallest difference you have a fair chance of catching.) Faul and colleagues are explicit that post-hoc "observed power" — computed after a non-significant result from the effect you happened to observe — is nearly useless, because it is just a restatement of the p-value and tells you nothing new.
Why underpowering is corrosive. Button and colleagues (2013), in an influential analysis of an entire research field, showed that chronically low power does more than miss real effects: it makes the significant findings that do emerge less likely to be true and more likely to exaggerate the effect size. This is the same mechanism described for course evaluation in Type S and Type M errors: in an underpowered comparison, a result only crosses the significance threshold when noise happens to inflate it, so the "significant" instructor difference you act on is systematically too large and sometimes points the wrong way.
Why it matters for course evaluation in practice
A-priori power analysis reshapes how a QA office plans and reports comparisons.
-
Set expectations before fielding. If you know in advance that a planned comparison of two 25-student sections can only detect a difference of about 0.8 standard deviations, you know before you start that anything subtler is invisible. You can then either pool across sections and semesters to raise the sample, or report the comparison as descriptive rather than inferential.
-
Report the minimum detectable effect alongside every null. The single most credibility-enhancing sentence a QA report can contain is: "With the responses received, the smallest difference this analysis could reliably detect was 0.6 points; the observed difference of 0.15 is well within the range we would expect from noise." That reframes a null result as an informed limit, not a failed test — and it pairs directly with the 4.2-vs-4.4 trap.
-
Decide what is worth measuring. Power analysis forces the prior question: what size of difference actually matters? A 0.1-point gap on a 5-point scale is rarely educationally meaningful even if detectable. Specifying a minimum effect of interest keeps evaluation focused on differences that would change a decision.
-
Justify pooling and multilevel designs. Because power rises with sample size, the analysis provides the quantitative case for aggregating across cohorts — which is exactly the structure that multilevel models and longitudinal designs like interrupted time series exploit.
Limitations and honest caveats
Power analysis is a planning tool, not an oracle, and it comes with real caveats.
-
You must guess the effect size. A-priori power depends on the true effect you are trying to detect — which you do not know. The honest practice is to power for the smallest effect that would matter (a policy choice) rather than an optimistic guess, and to treat the result as a floor, not a promise.
-
The conventional benchmarks are arbitrary. Cohen's small/medium/large labels and the 0.80 power target are useful conventions, not laws of nature. In high-stakes contexts (tenure, programme closure) you may want 90% or 95% power; for exploratory mid-cycle feedback, less.
-
Simple formulas assume simple designs. The textbook two-group calculation assumes independent observations, equal variances, and roughly normal data. Course-evaluation responses are ordinal, clustered (students nested in courses), and often skewed, all of which change the true power. Simulation-based power analysis, or power calculations built for multilevel and ordinal models, is preferable when the design departs from the textbook case.
-
Power is about detection, not importance. A well-powered study can detect a difference so small it is meaningless. Power analysis should always be paired with a judgement about practical significance and, where the goal is to confirm similarity, with equivalence testing or Bayes factors.
-
Post-hoc "observed power" is not a fix. After a non-significant result, computing power from the observed effect adds nothing; report the minimum detectable effect and confidence interval instead.
How Koji incorporates this
Koji is built to make the sample-and-detectability question visible up front rather than discovered too late.
-
Response-rate visibility and realistic planning. Because Koji tracks invitations, completions, and response rates in real time, a QA officer can see the sample they are actually going to get and reason about what it can detect — instead of running a comparison and being surprised by a null. This turns power from an afterthought into a planning input.
-
Reporting that states the detectable limit. Koji's analytics layer is designed to present differences with their uncertainty — confidence intervals and noise-aware framing — so that a small observed gap is shown as consistent with no real difference rather than reported as a finding. This is the practical expression of "report the minimum detectable effect," and it connects to the platform's broader fair-comparison tools such as empirical-Bayes shrinkage.
-
Pooling by design. Because responses are stored against stable question and course objects across semesters, Koji makes it straightforward to aggregate cohorts to reach the sample sizes power analysis says you need — the single most effective remedy for the chronic under-powering of individual small classes.
-
Depth where breadth is impossible. When a cohort is simply too small to power a quantitative comparison, Koji's AI-moderated conversational interviews extract far more signal per respondent than a Likert item does, and its automatic thematic analysis surfaces patterns from a handful of rich responses — a qualitative route to insight when the numbers can never reach significance.
Koji does not claim to make small classes statistically powerful, and it does not run a formal a-priori calculation for you automatically. What it is designed to do is stop teams from mistaking an underpowered null for a real finding. (Koji's core research platform at koji.so applies the same evidence-first, sample-aware reporting to product and customer research, where underpowered A/B and segment comparisons are just as common.)
Related Resources
- How Many Responses Do You Need for a Reliable Course Evaluation?
- Type S and Type M Errors in Small-Cohort Evaluation
- The 4.2 vs 4.4 Trap: Why Small Differences in Means Are Usually Noise
- Bayes Factors for Course-Evaluation Comparisons
- Equivalence Testing (TOST) for Course Evaluation
- Multilevel Models for Nested Course-Evaluation Data
References
- Cohen, J. (1992). A power primer. Psychological Bulletin, 112(1), 155–159. https://doi.org/10.1037/0033-2909.112.1.155
- Cohen, J. (1988). Statistical Power Analysis for the Behavioral Sciences (2nd ed.). Lawrence Erlbaum.
- Faul, F., Erdfelder, E., Lang, A.-G., & Buchner, A. (2007). G*Power 3: A flexible statistical power analysis program for the social, behavioral, and biomedical sciences. Behavior Research Methods, 39(2), 175–191. https://doi.org/10.3758/BF03193146
- Button, K. S., Ioannidis, J. P. A., Mokrysz, C., Nosek, B. A., Flint, J., Robinson, E. S. J., & Munafò, M. R. (2013). Power failure: why small sample size undermines the reliability of neuroscience. Nature Reviews Neuroscience, 14(5), 365–376. https://doi.org/10.1038/nrn3475
Related articles
How Many Responses Do You Need for a Reliable Course Evaluation?
Reliability of a class-average rating rises steeply with the number of students who respond. Generalizability theory — from Marsh''s syntheses to Dzakadzie (2026) — shows raters matter more than items, and a class of 12 cannot be evaluated like a class of 120.
The 4.2 vs 4.4 Trap: Why Small Differences in Course-Evaluation Means Are Usually Noise
Faculty and administrators routinely read meaning into tiny gaps between course-evaluation means. Boysen (2015) and Boysen et al. (2014) show this happens even when confidence intervals say the difference is nothing — and that warnings barely help. How to report uncertainty honestly.
Students Are Nested in Courses: Why Multilevel Models Beat Raw Averages for Course Evaluation
Course-evaluation data has a nested structure — students within sections within instructors — and a flat average ignores it. This guide explains how multilevel (hierarchical linear) models partition variance, why ignoring clustering understates uncertainty, and what it means for fair reporting.
Equivalence Testing (TOST): How to Show Two Instructors Really Do Score the Same
A non-significant t-test never proves two course-evaluation means are equivalent. Equivalence testing (TOST) does — here is what the method is, how to set a smallest effect size of interest, and how to use it for defensible decisions.