Is There Really No Difference, or Just No Evidence? Bayes Factors for Course-Evaluation Comparisons
A non-significant p-value cannot confirm two instructors scored the same — it only fails to reject. Bayes factors quantify evidence FOR the null as well as against it, distinguishing absence of evidence from evidence of absence in course-evaluation comparisons.
Koji Education Team
Product
In brief: When a course-evaluation comparison returns a non-significant p-value — two instructors, two cohorts, before-and-after a teaching change — you have not shown they are the same. A p-value can only fail to reject the null; it cannot support it. A Bayes factor can. It is the ratio of how well the data are predicted by the "no difference" hypothesis versus the "there is a difference" hypothesis, and it moves continuously in both directions: a Bayes factor of 8 means the data are eight times more likely under a real difference, while a Bayes factor of 1/8 means they are eight times more likely under no difference — genuine evidence of absence. Values near 1 mean the data are simply insensitive: you have not collected enough to conclude anything. For QA offices that constantly ask "did this actually change?", Bayes factors answer a question ordinary significance tests cannot.
The problem with "not significant"
Every quality-assurance team eventually writes a sentence like: "There was no significant difference between the new syllabus and the old one (p = 0.21), so the redesign had no effect on satisfaction." That sentence is a logical error, and a sophisticated reader will spot it immediately. A non-significant result is consistent with two very different worlds: (1) there really is no difference, or (2) there is a difference but your sample was too small, your data too noisy, or your test too weak to detect it. The p-value cannot tell these apart. As the old slogan goes, absence of evidence is not evidence of absence — and the ordinary t-test only ever delivers the former.
This is not a fringe concern. Higher-education decisions ride on "no difference" claims constantly: keeping an instructor because their scores are "no worse" than a colleague's, retiring a pedagogical intervention because it "made no difference," declaring two delivery modes "equivalent." Each of these is an argument for the null, and the standard toolkit is built only to argue against it.
There are two principled ways out. One is equivalence testing (TOST), which asks whether a difference is small enough to be practically negligible against a pre-set bound. The other, discussed here, is the Bayes factor, which asks how much the data shift the balance of evidence between the two hypotheses.
What the research says
The core idea. Rouder, Speckman, Sun, Morey, and Iverson (2009), in a paper that has become the standard reference, introduced a default Bayesian t-test whose Bayes factor lets a researcher accept as well as reject the null. The Bayes factor is the ratio of the marginal likelihood of the data under the two models. Crucially — and unlike the p-value — it is symmetric: it can favour the null, favour the alternative, or sit near 1 to indicate the data cannot discriminate. Rouder and colleagues also solved a practical problem: they specified a sensible default prior (a Cauchy distribution on the effect size) so that analysts without strong prior beliefs can still compute a well-behaved Bayes factor.
Why p-values needed replacing. Wagenmakers (2007) laid out the case against p-values that motivates the Bayesian alternative: p-values overstate evidence against the null, depend on hypothetical data never observed, and depend on the researcher's unstated sampling intentions (when they planned to stop collecting data). His BIC-based approximation showed that a Bayes-factor-style answer can be computed even from ordinary regression output — lowering the barrier to adoption.
Interpretation and tooling. Wagenmakers and colleagues (2018) made the machinery accessible through the free software JASP, and provide the now-standard interpretive bands: a Bayes factor between 1 and 3 is "anecdotal" (barely worth a mention), 3 to 10 is "moderate," and above 10 is "strong" evidence — reading in whichever direction the ratio points. Dienes (2014) made the applied case that Bayes factors are the right tool precisely when a study returns a null result, because they distinguish "the data support no effect" from "the data are inconclusive" — a distinction that determines whether you have learned anything at all. Dienes emphasises a practical rule: a Bayes factor between roughly 1/3 and 3 means your data are insensitive, and the honest report is "we cannot yet tell," not "there is no effect."
Why it matters for course evaluation in practice
Bayes factors change what you are allowed to conclude, and they map cleanly onto the recurring decisions a QA office faces.
-
Confirming equivalence, not just failing to find a difference. When a department redesigns a course and satisfaction is statistically unchanged, a Bayes factor of, say, 1/6 lets you say "the data provide moderate evidence that the redesign did not change satisfaction" — a defensible, positive statement. A Bayes factor near 1 tells you instead to withhold judgement and collect another cohort.
-
Small cohorts, honest answers. European programmes are full of small classes where p-values are chronically underpowered. The Bayes factor is candid about this: with 15 respondents it will usually land in the insensitive zone, correctly signalling that no strong conclusion — in either direction — is warranted. This complements the warnings in Type S and Type M errors and the 4.2-vs-4.4 trap.
-
Sequential monitoring without penalty. Because the Bayes factor does not depend on a fixed stopping rule the way a p-value does, you can watch the evidence accumulate as responses arrive and stop when it is decisive — useful for mid-cycle monitoring where waiting for a full cohort is costly.
-
Communicating uncertainty to non-statisticians. "The evidence is six-to-one that nothing changed" is far more intelligible to a teaching committee than "we failed to reject the null at alpha = 0.05," and it does not invite the fallacy of reading non-significance as proof of no effect.
Limitations and honest caveats
Bayes factors are not a free lunch, and a critical reader will press on the following.
-
The answer depends on the prior. The Bayes factor compares the data against a specified alternative — encoded in the prior on the effect size. Choose a very wide prior (huge effects are plausible) and you bias the comparison toward the null; choose a narrow one and you bias it toward the alternative. Rouder's default Cauchy prior is reasonable, but "reasonable" is a judgement. Best practice is a robustness check: report how the Bayes factor moves as the prior width changes, so readers see the conclusion is not an artefact of one arbitrary choice.
-
It is a relative measure. A Bayes factor tells you which of two specified models the data prefer — not whether either model is any good. If both the null and your chosen alternative are poor descriptions of the data, the ratio is still computable and still misleading.
-
Ordinal data still needs care. Course-evaluation responses are ordinal, bounded, and often skewed. A default Bayesian t-test assumes normality; for 1–5 Likert items the same cautions apply as for any averaging of ordinal scales, and an ordinal or non-parametric Bayesian model is preferable when the scale is coarse or the distribution lumpy.
-
It does not license fishing. Bayesian methods reduce but do not remove the danger of running many comparisons and reporting the decisive ones. Pre-registration of the specific comparisons that matter is still the honest practice.
-
"Evidence for the null" is evidence for a region, not a point. A Bayes factor favouring the null supports "the effect is small relative to your prior," not "the effect is exactly zero." That is usually what you want in evaluation — but say it precisely.
How Koji incorporates this
Koji is built so that "did this change?" is answered honestly rather than reflexively.
-
Reporting that distinguishes the three outcomes. Koji's analytics layer is designed to frame comparisons in the language of evidence — a real difference, evidence of no meaningful difference, or insufficient data to tell — rather than a bare significant/not-significant verdict. This directly discourages the "non-significant therefore equivalent" fallacy that Bayes factors exist to prevent, and it pairs naturally with the platform's equivalence-testing and shrinkage approaches to fair comparison.
-
Honesty about small samples. Because so many course cohorts are small, Koji is designed to flag when a comparison is simply under-powered — the "insensitive data" zone — instead of presenting a reassuring-looking null. That is the same candour a Bayes factor near 1 provides.
-
Richer evidence than a single number. A Bayes factor summarises the numeric signal; it does not explain why two courses feel the same or different to students. Koji's AI-moderated conversational interviews probe beyond the Likert rating with adaptive follow-ups, and its automatic thematic analysis of the open text supplies the qualitative texture that a "no meaningful difference" finding needs to be trusted — did students really experience the redesign as equivalent, or did they simply not notice it?
Koji does not claim to turn every evaluation into a Bayesian analysis automatically, and it does not pretend a Bayes factor removes the need for judgement about priors and practical significance. What it is designed to do is stop teams from mistaking silence for agreement. (Koji's core research platform at koji.so applies the same evidence-first reporting to product and customer research, where "no difference between segments" claims are just as consequential.)
Related Resources
- Equivalence Testing (TOST): How to Show Two Instructors Really Do Score the Same
- The 4.2 vs 4.4 Trap: Why Small Differences in Means Are Usually Noise
- Type S and Type M Errors in Small-Cohort Evaluation
- Permutation Tests for Course-Evaluation Comparisons
- Empirical-Bayes Shrinkage for Course-Evaluation Scores
- Ordinal Regression for Course-Evaluation Data
References
- Rouder, J. N., Speckman, P. L., Sun, D., Morey, R. D., & Iverson, G. (2009). Bayesian t tests for accepting and rejecting the null hypothesis. Psychonomic Bulletin & Review, 16(2), 225–237. https://doi.org/10.3758/PBR.16.2.225
- Wagenmakers, E.-J. (2007). A practical solution to the pervasive problems of p values. Psychonomic Bulletin & Review, 14(5), 779–804. https://doi.org/10.3758/BF03194105
- Wagenmakers, E.-J., Love, J., Marsman, M., et al. (2018). Bayesian inference for psychology. Part II: Example applications with JASP. Psychonomic Bulletin & Review, 25(1), 58–76. https://doi.org/10.3758/s13423-017-1323-7
- Dienes, Z. (2014). Using Bayes to get the most out of non-significant results. Frontiers in Psychology, 5, 781. https://doi.org/10.3389/fpsyg.2014.00781
Related articles
The 4.2 vs 4.4 Trap: Why Small Differences in Course-Evaluation Means Are Usually Noise
Faculty and administrators routinely read meaning into tiny gaps between course-evaluation means. Boysen (2015) and Boysen et al. (2014) show this happens even when confidence intervals say the difference is nothing — and that warnings barely help. How to report uncertainty honestly.
Stop Comparing Raw Averages: Empirical-Bayes Shrinkage for Course-Evaluation Scores
Why raw course-evaluation means mislead — especially for small classes — and how empirical-Bayes shrinkage (partial pooling), the method Kane and Staiger applied to teacher value-added, produces fairer, more accurate instructor estimates.
Is That Score Gap Real or Just Luck of the Draw? Permutation Tests for Course-Evaluation Comparisons
Comparing two instructors' evaluation averages with a t-test quietly assumes normal, equal-variance data you rarely have with small, skewed Likert samples. Permutation tests, formalised by Fisher and reviewed by Ernst, answer the comparison question by shuffling the data itself — with almost no distributional assumptions. Here is when and how to use them.
Equivalence Testing (TOST): How to Show Two Instructors Really Do Score the Same
A non-significant t-test never proves two course-evaluation means are equivalent. Equivalence testing (TOST) does — here is what the method is, how to set a smallest effect size of interest, and how to use it for defensible decisions.