Equivalence Testing (TOST): How to Show Two Instructors Really Do Score the Same
A non-significant t-test never proves two course-evaluation means are equivalent. Equivalence testing (TOST) does — here is what the method is, how to set a smallest effect size of interest, and how to use it for defensible decisions.
Koji Education Team
Product
In brief
Suppose two lecturers teaching parallel sections score 4.2 and 4.3 on the overall-rating item, and a t-test returns p = 0.28. The usual conclusion — "no significant difference, so they are equivalent" — is a logical error. A non-significant test means you failed to find a difference, which is not the same as finding no difference; the study may simply have lacked power. To make the positive claim that two instructors, sections, or delivery modes score practically the same, you need equivalence testing, most commonly the two one-sided tests (TOST) procedure (Schuirmann, 1987; Lakens, 2017). TOST asks a sharper, more honest question: is the observed difference small enough to fall inside a pre-specified band of "differences too small to matter"?
Answer box. Equivalence testing formally tests whether an effect is smaller than a smallest effect size of interest (SESOI). The two one-sided tests (TOST) procedure runs two one-sided t-tests against the lower and upper equivalence bounds; if both are significant, you can conclude the difference is statistically smaller than the bound — i.e., practically equivalent. In course evaluation this lets you make defensible positive claims (two instructors, sections, or modes score the same) that a non-significant conventional t-test cannot support, provided you set the equivalence bound before looking at the data.
What the research says
The procedure originates in Schuirmann's (1987) paper in the Journal of Pharmacokinetics and Biopharmaceutics, which compared the "power approach" (test the null of no difference, require adequate power) against the two one-sided tests procedure for assessing whether two drug formulations were bioequivalent. Schuirmann showed that TOST — testing whether the true difference is greater than a lower bound and less than an upper bound — corresponds to checking whether the (1 − 2α) confidence interval for the difference falls entirely inside the equivalence range, and has superior properties for the equivalence question. Bioequivalence regulation has used this logic ever since.
Daniel Lakens brought the method into mainstream behavioural science. Lakens (2017), "Equivalence tests: A practical primer for t tests, correlations, and meta-analyses" (Social Psychological and Personality Science), reframes TOST for psychologists and stresses the central design decision: you must specify a smallest effect size of interest (SESOI) — the smallest difference you would care about — before analysis. The equivalence bounds are set at ±SESOI. You then run two one-sided tests: one against the lower bound, one against the upper. If both reject, the difference is significantly smaller than the SESOI and you can declare statistical equivalence. Lakens, Scheel and Isager (2018), "Equivalence testing for psychological research: A tutorial" (Advances in Methods and Practices in Psychological Science), provide a step-by-step treatment, including how to justify a SESOI (from theory, from prior effect sizes, or from a smallest effect that is practically consequential) and how to interpret the four possible outcomes when a null-hypothesis test and an equivalence test are combined.
That combination is the conceptual payoff. Running both a conventional test and an equivalence test yields four states: (1) significant and equivalent — a difference too small to matter but detectable; (2) significant and not equivalent — a real, meaningful difference; (3) not significant and equivalent — good evidence of practical no-difference; and (4) not significant and not equivalent — the genuinely uninformative case where the data cannot distinguish anything, which conventional practice silently mislabels as "no difference."
Why it matters for course evaluation in practice
Course evaluation is saturated with implicit equivalence claims that conventional statistics cannot license. Making them explicit changes the quality of decisions.
-
Defensible personnel and comparison claims. When a committee says "the two candidates' teaching scores are basically the same" or "moving the course online didn't hurt evaluations," they are asserting equivalence. A non-significant difference does not support that assertion — with the small samples typical of single sections, almost nothing reaches significance, so "no significant difference" is nearly guaranteed and nearly meaningless. TOST provides a positive test with a pre-registered bound, which is far more defensible if a decision is ever challenged, and it dovetails with the cautions about misclassification and ranking discussed elsewhere in this knowledge base.
-
It forces the "how much matters?" conversation. Setting a SESOI compels a programme to decide, in advance, how many points on the evaluation scale constitute a difference worth acting on — say, 0.3 on a 5-point scale. That single act of judgement does more for interpretive discipline than any p-value, because it converts a vague "is there a difference?" into "is the difference big enough to matter?" This is the same shift that confidence-interval reporting encourages, made into a formal test.
-
It protects against over-reading noise. Year-over-year, section-to-section, and mode-to-mode wobble in evaluation means is mostly noise. Equivalence testing gives a principled way to say "this 0.1 change is within the band we agreed is trivial," pre-empting the regression-to-the-mean and small-difference over-interpretation that drives unnecessary interventions.
-
It improves modality and reform evaluations. Institutions frequently ask whether a redesign, a new delivery mode, or a larger class size harmed the student experience. Equivalence (or the related non-inferiority) framing — "the new format is no worse than the old by more than X" — is exactly the right question, and TOST answers it directly rather than resting on a failure to detect harm.
Limitations and honest caveats
-
The SESOI is a judgement, and the result depends on it. Equivalence bounds are not read off the data; they encode a value decision. A generous bound makes equivalence easy to declare and a strict one makes it hard. This is a feature — it makes the value judgement explicit — but it means results must always be reported alongside the bound and its justification, and pre-specification is essential to avoid choosing a bound that flatters a desired conclusion.
-
Equivalence testing needs adequate power too. Just as an underpowered study fails to detect real differences, an underpowered equivalence test fails to detect equivalence, landing in the uninformative "not significant, not equivalent" cell. Small single-section samples may be unable to establish equivalence within a tight bound; the honest report in that case is "inconclusive," not "equivalent."
-
Statistical equivalence is not identity. Concluding that two means fall within a trivial band does not mean the two courses were the same in any deeper sense — students, content, timing, and cohort composition differ. Equivalence on a scale score is a narrow claim about that score, not a global judgement of teaching quality, and should not be over-generalised.
-
Assumptions still apply. TOST for means rests on the usual t-test assumptions, and course-evaluation data are ordinal, often skewed, and clustered within sections. The ordinal-vs-interval debate and the case for multilevel or robust variants apply here as much as to any mean-based analysis, and a purist may prefer equivalence tests built on more appropriate models.
How Koji incorporates this
Koji for Education is designed to support the "is this difference big enough to matter?" question rather than only flagging statistical significance, so equivalence claims are made honestly.
-
Comparison reporting with pre-set meaningful-difference bands. Koji's analysis and reporting layer lets a quality team define, in advance, a smallest meaningful difference on the evaluation scale and reports section-to-section, mode-to-mode, and period-to-period comparisons against that band — surfacing whether a gap is inside the "trivial" range rather than presenting a bare, easily-misread p-value or mean gap. This operationalises the SESOI discipline that Lakens (2017) argues is the crux of a credible equivalence claim.
-
Confidence intervals and effect sizes over binary significance. Because TOST is equivalent to checking whether the confidence interval for a difference sits inside the equivalence bounds, Koji emphasises interval and effect-size reporting, giving decision-makers the raw material to judge equivalence and its uncertainty instead of a significant/not-significant verdict.
-
Honest "inconclusive" reporting. Where samples are too small to establish either a difference or equivalence, Koji is designed to report the comparison as underpowered and inconclusive rather than defaulting to "no difference," directly countering the logical error that motivates equivalence testing in the first place.
-
AI-moderated depth to complement the number. When two sections score equivalently, Koji's AI-moderated conversational interviews and thematic analysis can reveal whether the experience behind the equal scores differed — one section praised for pace, another for support — so an equivalence finding on a scalar rating is not mistaken for identity of teaching.
Koji frames these as designed to support rigorous equivalence reasoning, not to automate a decision. Koji's core research platform at koji.so applies the same interval-first, effect-size-aware reporting to product and customer research, where "these two versions performed the same" is an equivalence claim that a non-significant A/B test cannot, by itself, justify.
Frequently asked questions
Why doesn't a non-significant t-test prove two instructors score the same?
A non-significant result means you failed to detect a difference, not that no difference exists — the study may simply have lacked power. With the small samples typical of single sections almost nothing reaches significance, so "no significant difference" is nearly guaranteed and cannot support the positive claim of equivalence. Equivalence testing is needed to make that claim.
What is the two one-sided tests (TOST) procedure?
TOST (Schuirmann, 1987) runs two one-sided t-tests: one testing whether the difference is greater than a lower equivalence bound and one testing whether it is less than an upper bound. If both are significant, the difference is statistically smaller than the bound, so you can conclude practical equivalence. It is equivalent to checking whether the confidence interval for the difference lies entirely inside the equivalence range.
What is a smallest effect size of interest (SESOI)?
The SESOI is the smallest difference you would care about — for example 0.3 points on a 5-point evaluation scale. Equivalence bounds are set at plus and minus the SESOI. Lakens (2017) stresses it must be chosen before analysis, justified from theory, prior effect sizes, or practical consequence, and reported alongside the result because the conclusion depends on it.
Can an equivalence test be inconclusive?
Yes. Just as an underpowered study fails to detect real differences, an underpowered equivalence test fails to detect equivalence, producing a "not significant and not equivalent" result. With small single-section samples and a tight bound, the honest conclusion is "inconclusive," not "equivalent." Equivalence testing needs adequate power just like any other test.
Does statistical equivalence mean two courses are identical?
No. Concluding two means fall within a trivial band is a narrow claim about that one score; students, content, timing and cohort still differ. Equivalence on a scale score should not be over-generalised into a judgement that the courses or teaching were the same in any deeper sense.
How does Koji support equivalence reasoning?
Koji lets teams pre-set a smallest meaningful difference and reports comparisons against that band, emphasises confidence intervals and effect sizes over binary significance, reports underpowered comparisons as inconclusive rather than "no difference," and uses AI-moderated interviews to reveal whether the experience behind two equal scores actually differed.
Related resources
- Small mean differences and confidence intervals in course evaluation
- Is ranking instructors by evaluation scores fair?
- Misclassification when ranking instructors for personnel decisions
- Funnel plots for fair instructor comparison
- Interpreting and reporting student ratings responsibly
- Empirical Bayes shrinkage for instructor scores
References
- Schuirmann, D. J. (1987). A comparison of the two one-sided tests procedure and the power approach for assessing the equivalence of average bioavailability. Journal of Pharmacokinetics and Biopharmaceutics, 15(6), 657–680. https://doi.org/10.1007/BF01068419
- Lakens, D. (2017). Equivalence tests: A practical primer for t tests, correlations, and meta-analyses. Social Psychological and Personality Science, 8(4), 355–362. https://doi.org/10.1177/1948550617697177
- Lakens, D., Scheel, A. M., & Isager, P. M. (2018). Equivalence testing for psychological research: A tutorial. Advances in Methods and Practices in Psychological Science, 1(2), 259–269. https://doi.org/10.1177/2515245918770963
Related articles
Can You Fairly Rank Instructors by Their Course-Evaluation Scores? What Esarey & Valdes (2020) Proved
Even if course evaluations were unbiased, reliable, and valid, ranking instructors by their scores would still misclassify many good teachers. A walk through the Esarey & Valdes (2020) simulation and what it means for how you report and use evaluation data.
Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees
The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.
The 4.2 vs 4.4 Trap: Why Small Differences in Course-Evaluation Means Are Usually Noise
Faculty and administrators routinely read meaning into tiny gaps between course-evaluation means. Boysen (2015) and Boysen et al. (2014) show this happens even when confidence intervals say the difference is nothing — and that warnings barely help. How to report uncertainty honestly.
Even Fair Evaluations Can Rank Unfairly: Why Course-Evaluation Scores Misclassify Instructors
A simulation-based finding that even unbiased, reliable, and valid course-evaluation scores produce high misclassification rates when used to rank or compare instructors — and what it means for tenure, promotion, and merit decisions.