New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods9 min read

Is 4.2 Really Worse Than 4.4? Bayesian Estimation and Credible Intervals for Course Evaluations

Bayesian estimation answers the question institutions actually ask — how probable is it that this instructor is below standard? — with credible intervals and a region of practical equivalence instead of raw averages or p-values.

Koji Education Team

Product

Answer first. A course with a mean of 4.2 is not reliably "worse" than one at 4.4 unless the difference is large relative to the uncertainty in both numbers — and with small classes it rarely is. Bayesian estimation answers the question institutions actually care about ("how probable is it that this instructor is below standard, and by how much?") with a credible interval — a range that genuinely has, say, a 95% probability of containing the true value — and with a region of practical equivalence that lets you declare two scores "the same for practical purposes." It is a more honest framework than comparing raw averages or chasing p-values, especially when sample sizes are small.

What the research says

Frequentist confidence intervals and null-hypothesis tests answer a question few evaluation stakeholders are asking. A 95% confidence interval does not mean "there is a 95% probability the true score is in this range" — it is a statement about the long-run behaviour of the procedure. A p-value is not the probability that two instructors are equal. These are among the most persistently misinterpreted objects in applied statistics, and the misinterpretation is precisely the intuitive Bayesian reading.

Kruschke (2013), in "Bayesian estimation supersedes the t-test" (Journal of Experimental Psychology: General), set out an accessible Bayesian alternative — the BEST framework — that estimates the full posterior distribution of a group difference and reports a highest-density credible interval together with a region of practical equivalence (ROPE): a band around zero (or around a benchmark) inside which any difference is deemed too small to matter. If the credible interval falls entirely outside the ROPE, the difference is both real and material; if it falls entirely inside, the two are practically equivalent; if it straddles, the data are inconclusive. This maps almost perfectly onto the decisions a quality office makes about instructor scores.

Kruschke and Liddell (2018), "The Bayesian New Statistics" (Psychonomic Bulletin & Review), extended the argument to estimation, meta-analysis, and power, and made the case that Bayesian estimation delivers what practitioners think confidence intervals give them — direct probability statements about parameters — without the interpretive contortions.

The approach has been applied to teaching evaluation directly. Fouskakis, Petrakos, and Vavouras (2014) built a Bayesian hierarchical model (a beta regression with a Dirichlet prior on the coefficients) to compare teaching-quality indicators using student survey data from Panteion University in Athens, demonstrating that the Bayesian machinery handles the bounded, skewed nature of rating data and yields direct probability interpretations of which indicators students weight most. More broadly, the hierarchical Bayesian framework of Gelman and Hill (2007) is the natural home for evaluation data because it formalises partial pooling: an instructor with only a handful of ratings is estimated by borrowing strength from the wider distribution, so extreme small-sample means are pulled toward the average by exactly the amount the data justify.

Why it matters for course evaluation in practice

The "4.2 vs 4.4" problem is endemic. Institutions rank, flag, and sometimes reward on differences that are well inside the noise. A Bayesian credible interval makes the noise legible: with 15 respondents, the interval around a 4.2 may run from 3.7 to 4.6, comfortably overlapping the 4.4. The correct conclusion is "indistinguishable," and the framework says so plainly.

ROPE turns a philosophical debate into a policy lever. Instead of arguing about significance, a QA committee can decide in advance what difference is educationally meaningful — perhaps 0.3 on a 5-point scale — and encode it as a ROPE. This forces the useful conversation ("what size of difference would actually change a decision?") and pre-empts the ranking of instructors on trivial gaps.

Priors are a feature for small classes. A weakly-informative prior — for example, that most instructors score in the upper-middle of the scale, which institutional history supports — stabilises estimates for tiny cohorts without overwhelming genuine signal. This is the Bayesian counterpart to shrinkage, and it directly addresses the volatility that makes small-class averages untrustworthy.

It communicates uncertainty to non-statisticians. "There is an 80% probability this course is above the faculty median" is a sentence a dean can act on. "We reject the null hypothesis at p < 0.05" is not. The interpretability is not a cosmetic advantage; it changes what decisions get made.

This complements two techniques already in the toolkit. Empirical-Bayes shrinkage produces stabilised point estimates using the same borrowing-strength logic; full Bayesian estimation additionally delivers the whole posterior and the ROPE-based equivalence decision. Frequentist equivalence testing (TOST) reaches a similar "no meaningful difference" verdict from the other philosophical direction. Choosing among them is a matter of context and audience, not of one being universally right.

Limitations and honest caveats

  • Priors are a genuine assumption, and they can be contested. A prior that is too strong can dominate a small sample and manufacture a conclusion. Good practice uses weakly-informative priors, reports them explicitly, and runs a prior sensitivity analysis showing the conclusion does not hinge on the prior choice. A Bayesian analysis that hides its prior deserves the scepticism it will attract.
  • Bayesian machinery does not rescue biased data. If ratings are contaminated by gender bias, grading leniency, or non-response bias, a credible interval will be a beautifully-quantified estimate of a biased quantity. The framework improves inference, not validity. Garbage in, credible-interval-quantified garbage out.
  • Computation and competence cost. Full Bayesian models (MCMC) require software, tuning, and convergence checks that many QA offices are not staffed for. For routine reporting, simpler tools — bootstrap intervals, shrinkage — may deliver most of the benefit at lower operational risk.
  • The ROPE width is a value judgement, not a statistic. Setting the region of practical equivalence requires deciding what difference matters educationally. That is a legitimate policy decision, but it should be made transparently and consistently, not reverse-engineered to produce a desired verdict.
  • Bayesian estimation is not a bias-detection method. It quantifies uncertainty about a parameter; it does not tell you whether the parameter means what you hope. Triangulation with peer observation, learning outcomes, and qualitative evidence remains essential.

How Koji incorporates this

Koji's reporting philosophy is aligned with the estimation-not-testing stance of the Bayesian New Statistics: report the plausible range and the practical significance, not a bare average or a significance star.

  • Uncertainty-first reporting. Koji is designed to present evaluation results with explicit uncertainty rather than as decontextualised point estimates, so that a 4.2 from 12 students is visibly less certain than a 4.2 from 200. This directly implements the credible-interval intuition — small samples produce wide ranges, and the interface shows it.
  • Practical-significance framing. Rather than encouraging league tables built on hundredths of a point, Koji's analytics are built to foreground whether a difference is material, echoing the ROPE logic: below a meaningful threshold, two scores are reported as effectively equivalent.
  • Stabilised small-class estimates. Koji's approach to small-cohort reporting borrows strength across comparable courses so that a single seminar's mean is not read in isolation — the partial-pooling idea that underlies both hierarchical Bayes and shrinkage.
  • Triangulation by design. Because Bayesian estimation improves inference but not construct validity, Koji pairs quantitative scores with AI-moderated qualitative probing and structured multi-source evidence, so decisions rest on more than one number however well its uncertainty is quantified. These mechanisms are framed as designed to support honest interpretation, not to certify any single score as correct.

Research teams outside teaching face the same "is this difference real?" problem across product cohorts; Koji's core platform at koji.so applies the same uncertainty-aware reporting to customer and product studies.

Related resources

Frequently asked questions

What does a Bayesian credible interval actually mean?

A 95% credible interval is a range that, given your model and prior, has a 95% probability of containing the true value. This is the intuitive interpretation people wrongly attach to frequentist confidence intervals. It lets you make direct probability statements — "there is an 80% chance this course is above the median" — that stakeholders can act on.

How is this different from just comparing the averages?

Comparing raw averages ignores uncertainty. A 4.2 and a 4.4 from small classes usually have overlapping credible intervals, meaning the data cannot distinguish them. Bayesian estimation surfaces that overlap and, via a region of practical equivalence, lets you declare the two "the same for practical purposes" rather than ranking on noise.

What is a region of practical equivalence (ROPE)?

A ROPE is a band around a reference value inside which any difference is judged too small to matter educationally — say, ±0.3 on a 5-point scale. If a credible interval for a difference lies entirely inside the ROPE, the scores are practically equivalent; if entirely outside, the difference is real and material; if it straddles, the evidence is inconclusive. It converts a statistical question into a policy decision made in advance.

Do the priors just let analysts get the answer they want?

They can be abused, which is why priors must be stated explicitly and tested. Responsible practice uses weakly-informative priors grounded in institutional history and reports a sensitivity analysis showing the conclusion is not driven by the prior. A hidden or strong, unjustified prior is a red flag, not a normal part of the method.

Does Bayesian analysis fix biased evaluation scores?

No. It improves how honestly you quantify uncertainty about a number; it does nothing about whether the number is contaminated by gender bias, grading leniency, or non-response. A credible interval around a biased mean is still centred on a biased value. Validity requires triangulation with other evidence, not a better interval.

References

  • Kruschke, J. K. (2013). Bayesian estimation supersedes the t test. Journal of Experimental Psychology: General, 142(2), 573–603. https://doi.org/10.1037/a0029146
  • Kruschke, J. K., & Liddell, T. M. (2018). The Bayesian New Statistics: Hypothesis testing, estimation, meta-analysis, and power analysis from a Bayesian perspective. Psychonomic Bulletin & Review, 25(1), 178–206. https://doi.org/10.3758/s13423-016-1221-4
  • Fouskakis, D., Petrakos, G., & Vavouras, I. (2014). A Bayesian hierarchical model for comparative evaluation of teaching quality indicators in higher education. arXiv preprint arXiv:1404.1710. https://doi.org/10.48550/arXiv.1404.1710
  • Gelman, A., & Hill, J. (2007). Data Analysis Using Regression and Multilevel/Hierarchical Models. Cambridge University Press. https://doi.org/10.1017/CBO9780511790942