New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Methodology7 min read

Is a 0.3-Point Difference Real? Effect Size, Measurement Error, and the Over-Interpretation of Course Evaluations

"She scored 4.1, he scored 3.8—she's the stronger teacher." That sentence, repeated in thousands of promotion meetings, is usually a statistical error. Here is why small mean differences in course evaluations are mostly noise, and how to tell signal from it.

Koji Education Team

Product ·

Bottom line: A difference of a few tenths of a point between two instructors' average ratings is, in the overwhelming majority of cases, indistinguishable from measurement error. The research is unambiguous that faculty and administrators routinely treat these trivial gaps as meaningful—even when explicitly warned not to. Before a 0.2- or 0.3-point difference shapes a hiring, renewal, or promotion decision, it must clear three hurdles most institutions never check: statistical significance, practical (effect-size) significance, and the reliability of the instrument itself. Course evaluations that report only a mean fail all three by design.

The number on the report is not a fact—it is an estimate

Every average course-evaluation score is a sample statistic with uncertainty around it. The "4.1" is a point estimate; the truth lives somewhere inside a confidence interval whose width depends on how many students responded and how spread out their answers were. For a class of fifteen with a handful of responses, that interval can be wide enough to swallow most of the scale. (We unpack this in What Is a Fair Score for a Class of 12?.)

This is not a pedantic statistical footnote. It is the difference between a defensible decision and an indefensible one. When two instructors differ by 0.3 and their confidence intervals overlap heavily, the honest conclusion is we cannot distinguish them. Yet evaluation reports almost never display intervals. They display ranked means, and ranked means invite comparison the underlying data cannot support.

The evidence: people over-interpret small differences—even when warned

This is not speculation about how committees might misread data; it has been tested directly. Guy Boysen and colleagues studied how faculty and administrators interpret evaluation means, and the findings are consistent and uncomfortable. In Uses and Misuses of Student Evaluations of Teaching (Boysen, 2015) and related work on the over-interpretation of small mean differences, participants reliably judged one instructor as a better teacher than another on the basis of differences as small as a few tenths of a point—differences that statistical tests and confidence intervals indicated were not meaningful.

The most damning detail: in Boysen's work, the over-interpretation persisted even when participants were explicitly warned to avoid reading into small differences and were shown that the gaps were not statistically significant. The warning helped a little; it did not solve the problem. The pull to rank a 4.1 above a 3.8 is cognitive, not rational, and it does not switch off because a footnote tells it to.

Statistical significance is necessary but not sufficient

Suppose the difference is statistically significant—say, because the classes were large. You are still not done. Statistical significance tells you a difference is unlikely to be pure chance; it says nothing about whether the difference is large enough to matter. This is the distinction between statistical and practical significance, and it is where a great deal of evaluation analysis quietly goes wrong.

With a few hundred respondents, a difference of 0.1 on a 5-point scale can be "significant" at p < 0.05 and simultaneously trivial—an effect size so small that no reasonable observer would change a decision over it. Effect-size measures (such as Cohen's d) exist precisely to keep significance honest: they ask not "is there a difference?" but "how big, in standardised terms, is it?" A statistically significant but tiny effect is a real difference that does not justify a real consequence.

The reliability ceiling

There is a third hurdle, often forgotten. No measurement is perfectly reliable, and the reliability of an evaluation instrument caps how finely you can legitimately discriminate between scores. If a course-evaluation scale has a reliability coefficient of, say, 0.80, a meaningful chunk of the observed variance between two close scores is measurement error, not true difference. (For the difference between a consistent instrument and a valid one, see Are Course Evaluations Valid?.) The practical implication is stark: differences smaller than the instrument's error band are not interpretable, full stop. Reporting them to two decimal places creates a false precision that the measurement cannot honour.

A worked example

Imagine two lecturers in the same department. Lecturer A averages 4.1 across 40 respondents; Lecturer B averages 3.8 across 18 respondents. The report ranks A above B, and a reader concludes A is the better teacher. Now add what the report omits. With these sample sizes and typical response variance, both means carry confidence intervals roughly half a point wide—intervals that overlap substantially. The 0.3 gap is well within the range you would expect from sampling noise alone. Factor in that B's smaller sample makes its estimate shakier still, and that neither figure has been corrected for class size, discipline, or the regression-to-the-mean effect, and the "A is better" conclusion has no statistical footing whatsoever. The two are, on this evidence, indistinguishable. The only honest report is one that says so—ideally by drawing the intervals so the overlap is impossible to miss, rather than printing two tidy decimals that invite a ranking the data cannot support.

Why this matters most at the top of the stakes ladder

The over-interpretation of small differences is harmless when nobody acts on it. It becomes a fairness problem the moment these numbers drive tenure and promotion decisions, contract renewals, or merit pay. A 0.2-point gap that is statistically, practically, and psychometrically meaningless can still end a career if a committee treats the higher number as proof of superior teaching. Add the documented biases in raw ratings—and the regression-to-the-mean artefact whereby last year's low scorer "improves" for purely statistical reasons—and the case for treating small mean differences as decisive collapses entirely.

"But we have to make decisions—are you saying the numbers are useless?"

No. This is the strongest and fairest objection, and the answer is not nihilism. Course evaluations carry real signal; the problem is the resolution at which institutions read them. The defensible posture is not "ignore the data" but "respect its grain." Treat evaluation scores as detecting broad bands—a clearly struggling course versus a clearly thriving one—rather than fine rankings. A 2.6 and a 4.4 are telling you something. A 4.1 and a 3.8 almost never are. And no single number, however large the gap, should be the sole evidence for a high-stakes judgement; it should be triangulated with peer review, teaching artefacts, and qualitative feedback.

A second objection: "Then just report confidence intervals and effect sizes." Yes—and you should. But intervals on a number that mostly measures satisfaction or surprise rather than learning still answer a shallow question precisely. Better statistics on a thin construct is progress; it is not the destination.

How Koji changes what you are measuring, not just how you report it

Koji does not pretend that better arithmetic rescues a thin instrument. Its contribution is to widen the evidence base so that decisions rest on more than a contestable decimal:

  • Conversational, AI-moderated interviews replace a lone Likert mean with structured qualitative depth—why a course worked or did not—so committees are not forced to adjudicate 0.3-point gaps in the absence of anything richer.
  • Automatic thematic analysis turns open-text responses into weighted, recurring themes, giving evaluators a second, qualitatively-grounded signal to set against the numbers rather than ranking on the numbers alone.
  • Quality scoring and standardised moderation mean every respondent is asked equivalently well-formed questions, reducing the noise that inflates measurement error in the first place.
  • Programme- and institution-level reporting is designed to surface bands and patterns over time, not to invite spurious instructor-to-instructor rankings on trivial differences.
  • Mixing structured question types—scale items alongside open-ended probing—lets you keep the comparability of a number while refusing to let the number stand alone.

The same engine powers koji.so for product and customer research, where confusing a statistically significant but trivial metric movement for a real one wastes roadmaps the same way it wastes careers in academia. The discipline is identical: ask whether the difference is real, whether it is large, and whether the instrument can even see it.

The takeaway

Before any course-evaluation difference influences a decision, ask three questions in order. Is it statistically distinguishable from chance? Is it large enough, as an effect size, to matter? And is it bigger than the instrument's own error? Most small mean differences fail at least one of these tests. Reading them as fact is not rigour—it is the appearance of rigour. The genuinely evidence-led move is to measure deeper and rank less.

Want evaluation evidence that is built to be triangulated, not over-read? Explore Koji for Education.