New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Methodology9 min read

Is a 4.2 Really Better Than a 3.9? The Measurement Error Hiding in Your Course-Evaluation League Table

Institutions routinely rank instructors and courses on decimal-point differences in mean ratings. But a mean carries a margin of error, and most of the gaps that decide promotions and awards fall well inside it. The evidence that faculty and administrators over-read small differences is direct — and unflattering.

Koji Education Team

Product · August 23, 2026

Bottom line up front: A course-evaluation mean is not a fixed fact; it is an estimate with a margin of error. When one instructor scores 4.2 and another 3.9, the difference is very often smaller than the measurement and sampling error around each number — which means the ranking is noise dressed as signal. Yet institutions build league tables, hand out awards, and inform promotion cases on exactly these differences. The research on how faculty and administrators read evaluation means is direct, replicated, and uncomfortable: we over-interpret small gaps, even when told not to.

A mean is an estimate, not a verdict

Every observed score sits inside a band of uncertainty. Two sources widen that band.

The first is measurement error. Classical test theory expresses it as the standard error of measurement: SEM = SD × √(1 − reliability). The lower the reliability of the instrument and the more the responses vary, the wider the band around the observed mean. A rating scale that is only moderately reliable — which is typical for a handful of items answered by a modest class — produces means that could easily have landed a few tenths higher or lower by chance alone. (Reliability is a precondition here, not a nicety; see reliability versus validity and why Cronbach's alpha is so often misreported.)

The second is sampling error. The students who respond are a sample, and small samples give unstable means. A 25-student module with a 40% response rate is a mean built on ten voices; move two of them and the decimal moves too. This is the same fragility that makes low response rates so treacherous — except here the problem is not only who answered but how few.

Put the two together and the honest way to report a 3.9 is not "3.9" but "3.9, give or take" — with a confidence interval wide enough that it overlaps the 4.2 next to it. When intervals overlap, you cannot say one is higher.

The evidence: we read noise as skill

This is not a theoretical worry. Guy Boysen and colleagues tested how real faculty and administrators interpret evaluation means, and the results are consistent across studies.

In one study, participants judged teaching quality from evaluation means where the differences between teachers were smaller than 0.30 on a five-point scale — differences with no statistical or practical meaning. They still rated teachers with the higher means as more qualified and assigned them higher awards (Boysen, 2015, Teaching of Psychology). A companion paper found that faculty and administrators interpreted small differences as meaningful even when confidence intervals and significance tests explicitly indicated no real difference (Boysen et al., "Statistical knowledge and the over-interpretation of student evaluations of teaching," Assessment & Evaluation in Higher Education).

Worse, the fix is only partly effective. When Boysen tested whether warning people about overinterpretation helped, warnings eliminated the bias when participants compared two teachers — but did not fully stop department heads from reading meaning into small fluctuations within a single teacher's scores over time (Boysen, 2015, Scholarship of Teaching and Learning in Psychology). The instinct to explain a 0.2 dip with a story about the instructor is remarkably hard to switch off.

Why the league table is the wrong object

Ranking instructors by mean rating fails for reasons that compound:

  • The gaps are inside the error band. Most rank positions in a typical department separate means that are statistically indistinguishable. The order is largely random reshuffling of noise.
  • It invites forbidden comparisons. Ranking encourages comparing a 12-student seminar with a 300-student lecture, an elective with a required gateway course — settings whose means are not on the same footing. Reading a group-level number as if it spoke about the individual is a version of the ecological fallacy.
  • It multiplies false alarms. Scanning a dashboard of many courses for the "low" ones is a multiple-comparisons problem: with enough courses, some will look bad by chance, and the base rate of genuine problems is low enough that most flags are false.

What to do instead

None of this means evaluation data is worthless — it means it must be read at the right resolution.

  1. Report uncertainty, not just the point. Publish confidence intervals or, at minimum, the number of respondents alongside every mean. If two intervals overlap, treat the scores as equal.
  2. Aggregate before you compare. Means are more reliable in aggregate. Combining across items and across a teacher's courses reduces the random fluctuation that drives overinterpretation — Boysen's own recommendation.
  3. Stop ranking on decimals. Use broad bands ("clearly strong / typical / needs a look") rather than ordinal ranks, and never let a fractional difference carry a personnel decision on its own.
  4. Treat the number as a screening question, not an answer. A low band is a prompt to go and understand why — through open text and follow-up — not a conclusion. Averaging ordinal Likert data has its own well-known problems, and Rasch/IRT approaches exist precisely because raw means are shaky.

But doesn't this just make evaluation unusable?

The natural objection: if we cannot trust the decimals, and we cannot rank, what is the data even for?

The answer is that the numbers were never the point — they are a triage signal, and triage does not require decimal precision. Knowing that a course sits in the bottom band across several cohorts, with wide agreement, is genuinely informative and worth acting on. What is not informative is that one instructor scored 4.18 and another 4.06 this semester. The discipline is to use the mean to decide where to look, and to use richer evidence to decide what is actually happening. That preserves everything evaluation is good for while discarding the false precision that harms careers and, ironically, teaching — because chasing a spurious decimal pushes instructors toward leniency and lighter workloads rather than better learning.

Where Koji fits

Koji is designed to move the decision away from the fragile decimal and toward the evidence underneath it. Its AI-moderated conversational interviews gather structured open-text at scale, and its automatic thematic analysis turns that into quote-anchored themes — so a course flagged by a wide-interval mean can be understood, not just ranked. Reporting is built to show patterns across cohorts rather than to crown a monthly winner, which is exactly the aggregate-before-you-compare discipline the evidence calls for. The same AI interview engine underpins general user and customer research on koji.so, where the same rule holds: a small sample's average is a starting question, not a finding.

To be clear about the limits: Koji does not make a small sample large or a noisy mean precise — no tool can repeal sampling error. What it does is reduce the temptation to over-read a single number by putting the reasons in front of the reader, and surface the qualitative signal that a league table throws away. For an evaluation office, that is the difference between managing a spreadsheet of decimals and understanding what students are actually experiencing.

Frequently asked questions

What is the standard error of measurement? It is the margin of error around an observed score, calculated as the standard deviation multiplied by the square root of one minus the instrument's reliability. It tells you how much a mean might have varied by chance, which is why a single evaluation mean should be read as a band, not a point.

Is a 4.2 statistically better than a 3.9? Usually not. In typical class sizes and reliabilities, the confidence intervals around two means that close overlap, which means you cannot conclude one is genuinely higher. Treating such a gap as a real difference is reading noise as signal.

What does the research say about how people read evaluation means? Studies by Boysen and colleagues found that faculty and administrators judged teachers with higher means as more qualified even when differences were under 0.30 on a five-point scale, and continued to over-interpret small gaps even when told the differences were not statistically significant.

Do warnings fix the problem? Only partly. Warning people helped when they compared two teachers, but did not fully prevent department heads from reading meaning into small fluctuations within one teacher's scores over time.

Can you ever compare courses fairly on their means? Only cautiously: report response counts and confidence intervals, aggregate across items and cohorts, use broad bands rather than ranks, and never compare very different course types as if their means were equivalent.

So should we abandon course-evaluation numbers? No — use them as a triage signal to decide where to look, not as a precise verdict. The number tells you a course may need attention; open text and follow-up tell you why.


Want evaluation that explains the number instead of just ranking it? See how Koji for Education works.