The 4.2 vs 4.4 Trap: Why Small Differences in Course-Evaluation Means Are Usually Noise
Faculty and administrators routinely read meaning into tiny gaps between course-evaluation means. Boysen (2015) and Boysen et al. (2014) show this happens even when confidence intervals say the difference is nothing — and that warnings barely help. How to report uncertainty honestly.
Koji Education Team
Product
In brief: A 4.2 versus a 4.4 on a course-evaluation form almost never represents a real difference in teaching, yet faculty and administrators consistently interpret such gaps as meaningful. Boysen (2015) and Boysen, Kelly, Raesly and Casner (2014) demonstrated this over-interpretation experimentally — it persists even when confidence intervals and significance tests indicate no difference, and even after explicit warnings. The practical fix is to stop reporting bare means: report distributions, confidence intervals, and the number of responses, and refuse to rank on differences smaller than measurement error.
The most common misuse of course-evaluation data
More damage is probably done by the interpretation of course-evaluation scores than by the scores themselves. A department head compares two lecturers at 4.2 and 4.4 and concludes one is the stronger teacher. A promotion committee notes a candidate "dipped" from 4.5 to 4.3 and asks them to explain. A programme report flags a module that fell "below the faculty average" of 4.1 by scoring 4.0. In every case, a difference well inside the margin of error is being treated as signal. This is not a fringe error; the research shows it is the default behaviour of trained academics.
What the research says
The central evidence comes from Guy Boysen. In Boysen, Kelly, Raesly and Casner (2014), "The (mis)interpretation of teaching evaluations by college faculty and administrators" (Assessment & Evaluation in Higher Education, 39(6), 641-656), participants were shown course-evaluation means and asked to interpret them. Faculty interpreted small differences between means as indicating genuine differences in teaching quality even when accompanying confidence intervals and statistical tests indicated the absence of a meaningful difference. Differences explicitly labelled non-significant still shifted perceptions of teaching ability and qualifications.
Boysen (2015), "Uses and Misuses of Student Evaluations of Teaching: The Interpretation of Differences in Teaching Evaluation Means Irrespective of Statistical Information" (Teaching of Psychology, 42(2)), extended this. Across studies, participants drew conclusions about teaching from trivially small mean differences regardless of the statistical information provided. A particularly sobering finding: teachers from disciplines with very different levels of statistics training showed no difference in their tendency to over-interpret — statistical knowledge alone did not inoculate them.
Boysen also tested whether warnings help, in work on "Preventing the Overinterpretation of Small Mean Differences in Student Evaluations of Teaching" (Scholarship of Teaching and Learning in Psychology, 1(2), 150-162). The results were mixed and modest: an explicit warning could eliminate over-interpretation when participants compared two different teachers, but warnings only partially reduced the tendency when department heads interpreted variation within a single teacher's set of means. In other words, you cannot caption your way out of the problem.
This sits inside a broader methodological consensus. Work on the reliability of student ratings (generalizability theory) and on ranking instructors fairly (Esarey & Valdes, 2020) reaches the same destination from a different direction: the measurement error around a course-evaluation mean is large enough that fine-grained rankings and small gaps are not interpretable.
Why it matters for course evaluation in practice
The implication is concrete and actionable: a course-evaluation mean reported without a measure of uncertainty is an invitation to misuse. Three practices follow.
Report the distribution, not just the average. A 4.0 built from a tight cluster of 4s is a different reality from a 4.0 built from a bimodal split of 2s and 5s, yet they share a mean. Histograms and the percentage in the top categories tell committees what the average conceals.
Always attach a confidence interval and the response count. With twenty responses, the 95% confidence interval around a mean on a five-point scale is wide — often ±0.3 or more depending on the spread. If two intervals overlap heavily, the means are, for decision-making purposes, the same. Reporting n alongside the interval also exposes the small samples where a single response moves the average.
Set a "minimum meaningful difference" and enforce it. Because Boysen shows that people over-interpret even when given the statistics, the durable safeguard is procedural: institutional policy that explicitly forbids ranking, flagging, or personnel inferences based on differences within the margin of error. The discipline has to be built into the reporting template, not left to the reader's self-control.
For accreditation, this is also a credibility issue. A quality process that ranks staff on 0.1 gaps will not withstand scrutiny from a reviewer who knows the measurement literature; one that reports uncertainty and triangulates demonstrates methodological maturity.
Limitations and honest caveats
A careful reader should note the boundaries of this evidence. Boysen's studies are largely vignette experiments with academic participants in (predominantly North American) psychology and related fields; the exact magnitude of over-interpretation may differ in other systems and disciplines. The studies establish that the bias is robust and resistant to warnings, not a precise population rate.
There is also a genuine tension: insisting on uncertainty can tip into nihilism — "the numbers mean nothing, so ignore them." That is not the lesson. Large, consistent differences across multiple cohorts, supported by open-text evidence and other sources, can be informative. The point is calibration: small single-cohort gaps are noise; stable multi-source patterns are signal. Finally, confidence intervals themselves assume reasonably behaved data; with heavily skewed ceiling-effect distributions (common in SET), they are an approximation, and distribution-aware reporting matters more than any single interval.
How Koji incorporates this
Koji is designed to make the honest interpretation the easy interpretation.
- Uncertainty-aware reporting by default. Rather than surfacing a bare mean, Koji's reporting is built to show distributions, response counts, and the spread behind a score, so committees see whether a 4.2/4.4 gap is real or noise. This directly targets the false-precision behaviour Boysen documents.
- Triangulation over single numbers. Because Koji captures structured ratings and probed open-text responses, a small mean difference is automatically set against qualitative evidence — discouraging conclusions drawn from one fragile average. Automatic thematic analysis identifies whether the reasons differ, not just the digits.
- Cohort comparison framed as patterns, not point gaps. Koji is oriented toward stability across cohorts rather than one-off rankings, aligning with the generalizability-theory lesson that one class is not enough.
- Bias-aware presentation for committees. By presenting scores with their context rather than as a leaderboard, Koji is designed to mitigate (not eliminate) the over-interpretation that persists even among statistically trained readers.
No tool can stop a determined misreader, and Koji does not claim to — but defaulting to distributions and uncertainty removes the most common path to misuse. The same uncertainty-honest reporting philosophy underpins the core Koji research platform at koji.so, where product and customer teams face the identical temptation to over-read small movements in a metric.
A reporting template that resists misuse
Because Boysen shows that exhortation does not work, the durable remedy is to redesign the report so that over-interpretation becomes difficult. A defensible course-evaluation summary contains, for each item, four elements rather than one. First, the mean with its 95% confidence interval, written so the interval is impossible to ignore — for example "4.2 (95% CI 3.9-4.5, n = 22)" rather than a lone "4.2". Second, the full distribution, as a small histogram or the percentage in each scale point, so a reader can see whether a 4.0 is a tight consensus or a polarised average of 2s and 5s. Third, the response count and rate, which signals at a glance where the numbers rest on too few students to support any inference. Fourth, a comparison rule stated in the template itself: differences are reported as meaningful only when intervals do not substantially overlap and the pattern holds across cohorts.
The effect of this layout is to make the honest reading the path of least resistance. A department head who sees "4.2 (3.9-4.5)" beside "4.4 (4.1-4.7)" is far less likely to declare one lecturer superior than one who sees a naked "4.2" next to "4.4". Pairing the table with a short standing caption — that gaps within the margin of error must not drive ranking or personnel decisions — adds a second, procedural guard. None of this is sophisticated statistics; it is disciplined presentation. The research is unambiguous that the alternative, a leaderboard of bare means, will be misread even by readers who know better, and that a quality process built on such misreadings will not survive informed scrutiny.
Related Resources
- Can You Fairly Rank Instructors by Their Course-Evaluation Scores?
- Interpreting and Reporting Student Ratings Responsibly
- How Many Responses Do You Need for a Reliable Course Evaluation?
- Generalizability Theory and the Reliability of Student Ratings
- Cronbach's Alpha and Course-Evaluation Reliability
- Measurement Invariance: Can You Compare Scores Across Groups?
References
- Boysen, G. A., Kelly, T. J., Raesly, H. N., & Casner, R. W. (2014). The (mis)interpretation of teaching evaluations by college faculty and administrators. Assessment & Evaluation in Higher Education, 39(6), 641-656. https://doi.org/10.1080/02602938.2013.860950
- Boysen, G. A. (2015). Uses and Misuses of Student Evaluations of Teaching: The Interpretation of Differences in Teaching Evaluation Means Irrespective of Statistical Information. Teaching of Psychology, 42(2), 109-118. https://doi.org/10.1177/0098628315569922
- Boysen, G. A. (2015). Preventing the Overinterpretation of Small Mean Differences in Student Evaluations of Teaching: An Evaluation of Warning Effectiveness. Scholarship of Teaching and Learning in Psychology, 1(2), 150-162. https://doi.org/10.1037/stl0000017
- Esarey, J., & Valdes, N. (2020). Unbiased, reliable, and valid student evaluations can still be unfair. Assessment & Evaluation in Higher Education, 45(8), 1106-1120. https://doi.org/10.1080/02602938.2020.1724875
Related articles
How Many Responses Do You Need for a Reliable Course Evaluation?
Reliability of a class-average rating rises steeply with the number of students who respond. Generalizability theory — from Marsh''s syntheses to Dzakadzie (2026) — shows raters matter more than items, and a class of 12 cannot be evaluated like a class of 120.
Can You Fairly Rank Instructors by Their Course-Evaluation Scores? What Esarey & Valdes (2020) Proved
Even if course evaluations were unbiased, reliable, and valid, ranking instructors by their scores would still misclassify many good teachers. A walk through the Esarey & Valdes (2020) simulation and what it means for how you report and use evaluation data.
Generalizability Theory and the Reliability of Student Ratings: Why One Class Is Not Enough
Reliability is not one number. Generalizability theory (Gillmore, Kane & Naccarato 1978; Marsh 1984) decomposes the variance in student ratings into student, teacher, course and occasion components — and shows that a single class can be reliable for the course yet a poor estimate of the teacher. What that means for fair evaluation.
Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees
The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.