Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees
The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.
Koji Education Team
Product
Answer box
The most damaging errors in student evaluation happen after the data is collected, in how it is read. Drawing on more than eight decades of research, Linse (2017) shows that administrators and personnel committees routinely misuse student ratings — comparing instructors against a single institutional average, treating tiny mean differences as meaningful, and ignoring the qualitative context behind the numbers. Boysen (2015) demonstrated experimentally that even statistically trained faculty interpret small, non-significant differences as real, even after being explicitly warned not to. The fix is procedural: report distributions and uncertainty rather than bare means, compare like-with-like, never rank on differences smaller than the noise, and always pair numbers with the open-text reasoning that explains them.
What the research says
Linse (2017), writing in Studies in Educational Evaluation, distils the literature into guidance for the people who actually make decisions: faculty serving as administrators and on evaluation committees. Her central argument is that student ratings can be useful evidence of teaching when interpreted correctly, but are routinely interpreted incorrectly in ways that disadvantage instructors and produce indefensible decisions. Her key, research-grounded recommendations include:
- Do not compare an instructor's mean to a single global benchmark. Ratings vary systematically with factors outside the instructor's control — class size, discipline, course level, whether the course is required or elective. A fair comparison is against similar courses, not the whole institution.
- Stop treating small mean differences as meaningful. The difference between a 4.2 and a 4.4 is almost always within the margin of error; ranking faculty on such gaps is statistically illiterate.
- Use ranges, distributions and multiple years, not a single decontextualised number.
- Triangulate. Student ratings should be one of several sources (peer observation, self-reflection, teaching materials), never the sole measure of teaching effectiveness.
Boysen (2015), in Scholarship of Teaching and Learning in Psychology, supplies the experimental teeth. He gave 225 English, mathematics and psychology teachers statistical reports for a fictional instructor and asked them to interpret small differences in the ratings. The teachers reliably judged trivial, non-significant differences as indicating a genuine need to improve a course — and, strikingly, statistical training did not help: despite the three disciplines differing sharply in statistics background, all over-interpreted equally, and an explicit warning to avoid over-interpretation in related work did little to prevent it. The lesson is sobering: you cannot rely on evaluators' good judgement or training to stop over-interpretation; you must engineer it out of the reporting format itself.
Stark & Freishtat (2014), in their widely-circulated "An Evaluation of Course Evaluations," reinforce both points from a statistician's standpoint. They argue that averaging ordinal Likert responses is itself questionable, that response rates and distributions matter far more than the mean, and that small differences in averages are essentially meaningless for personnel decisions. Their recommendation — report the distribution of scores and the response rate, and avoid ranking — converges precisely with Linse and Boysen.
The three sources form a tight evidentiary chain: Linse says here is how to interpret ratings correctly; Boysen shows people fail to do it even when warned; Stark & Freishtat prescribe report formats that make the failure harder. The agreement across a practitioner review, an experiment, and a statistical critique is what makes the guidance authoritative.
Why it matters for course evaluation in practice
For European QA officers, deans and programme directors, this body of work reframes the central risk. The instrument can be well-designed and the bias literature fully respected, yet decisions can still be unfair if the reporting layer invites misreading. Three habits do most of the harm:
- Single-benchmark ranking. Sorting a department's instructors by mean against one institutional average rewards those who happen to teach small, elective, advanced classes and penalises those teaching large required quantitative modules — independent of teaching quality.
- Spurious precision. Reporting "4.27" implies a precision the data does not have. Boysen shows readers will act on the second decimal place even when it is noise.
- Numbers without narrative. A mean with no open-text context cannot tell a committee why a course scored as it did, so reviewers fill the gap with assumption — often the very biases documented for gender, race and difficulty.
The defensible alternative is a reporting standard: show the full distribution and the number of respondents; benchmark only against comparable courses; attach confidence/uncertainty so trivial gaps are visibly trivial; suppress or caveat results below a response-rate threshold; and present open-text themes alongside every score. This is not "going easy" on teaching evaluation — it is making it rigorous enough to survive scrutiny.
Limitations and honest caveats
- Guidance is largely US-derived. Linse, Boysen and Stark & Freishtat write primarily about US institutions; European norms (ESG/ENQA, NVAO, national frameworks) differ in governance and culture. The principles transfer; specific thresholds should be set locally.
- "Compare like-with-like" needs enough data. Fair benchmarking against similar courses requires a large enough comparison pool. In small departments, comparable groups may be too small to be stable — a real constraint, not a solved problem.
- Distributions can be misread too. Showing distributions helps, but untrained readers may over-interpret a single low cluster. Reporting format reduces error; it does not eliminate the need for evaluator literacy.
- Boysen's stimuli were fictional reports. The experiment shows a robust tendency to over-interpret; real committees with discussion and accountability may behave somewhat differently. The direction of the finding is clear; the exact real-world magnitude is uncertain.
- No reporting standard fixes a biased instrument. Responsible interpretation is necessary but not sufficient; it must sit on top of a valid, bias-aware instrument.
How Koji incorporates this
Koji treats interpretation as a first-class design problem, not an afterthought left to committees:
- Distributions and response rates by default. Koji's reporting surfaces the full distribution of responses and the number of respondents prominently, rather than leading with a single mean — operationalising the Stark & Freishtat / Linse recommendation directly.
- Comparable-context benchmarking. Reports are designed to compare a course against similar offerings (level, discipline, size, format), discouraging the single-institutional-average ranking Linse identifies as a core error.
- Uncertainty made visible. By presenting ranges and flagging when differences are within expected variation, Koji makes trivial gaps look trivial — the structural guard Boysen shows is necessary because warnings alone fail.
- Every score paired with reasoning. Because Koji collects AI-moderated open-text responses and runs automatic thematic analysis, each numeric result arrives with the why attached, so committees read meaning rather than imputing it.
- Triangulation-ready exports. Koji is built to be one evidence stream among several, exporting structured outputs that sit alongside peer observation and self-evaluation rather than standing in as the sole measure.
- Response-rate caveats. Low-response results can be flagged so they are read with appropriate caution instead of being treated as equivalent to well-sampled ones.
Stated honestly, Koji cannot force a committee to reason well — but it can make the correct interpretation the path of least resistance and the incorrect one visibly unsupported. The same reporting philosophy applies to customer and product research on koji.so, where over-reading a small metric difference is an equally common and costly mistake.
A checklist for evaluation committees
The research converges on a short, operational checklist that any personnel or quality-assurance committee can adopt without specialist statistics. Use it as a gate before any student-rating number influences a decision:
- Is the response rate adequate, and is it shown? Treat results below your agreed threshold as indicative only, and never compare a 30%-response course against a 90%-response one as equals.
- Are you looking at the distribution, not just the mean? A 4.0 built from a tight cluster and a 4.0 hiding a bimodal split are completely different teaching stories. Read the shape.
- Is the comparison like-with-like? Benchmark against similar courses — same level, discipline, size and required/elective status — not against a single institutional average (Linse 2017).
- Is the difference bigger than the noise? If two instructors differ by a couple of tenths of a point, treat them as tied. Boysen (2015) shows even trained readers wrongly act on such gaps, so the format must make trivial differences look trivial.
- Is there qualitative context? Every score should arrive with open-text themes that explain why; numbers without narrative invite biased guesswork.
- Is this one source among several? Student ratings should sit alongside peer observation, self-reflection and teaching materials, never stand alone.
- Is there enough evidence over time? A single semester is rarely sufficient for a high-stakes judgement about an individual.
A committee that can answer "yes" to all seven has a decision that will survive appeal, audit and accreditation scrutiny. One that cannot is exposed — not because the survey was bad, but because the interpretation was. Embedding this checklist in policy is the single highest-leverage change most institutions can make.
Related Resources
- Can You Fairly Rank Instructors by Their Scores?
- How Many Responses Do You Need for a Reliable Course Evaluation?
- What Do Student Evaluations Actually Measure?
- Do Student Evaluations Measure Learning?
- Text Analytics for Open-Ended Student Comments
- Does Grading Leniency Inflate Student Evaluations?
References
- Linse, A. R. (2017). Interpreting and using student ratings data: Guidance for faculty serving as administrators and on evaluation committees. Studies in Educational Evaluation, 54, 94–106. https://doi.org/10.1016/j.stueduc.2016.12.004
- Boysen, G. A. (2015). Significant interpretation of small mean differences in student evaluations of teaching despite explicit warning to avoid overinterpretation. Scholarship of Teaching and Learning in Psychology, 1(2), 150–162. https://doi.org/10.1037/stl0000017
- Stark, P. B., & Freishtat, R. (2014). An evaluation of course evaluations. ScienceOpen Research. https://doi.org/10.14293/S2199-1006.1.SOR-EDU.AOFRQA.v1
Related articles
How Many Responses Do You Need for a Reliable Course Evaluation?
Reliability of a class-average rating rises steeply with the number of students who respond. Generalizability theory — from Marsh''s syntheses to Dzakadzie (2026) — shows raters matter more than items, and a class of 12 cannot be evaluated like a class of 120.
Text Analytics for Open-Ended Student Comments: What NLP Can and Cannot Tell You
Natural-language processing can turn thousands of free-text course-evaluation comments into themes and sentiment at scale — but the research (Cunningham-Nelson 2019; Sunar & Khalid 2023) shows where automated analysis is reliable and where human judgement is still required.
Can You Fairly Rank Instructors by Their Course-Evaluation Scores? What Esarey & Valdes (2020) Proved
Even if course evaluations were unbiased, reliable, and valid, ranking instructors by their scores would still misclassify many good teachers. A walk through the Esarey & Valdes (2020) simulation and what it means for how you report and use evaluation data.
Do Student Evaluations Measure Learning? The Uttl Meta-Analysis Revisited
A research-grounded reading of the Uttl, White & Gonzalez (2017) meta-analysis on the SET–learning relationship, with implications for European course evaluation and quality-assurance policy.