An Evaluation of Course Evaluations: What the Berkeley Statisticians Concluded About Reporting Student Ratings
Two Berkeley statisticians argued that averaging student-rating scores to rank instructors for promotion and tenure should be abandoned for both statistical and substantive reasons. What Stark & Freishtat (2014) actually recommend — and how to report ratings responsibly instead.
Koji Education Team
Product
In brief: In An Evaluation of Course Evaluations (2014), Berkeley statisticians Philip Stark and Richard Freishtat argued that the standard practice of averaging student-rating scores and ranking instructors by the mean for promotion and tenure decisions should be abandoned — for statistical reasons (ordinal data should not be averaged; small samples and non-response make means unstable; tiny differences are treated as meaningful) and substantive reasons (ratings are confounded by factors unrelated to teaching). Their constructive alternative: report distributions, not just averages, stop over-interpreting small gaps, and treat student ratings as one input alongside peer review and teaching portfolios.
What the research says
Most institutions reduce a course evaluation to a single number — the mean of a set of Likert items — and then compare that number across instructors, against a departmental average, or against a fixed threshold. Stark and Freishtat's contribution was to subject that practice — the averaging and ranking, not just the survey — to statistical scrutiny.
The anchor: Stark & Freishtat (2014)
Published open-access in ScienceOpen Research, "An Evaluation of Course Evaluations" makes a layered argument. The headline claim is blunt: relying on averages of student teaching-evaluation scores as the primary measure of teaching effectiveness for personnel decisions "should be abandoned for substantive and statistical reasons." The supporting points, paraphrased:
- Averaging ordinal ratings is statistically improper. A 1–5 rating scale is ordinal: the distance from "1" to "2" is not guaranteed to equal the distance from "4" to "5". Computing a mean (and the difference between two means) treats ordinal categories as if they were equally-spaced numbers on an interval scale. The median and the full distribution are more defensible summaries.
- Small differences are noise, not signal. With the typical class sizes and response rates, the difference between, say, a 4.1 and a 4.3 is well within the range of random variation. Ranking or thresholding instructors on such gaps manufactures precision that the data do not support.
- Non-response and small samples bias the mean. When only a fraction of a class responds — and responders differ from non-responders — the average is a biased estimate of the class's view, and that bias is unmeasured.
- Ratings are confounded. Scores are associated with class size, discipline, electivity, grading, and instructor demographics — so a difference in means can reflect the course or the instructor's identity rather than teaching quality. (Stark's related work with Boring and Ottoboni, 2016, demonstrated gender bias directly.)
- The constructive recommendation. Because of all this, ratings should not stand alone. Stark and Freishtat advocate teaching portfolios and peer observation of teaching as primary routes to measuring and improving teaching, with student feedback used for formative improvement and reported as distributions rather than reduced to a single rankable average.
Corroborating evidence
The statistical case does not rest on this paper alone. Boring, Ottoboni and Stark (2016) showed experimentally and observationally that student ratings can measure students' gender bias rather than teaching effectiveness. Esarey and Valdes (2020) demonstrated through simulation that even when ratings are weakly and unbiasedly related to teaching quality, using them to rank or reward faculty produces a high rate of unfair decisions, because the noise swamps the signal at the margins where decisions are actually made. And Uttl, White and Gonzalez (2017), re-analysing the multisection validity literature, found that once study quality and sample size are accounted for, the rating–learning correlation is close to zero. Three independent lines — measurement, simulation, meta-analysis — converge on the same caution.
Why it matters for course evaluation in practice
For a quality-assurance office, institutional-research team, or promotion committee, the implications are concrete and immediately actionable:
- Stop ranking instructors on raw means. A sorted list of mean scores invites exactly the over-interpretation Stark and Freishtat warn against. If you must summarise, report the median and the full distribution (the proportion in each scale category), and show the response rate beside every figure.
- Attach uncertainty. A mean without a confidence interval or distribution is a point estimate masquerading as a fact. Where two instructors' intervals overlap heavily, treat them as indistinguishable.
- Set a "minimum interpretable difference". Decide in advance how large a gap must be before it is treated as meaningful, and refuse to act on smaller ones. This pre-commitment defends against reading patterns into noise.
- Demote ratings to one input. Use them formatively (to help instructors improve) and as one strand of a multi-source dossier — peer observation, teaching portfolio, learning-outcome evidence — for summative decisions. This is the heart of the Stark–Freishtat recommendation.
- Report response context. A 4.5 from 12 of 200 students is not a 4.5; the non-response makes it uninterpretable as a class average. Display N, response rate, and distribution, not just the mean.
A concrete reporting contrast shows what changes in practice. The old report reads: "Instructor mean 4.1 (department mean 4.3)" — an invitation to conclude the instructor is "below average". The Stark–Freishtat-aligned report reads: "N = 38 of 95 (40% response); median 4; distribution 5★ 30% / 4★ 38% / 3★ 20% / 2★ 8% / 1★ 4%; difference from department median not distinguishable given response rate and spread." The second version makes the uncertainty visible, refuses to manufacture a ranking from a 0.2 gap, and shows the non-response that makes the "average" an unreliable estimate of the class's view in the first place. Nothing is hidden, but nothing is overclaimed either — which is precisely the standard the paper sets.
Limitations and honest caveats
A careful reader should weigh several counterpoints:
- It is a position paper, not a new empirical study. An Evaluation of Course Evaluations synthesises statistical principles and prior findings; it does not present a fresh dataset. Its force comes from the soundness of the argument and the corroborating literature, not from new data.
- The ordinal-data objection is contested. Many psychometricians argue that, in practice, averaging well-constructed multi-point scales is robust and that means of Likert composites behave acceptably as interval data. The strict ordinal purism is a defensible methodological stance, not a settled universal truth.
- Alternatives have their own problems. Peer observation suffers from low inter-rater reliability and small samples; teaching portfolios are labour-intensive and hard to standardise. "Use portfolios and peer review instead" trades one set of measurement problems for another, as the convergent-validity literature shows.
- "Abandon averages" can be over-read. The authors target high-stakes ranking on small differences, not all quantitative summary. Distributions are still numbers; the recommendation is to summarise honestly, not to refuse measurement.
- Context matters. With large, high-response samples and modest stakes, a mean is far less dangerous than in the small-class, high-stakes, marginal-decision setting the paper foregrounds.
The robust, defensible takeaway is procedural: report distributions and uncertainty, do not rank on small differences, and never let a single average drive a career decision.
How Koji incorporates this
Koji's analysis and reporting layer is designed around exactly the principles Stark and Freishtat advocate — turning their critique into default behaviour rather than an optional discipline. These features mitigate the misuse they identify; they cannot, of course, eliminate the human temptation to rank.
- Distribution-first reporting. Koji is built to surface the full distribution and the median, the response rate, and the sample size — not a lone average. Committees see the shape of the responses, so a bimodal or thin-response result cannot hide behind a tidy mean.
- Qualitative depth that a number cannot carry. Koji's AI-moderated conversational interviews generate rich open-text that is automatically thematically analysed and quality-scored, giving the formative signal Stark and Freishtat say student feedback is best suited for — the "why" behind the rating.
- Triangulation by design. Because Koji structures evidence across cohorts, time, and question types (
scale,open_ended,single_choice,ranking,yes_no), it is suited to being one strand in a multi-source dossier alongside peer observation and portfolios, rather than the sole metric. - Bias-aware framing. Reporting is designed to discourage over-interpretation of small gaps and to present scores with their uncertainty, directly addressing the "small differences are noise" problem.
- Closing-the-loop action tracking reinforces formative use: the question becomes "what did we change and did it help", not "who ranked highest".
Koji is designed to mitigate the reporting failures Stark and Freishtat diagnose by defaulting to distributions, context, and qualitative depth — not to promise that any single output is an unimpeachable measure of teaching. The core research platform at koji.so applies the same distribution-first, interview-rich philosophy to product and customer research, where averaging a five-point satisfaction scale into a single rank is an equally common and equally fragile habit.
Related Resources
- Can You Fairly Rank Instructors by Their Course-Evaluation Scores? Esarey & Valdes (2020)
- The 4.2 vs 4.4 Trap: Why Small Differences in Means Are Usually Noise
- Interpreting and Reporting Student Ratings Responsibly: Linse (2017)
- Peer Observation vs Student Evaluations: What Each Actually Measures
- Generalizability Theory and the Reliability of Student Ratings
- Do Student Evaluations Measure Learning? The Uttl Meta-Analysis Revisited
References
- Stark, P. B., & Freishtat, R. (2014). An Evaluation of Course Evaluations. ScienceOpen Research. https://doi.org/10.14293/S2199-1006.1.SOR-EDU.AOFRQA.v1
- Boring, A., Ottoboni, K., & Stark, P. B. (2016). Student evaluations of teaching (mostly) do not measure teaching effectiveness. ScienceOpen Research. https://doi.org/10.14293/S2199-1006.1.SOR-EDU.AETBZC.v1
- Esarey, J., & Valdes, N. (2020). Unbiased, reliable, and valid student evaluations can still be unfair. Assessment & Evaluation in Higher Education, 45(8), 1106–1120. https://doi.org/10.1080/02602938.2020.1724875
- Uttl, B., White, C. A., & Gonzalez, D. W. (2017). Meta-analysis of faculty's teaching effectiveness: Student evaluation of teaching ratings and student learning are not related. Studies in Educational Evaluation, 54, 22–42. https://doi.org/10.1016/j.stueduc.2016.08.007
Related articles
Can You Fairly Rank Instructors by Their Course-Evaluation Scores? What Esarey & Valdes (2020) Proved
Even if course evaluations were unbiased, reliable, and valid, ranking instructors by their scores would still misclassify many good teachers. A walk through the Esarey & Valdes (2020) simulation and what it means for how you report and use evaluation data.
Generalizability Theory and the Reliability of Student Ratings: Why One Class Is Not Enough
Reliability is not one number. Generalizability theory (Gillmore, Kane & Naccarato 1978; Marsh 1984) decomposes the variance in student ratings into student, teacher, course and occasion components — and shows that a single class can be reliable for the course yet a poor estimate of the teacher. What that means for fair evaluation.
Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees
The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.
The 4.2 vs 4.4 Trap: Why Small Differences in Course-Evaluation Means Are Usually Noise
Faculty and administrators routinely read meaning into tiny gaps between course-evaluation means. Boysen (2015) and Boysen et al. (2014) show this happens even when confidence intervals say the difference is nothing — and that warnings barely help. How to report uncertainty honestly.