New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
analysis-reporting9 min read

Can You Fairly Rank Instructors by Their Course-Evaluation Scores? What Esarey & Valdes (2020) Proved

Even if course evaluations were unbiased, reliable, and valid, ranking instructors by their scores would still misclassify many good teachers. A walk through the Esarey & Valdes (2020) simulation and what it means for how you report and use evaluation data.

Koji Research Desk

Education Research

In short: Even granting student evaluations the most favourable assumptions the research literature allows — that they are unbiased, highly reliable, and moderately correlated with real teaching quality — you still cannot use them to fairly rank or compare individual instructors. A 2020 simulation by Esarey and Valdes shows that a comparison between two instructors with a sizeable gap in scores still fails to identify the better teacher a large share of the time, and that more than a quarter of faculty scoring at or below the 20th percentile are actually above-median teachers. The problem is not (only) bias; it is statistical noise. Evaluations are valuable formative evidence. They are not a ranking instrument.

The question this article answers

Quality-assurance offices and personnel committees routinely line up instructors by their mean evaluation score and treat the ordering as meaningful: this lecturer scored 4.2, that one 3.8, so the first is the better teacher. Is that inference defensible? The most rigorous answer comes not from the bias literature — which is itself damning — but from a study that granted evaluations every benefit of the doubt and still found ranking indefensible.

What the research says

Justin Esarey and Natalie Valdes, writing in Assessment & Evaluation in Higher Education (2020), took a deliberately generous approach. Rather than argue that student evaluations of teaching (SET) are biased — a case made extensively elsewhere — they asked a sharper question: suppose evaluations were as good as their defenders claim. Would they then be fair to use for high-stakes comparisons?

They built a Monte Carlo simulation in which SET scores were, by construction, free of systematic bias (no gender, accent, or discipline penalty), highly reliable, and correlated with true teaching quality at the upper end of what the empirical literature reports (correlations in the region of r = 0.4–0.5, consistent with the most optimistic validity studies). They then simulated large numbers of instructors, drew evaluation scores for them, and asked how often the scores correctly identified the better teacher.

The results were stark. Even with a large observed difference in SET scores between two instructors, the higher-scoring instructor was frequently not the better teacher. Across realistic conditions, a substantial fraction of pairwise comparisons pointed to the wrong person. More memorably, Esarey and Valdes reported that more than a quarter of instructors whose evaluations placed them at or below the 20th percentile were, in reality, above the median in instructional quality. A score that looks like a red flag is, for many faculty, a coin-flip artefact of who happened to be in the room and how the random draw fell.

The mechanism is ordinary sampling noise compounded by an imperfect correlation. Because each instructor is evaluated by a finite, variable group of students, and because even a "valid" instrument only partially tracks teaching quality, the score is a blurry estimate of the underlying construct. Ranking sharpens blurry estimates into false precision.

Corroborating and contrasting work

Esarey and Valdes do not stand alone. Guy Boysen (2015), in Scholarship of Teaching and Learning in Psychology, showed experimentally that faculty and administrators routinely read meaning into small, statistically non-significant differences in evaluation means — and that explicit warnings not to do so did not stop them. Statistical training made no difference; the over-interpretation is cognitive, not informational. Philip Stark and Richard Freishtat (2014), in their widely cited An Evaluation of Course Evaluations, demonstrated that reducing ordinal rating data to a class mean and comparing means across small classes is statistically inappropriate, and recommended reporting distributions rather than averages. And Uttl, White and Gonzalez (2017), re-analysing decades of multisection studies, found that once you control for prior ability, the correlation between evaluations and actual learning is close to zero — which only widens the gap between what a ranking implies and what it can support.

The throughline: the score contains real information, but far less, and far noisier, than a ranked list pretends.

Why it matters for course evaluation in practice

If you operate a quality-assurance process, three practices are quietly indefensible in light of this evidence:

  1. Sorting instructors into a league table and treating position as a quality signal.
  2. Flagging "low performers" by percentile without a margin of uncertainty — Esarey and Valdes show a quarter of the bottom quintile are above-median teachers.
  3. Reporting a single mean as the headline number, which invites exactly the small-difference over-interpretation Boysen documented.

None of this means evaluations are worthless. It means their legitimate use is formative and diagnostic — surfacing themes, identifying courses that need a conversation, tracking a programme over time — not summative and comparative at the level of the individual. The American Sociological Association, among others, has formally recommended against using SET scores to compare individual faculty for personnel decisions, citing precisely this body of work.

Limitations and honest caveats

A critical reader should hold this evidence to its own standard.

  • It is a simulation. Esarey and Valdes assume a data-generating process; if real evaluations were more tightly correlated with teaching quality than the optimistic literature suggests, the error rates would fall. But the simulation already used optimistic inputs — the realistic case is worse, not better.
  • "True teaching quality" is itself contested. The simulation treats it as a latent variable; defining and measuring it is the hard problem the whole field circles around.
  • Reliability is conditional. Marsh and others have shown evaluations can be highly reliable with enough raters per class. The noise problem is most acute in small classes and at low response rates — which is exactly where percentile-ranking does most damage.
  • Aggregate use is more defensible than individual use. Programme-level or multi-cohort trends average out much of the noise that wrecks individual comparison. The argument here is against ranking people, not against using evaluation data at all.

Acknowledging these limits is not a retreat. It sharpens the conclusion: evaluations are decent evidence about courses and programmes over time, and poor evidence for ranking individuals at a point in time.

How Koji incorporates this

Koji is designed so that the unit of insight is the theme, not the league-table position — which is the practical answer to the Esarey–Valdes problem.

  • Distributions over averages. Koji reporting is built to surface the spread and shape of responses, not a single decontextualised mean, directly addressing the Stark–Freishtat critique and the Boysen over-interpretation trap.
  • Uncertainty-aware reporting. Where response counts are low, Koji is designed to flag that a result is within the range of noise rather than presenting a precise-looking number that invites false ranking. The aim is to make "we cannot reliably distinguish these two courses" a visible, first-class output.
  • Qualitative triangulation instead of scalar ranking. Koji''s AI-moderated conversational interviews probe why a student rated as they did, producing structured, thematically analysed evidence about specific aspects of a course. This shifts the decision basis from "who is higher on the list" to "what concretely should change", which is the use the evidence actually supports.
  • Structured, multi-signal questions. By combining scale items with open_ended, single_choice, and ranking questions and analysing them together, Koji avoids resting a judgement on one fragile average.
  • Programme- and cohort-level aggregation. Koji is built to track themes across cohorts over time, where the law of large numbers makes evaluation evidence genuinely informative — the defensible end of the spectrum.

None of this eliminates sampling noise; nothing can. Koji is designed to keep that noise visible and to steer users away from the single inference — ranking individuals on a mean — that the research most clearly rules out. Koji''s core research platform at koji.so applies the same distribution-first, theme-led analysis to product and customer research, where over-reading a single average is an equally common error.

A practical reporting checklist

If you take one operational change from this evidence, make it the way evaluation results are presented, because that is where the Esarey–Valdes and Boysen findings bite. A defensible report does five things. First, it shows the full distribution of responses for each item, not just the mean — a 4.0 built from straight 4s means something different from a 4.0 split between 2s and 5s. Second, it attaches a measure of uncertainty (a confidence interval, an interquartile range, or simply the response count) so that two courses one-tenth of a point apart are visibly indistinguishable. Third, it suppresses rankings of individual instructors entirely, replacing "position in the cohort" with "themes and changes". Fourth, it aggregates to the level the data can support — programme and multi-cohort trends rather than single small classes. Fifth, it pairs every quantitative summary with the qualitative evidence that explains it, so decisions rest on reasons, not on a contested decimal. A report that does these five things is far harder to misuse than a league table, and far more useful to the person who actually has to improve the course.

Related Resources

References

  • Esarey, J., & Valdes, N. (2020). Unbiased, reliable, and valid student evaluations can still be unfair. Assessment & Evaluation in Higher Education, 45(8), 1106–1120. https://doi.org/10.1080/02602938.2020.1724875
  • Boysen, G. A. (2015). Significant interpretation of small mean differences in student evaluations of teaching despite explicit warning to avoid overinterpretation. Scholarship of Teaching and Learning in Psychology, 1(2), 150–162. https://doi.org/10.1037/stl0000017
  • Stark, P. B., & Freishtat, R. (2014). An evaluation of course evaluations. ScienceOpen Research. https://doi.org/10.14293/S2199-1006.1.SOR-EDU.AOFRQA.v1
  • Uttl, B., White, C. A., & Gonzalez, D. W. (2017). Meta-analysis of faculty''s teaching effectiveness: Student evaluation of teaching ratings and student learning are not related. Studies in Educational Evaluation, 54, 22–42. https://doi.org/10.1016/j.stueduc.2016.08.007