New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
analysis-reporting10 min read

Even Fair Evaluations Can Rank Unfairly: Why Course-Evaluation Scores Misclassify Instructors

A simulation-based finding that even unbiased, reliable, and valid course-evaluation scores produce high misclassification rates when used to rank or compare instructors — and what it means for tenure, promotion, and merit decisions.

Koji Education Team

Product

In brief

Even if a course-evaluation score were perfectly unbiased, reliable, and a valid signal of teaching quality, using it to rank instructors against each other still produces a high rate of wrong decisions. Simulation evidence shows that more than a quarter of instructors flagged in the bottom fifth of scores are actually above-average teachers, and that even large score gaps fail to reliably identify the better teacher. The problem is not (only) bias — it is the statistics of comparing noisy, imperfectly correlated measures, and it has direct consequences for tenure, promotion, and merit.

What the research says

The pivotal study is Esarey and Valdes (2020) in Assessment & Evaluation in Higher Education. Rather than re-litigate whether course evaluations are biased, the authors granted them every benefit of the doubt and asked a different question: if student evaluations of teaching (SET) were a best-case measure — moderately correlated with true teaching quality, highly reliable, and non-discriminatory — would it then be fair to use them to compare faculty? Using a Monte Carlo simulation in which the "true" teaching quality of each instructor is known and SET is a clean, unbiased signal of it, they tracked how often score-based decisions matched reality.

The results are sobering even under these idealized assumptions. They report that more than a quarter of faculty whose evaluations sit at or below the 20th percentile are in fact above the median in instructional quality. A score that places an instructor in the "bottom fifth" — the kind of result that triggers remediation, denies merit pay, or counts against tenure — is wrong about that instructor being below-average more than a quarter of the time. Worse, they show that even a large gap in SET scores between two instructors fails to reliably identify which is the better teacher in a pairwise comparison. Their recommendation is not to discard SET but to stop using it as a standalone ranking instrument and instead combine several imperfect measures of teaching.

This builds on long-standing methodological warnings. Stark and Freishtat (2014), in An Evaluation of Course Evaluations, argue that averaging ordinal response categories (turning "good/very good/excellent" into a mean) is not statistically justified, that response rates below 100% make estimates less representative and less reliable, and that comparing an instructor's average against a department average over-reads noise as signal. Boysen, Kelly, Raesly, and Casner (2014) supplied the behavioural complement: across three studies, faculty and administrators treated mean differences small enough to fall within the margin of error as meaningful, rewarding or penalizing instructors on the basis of differences that statistics cannot distinguish from zero. Boysen's follow-up work (2015) confirmed that decision-makers apply a "higher is better" heuristic irrespective of statistical information.

Why is even a reliable instrument so error-prone for ranking? The answer lies in generalizability theory, applied to student ratings as far back as Gillmore, Kane, and Naccarato (1978), who decomposed ratings variance into instructor and course components and showed that a dependable estimate of an instructor effect requires data aggregated across multiple sections or many responses. A single course's mean is a noisy draw; ranking on noisy draws guarantees frequent reversals. Professional bodies have since formalized the caution: the American Sociological Association's 2019 statement on student evaluations of teaching, endorsed by numerous scholarly organizations, urges institutions not to use SET to compare individual faculty to one another or to a department average, and to embed them in a holistic, multi-measure, pattern-over-time framework.

Why it matters for course evaluation in practice

The temptation in any quality-assurance or personnel process is to convert a distribution of student opinion into a single comparable number and then rank. Esarey and Valdes show this is precisely where the instrument breaks — not because students are malicious or biased (the simulation assumes they are neither) but because a moderately-correlated, finitely-reliable signal cannot support fine-grained comparisons.

For practice, several rules follow. Do not rank instructors by mean score, especially not into rewarded and penalized groups; the misclassification rate makes this unfair to a large minority of staff. Treat small differences as no difference: if two instructors are separated by a few tenths of a point with overlapping confidence intervals, the correct conclusion is "indistinguishable," not "one is better." Aggregate before deciding: a pattern across several courses and terms is far more dependable than any single course mean, consistent with generalizability theory. And never let SET stand alone in tenure, promotion, or contract-renewal decisions — combine it with peer observation, teaching portfolios, and evidence of student learning, as the Esarey-Valdes recommendation and the ASA statement both insist.

Limitations and honest caveats

A rigorous reader will note that Esarey and Valdes is a simulation: its conclusions depend on the assumed correlation between SET and true quality, the assumed reliability, and the distributional choices. If SET were more strongly correlated with quality than modelled, misclassification would fall; if less (as the bias literature suggests in the real world), it would rise — so the simulation is arguably optimistic. The study quantifies error under best-case assumptions, not a measured real-world rate. Stark and Freishtat's strictures on averaging ordinal data are contested by some psychometricians who treat well-constructed multi-item scales as approximately interval. Gillmore et al.'s generalizability coefficients are specific to their institution and instrument. And none of this implies SET carries no information: a consistently low pattern across many courses, corroborated by other evidence, is meaningful. The defensible conclusion is narrow but firm — SET should not be used as a precise ranking device, and small or single-course differences should not drive high-stakes decisions.

How Koji incorporates this

Koji for Education is designed so that the statistics of uncertainty are visible in the product, not hidden behind a single decontextualized average.

  • Uncertainty-aware reporting. Koji's reports are designed to present distributions, response rates, and the imprecision of small samples rather than a bare mean, discouraging the "higher is better" heuristic that Boysen et al. documented. Where a difference is within the margin of error, the reporting frames it as indistinguishable.
  • Aggregation across cohorts and terms. Because Koji organizes feedback longitudinally, it supports the pattern-over-time, multi-section view that generalizability theory requires for a dependable instructor estimate — the opposite of ranking on one noisy course mean.
  • Richer signal than a number. Koji's AI-moderated conversational interviews and automatic thematic analysis produce qualitative, behaviourally specific evidence that complements the quantitative score, supporting the multi-measure approach Esarey and Valdes and the ASA both recommend instead of standalone ranking.
  • Structured triangulation. Koji is positioned as one input among peer review, portfolios, and learning evidence, with reporting that explicitly discourages using student feedback alone for tenure, promotion, or merit comparisons.

These are framed as safeguards, not guarantees: Koji cannot make a noisy comparison precise, but it is designed to mitigate misuse by surfacing uncertainty and steering institutions away from fine-grained ranking. Koji's core research platform at koji.so applies the same discipline to product and customer research, where treating a small difference in a satisfaction metric as a real effect is an equally common and costly error.

A worked intuition for the misclassification result

It helps to see why the result is not paradoxical. Imagine teaching quality and the evaluation score share a moderate correlation — each instructor's score is their true quality plus a chunk of noise from cohort, timing, room, and chance. Sort instructors by score and the bottom slice will be a mix: some genuinely weak teachers, and some good teachers who drew a bad term. The weaker the correlation and the smaller the sample behind each score, the larger that second group becomes. Esarey and Valdes simply quantified that group under generous assumptions and found it uncomfortably large — large enough that a "bottom-quintile" label is wrong about below-average teaching more than a quarter of the time.

The remedy is not a better single number but more numbers and more kinds of evidence. Aggregating an instructor's scores across several courses and terms shrinks the noise term, which is why generalizability theory treats multi-section data as the unit of dependable inference. Pairing the quantitative signal with peer observation and direct evidence of learning adds independent measures whose errors are unlikely to line up, so a consensus across them is far more trustworthy than any one alone. And reporting confidence intervals rather than bare means makes the central discipline visible: when two instructors' intervals overlap, the honest reading is "we cannot tell them apart" — precisely the conclusion the misclassification mathematics demands.

Related resources

References

Related articles

analysis-reporting

How Many Responses Do You Need for a Reliable Course Evaluation?

Reliability of a class-average rating rises steeply with the number of students who respond. Generalizability theory — from Marsh''s syntheses to Dzakadzie (2026) — shows raters matter more than items, and a class of 12 cannot be evaluated like a class of 120.

analysis-reporting

Can You Fairly Rank Instructors by Their Course-Evaluation Scores? What Esarey & Valdes (2020) Proved

Even if course evaluations were unbiased, reliable, and valid, ranking instructors by their scores would still misclassify many good teachers. A walk through the Esarey & Valdes (2020) simulation and what it means for how you report and use evaluation data.

analysis-reporting

Generalizability Theory and the Reliability of Student Ratings: Why One Class Is Not Enough

Reliability is not one number. Generalizability theory (Gillmore, Kane & Naccarato 1978; Marsh 1984) decomposes the variance in student ratings into student, teacher, course and occasion components — and shows that a single class can be reliable for the course yet a poor estimate of the teacher. What that means for fair evaluation.

analysis-reporting

Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees

The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.