New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Evaluation bias9 min read

Is Your Course Evaluation Biased? It Depends on How You Analyse It — and That Is the Whole Problem

Two analysts can take the same course-evaluation dataset and reach opposite conclusions about bias, simply by making different defensible choices about controls, subsets, and outcomes. Specification-curve and multiverse analysis expose that fragility — and change what institutions should demand before acting on a number.

Koji Education Team

Product ·

Bottom line up front: Whether a course-evaluation dataset "shows bias" often depends less on the data than on the analyst's choices — which variables to control for, which students to include, how to define the outcome. Each of those choices is individually defensible, but together they form hundreds of plausible analyses that can point in different directions. Specification-curve and multiverse analysis make this fragility visible instead of hiding it behind a single number. For institutions making personnel and curriculum decisions, the lesson is blunt: demand robustness across reasonable analyses, not a lone point estimate.

The garden of forking paths in your evaluation data

Suppose you want to know whether female lecturers in your faculty receive lower evaluations than male lecturers after accounting for everything else. You sit down with the data and immediately face a cascade of choices. Do you control for class size? For course level, discipline, whether the course is compulsory? Do you include the lecturers with only a handful of responses, or set a minimum? Do you model the mean rating, the proportion of top-box scores, or the full ordinal distribution? Do you treat students as independent or nest them within courses?

Every one of these is a reasonable, arguable decision. And every combination produces a slightly different analysis. A dozen binary choices generate thousands of distinct, individually-defensible analytic paths — Gelman and Loken's "garden of forking paths." The unsettling consequence: an analyst who wants to find bias can usually find a defensible path to it, and an analyst who wants to dismiss it can usually find a defensible path away from it. Neither has to cheat. They just have to stop at the specification that agrees with them.

This is not a hypothetical worry for course evaluation specifically. It is why the published literature can simultaneously contain robust natural-experiment evidence of gender bias — Boring, Ottoboni and Stark's analysis of 23,001 evaluations found student ratings were more strongly related to an instructor's perceived gender and to students' grade expectations than to learning measured by anonymously graded final exams (Boring, Ottoboni & Stark, 2016, ScienceOpen) — and other studies reporting no significant effect. Some of that disagreement is genuinely different contexts. Some of it is different forks in the same kind of garden.

Specification curve and the multiverse: show all the paths

The methodological response, developed in psychology and economics precisely to tame this problem, is to stop reporting one analysis and start reporting all reasonable analyses.

Specification-curve analysis (Simonsohn, Simmons & Nelson, 2020, Nature Human Behaviour) proceeds in three steps: (1) enumerate the set of theoretically justified, statistically valid, non-redundant specifications; (2) plot the estimate from every specification, sorted, so readers can see how the result moves as choices change; (3) conduct joint inference across the whole set rather than cherry-picking one. The closely related multiverse analysis (Steegen, Tuerlinckx, Gelman & Vanpaemel, 2016, Perspectives on Psychological Science) does the same for data-processing choices, computing the result under every reasonable combination of decisions.

Applied to a bias question, the output is not "bias = yes/no." It is a picture: across 500 defensible ways of analysing this dataset, the estimated gender effect is negative in 470 of them, statistically significant in 300, and never meaningfully positive. That is a far more honest and decision-useful summary than a single regression coefficient with a p-value — and it is very hard to game, because the analyst no longer chooses the one path that gets reported.

Even a robust, valid instrument can still be unfair

There is a second, deeper reason institutions should stop trusting single numbers, and it survives even if your instrument is psychometrically excellent. Esarey and Valdes modelled evaluations under deliberately ideal conditions — assuming the ratings were unbiased, reliable and valid — and still found an alarmingly high error rate when the scores were used to rank instructors: under realistic assumptions, the process identified the wrong instructor as the superior teacher a substantial share of the time (Esarey & Valdes, 2020, Assessment & Evaluation in Higher Education). The reason is statistical, not psychological: small samples per instructor and real measurement noise make fine-grained rankings unreliable no matter how good the underlying instrument.

Combine the two findings and the message is stark. The analytic path you pick determines whether you see bias, and even a clean instrument produces unreliable rankings. A course-evaluation mean, reported alone and acted on directly, is standing on far less solid ground than the two-decimal-place precision suggests. This compounds the more basic validity problem that student ratings correlate near zero with actual learning once prior ability is controlled (Uttl, White & Gonzalez, 2017) — a point we develop in our overview of reliability versus validity.

"But doesn't this just make everything unknowable?"

This is the fair objection: if any conclusion can be reached by some defensible path, does robustness analysis simply license nihilism — throw up your hands, trust nothing, act on gut? No, and it is important to say why.

Multiverse and specification-curve analyses are the opposite of nihilism. They discipline uncertainty rather than surrendering to it. When the great majority of reasonable specifications agree, you have unusually strong grounds for a conclusion — stronger than any single analysis could give, because you have shown the result is not an artefact of one lucky choice. When specifications disagree, that is genuine information: it tells you the effect is fragile and that a high-stakes decision resting on it is unsafe. Either way you know more, not less. What robustness analysis forbids is the false confidence of a lone estimate — and for personnel decisions, false confidence is the expensive failure mode. It also complements the routine multiple-comparisons discipline every evaluation dashboard needs, which we cover in false flags on the dashboard.

What this means in practice

You do not need a research team to apply the underlying discipline:

  • Never act on a single coefficient or mean for a high-stakes decision. Ask what happens under the two or three most obviously reasonable alternative choices before drawing a conclusion.
  • Pre-specify the analysis for recurring questions (annual bias monitoring, flagging modules) so choices are made before the results are seen, closing the garden of forking paths.
  • Report the spread, not just the point. Confidence intervals, distributions across specifications, and honest "this is fragile" flags belong in the report that reaches a promotion panel.
  • Treat convergence as the bar for action. Change something when the conclusion holds across reasonable analyses — not when one specification happens to reach significance.

Where Koji fits

Robustness analysis lives or dies on the quality of the underlying evidence, and this is where the design of the instrument matters. Where a bare Likert mean gives you a single fragile number to slice, Koji for Education gathers richer, multi-dimensional evidence — AI-moderated conversational interviews that capture the reasoning behind a rating, six structured question types, and automatic thematic analysis of open text — so a conclusion can be corroborated across qualitative themes and quantitative items rather than resting on one coefficient. Its quality scoring and standardised, bias-aware moderation reduce the measurement noise that makes fine-grained rankings unreliable in the first place, and programme- and institution-level reporting is built to show distributions and context rather than a lone decimal, nudging committees toward the convergence-based judgements this literature demands. Koji does not claim to eliminate analytic uncertainty — no platform can repeal the garden of forking paths — but better, richer evidence is what makes an honest robustness check possible.

Teams doing broader institutional research can apply the same evidence-rich interview engine through the main Koji platform, keeping one defensible methodology across course evaluation and wider studies.

Stop betting personnel decisions on a single fragile number. See how Koji for Education gathers evidence robust enough to survive an honest second look.