New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Methodology9 min read

Simpson's Paradox in Course Evaluation: How a Department Average Can Lie About Every Programme It Contains

Aggregate course-evaluation means can reverse the direction of a difference that holds within every subgroup. Here is why Simpson's paradox is a routine feature of evaluation data — and how to report scores so it cannot mislead your quality decisions.

Koji for Education

Research & Editorial Team · June 22, 2026

Bottom line up front: When you average course-evaluation scores across groups — disciplines, cohorts, delivery modes, the gender of the instructor — the direction of a difference can reverse. A department can post a higher overall mean than its neighbour while scoring lower in every single programme inside it. This is Simpson's paradox, and in course evaluation it is not an exotic edge case. It is a predictable consequence of the unbalanced, observational data that student feedback produces. If your quality-assurance dashboards rank units on raw aggregate means, you are at risk of acting on conclusions the disaggregated data flatly contradict.

The paradox in one famous example

The cleanest illustration comes not from education but from admissions. In 1973, the University of California, Berkeley admitted roughly 44% of male graduate applicants and 35% of female applicants. The aggregate gap looked like discrimination against women. When the statistician Peter Bickel and colleagues disaggregated by department, the apparent bias dissolved — and in four of six large departments, women were admitted at higher rates than men. The explanation, published in Science in 1975, was that women disproportionately applied to the most competitive departments, which had low admission rates for everyone (Bickel, Hammel & O'Connell, 1975). The department someone applied to was a lurking variable correlated with both applicant gender and admission odds. Pool over it, and the relationship flips.

That reversal is Simpson's paradox: an association present in aggregated data disappears or inverts when the data are split by a confounding subgroup. It is a mathematical fact about weighted averages, not a statistical anomaly you can hope to avoid by collecting more data.

Why course evaluation is unusually prone to it

Course-evaluation data has exactly the structure that breeds the paradox: groups of very different sizes, compared on a shared scale, where group membership is correlated with something that also moves the score.

  • Disciplines rate differently. It is well established that quantitative and STEM courses receive systematically lower ratings than humanities and arts courses, independent of teaching quality. A department that happens to teach more service mathematics will look worse in aggregate even if every one of its modules out-rates the comparator.
  • Class size is correlated with ratings. Larger classes tend to draw lower scores. A unit weighted toward big first-year lectures carries a structural disadvantage in any pooled mean.
  • Bias is not uniform. The natural experiment by Boring, Ottoboni and Stark (2016), drawing on 23,001 evaluations at a French university, found that gender bias in student evaluations varies by discipline and by student gender. When the size and direction of a bias depend on subgroup, pooling across subgroups can manufacture a difference that exists in none of them — or erase one that exists in all of them.
  • Response rates differ. Online evaluations, optional modules, and certain cohorts return very different participation. Each subgroup mean is built on a different base, so the weights in the pooled average are not the weights you would choose deliberately.

Put concretely: imagine Programme A and Programme B each teach the same two course types, "small seminar" and "large lecture." In both course types, A out-scores B. But A teaches mostly large lectures (which score low for everyone) and B teaches mostly seminars (which score high for everyone). Pool the two course types together and B's overall mean can exceed A's — even though A is better in every comparison that actually holds the course type constant. A dean ranking departments on the aggregate would reward the worse performer.

The deeper point: which average is "true"?

It is tempting to declare the disaggregated picture correct and the aggregate wrong. That is not quite right, and the nuance matters for a methodology-literate audience. Neither number is false; they answer different questions. The aggregate answers "what did the typical student in this unit experience?" The disaggregated comparison answers "holding course type constant, which unit rates higher?" The paradox is dangerous only when you let an aggregate answer a question it cannot — chiefly, causal questions about teaching quality, fairness, or who deserves a personnel consequence.

Resolving the paradox is therefore not a statistical trick but a question of causal structure: you must condition on the confounders that are not part of what you are trying to measure (discipline, class size, mode) while being careful not to condition away the thing you do care about. This is why the honest answer to "which department is better?" is usually "compared how, and adjusting for what?"

But doesn't disaggregating just let you cherry-pick?

The strongest objection runs the other way. If splitting the data can reverse a conclusion, can't a motivated administrator slice until they get the answer they wanted? Yes — and this is the real risk of naive subgroup analysis. The discipline of avoiding Simpson's paradox is not "always disaggregate." It is "decide, in advance and on substantive grounds, which variables are genuine confounders, and adjust for those." Discipline, class size, and delivery mode have a clear theoretical claim to be confounders of a teaching-quality comparison. Inventing post-hoc subgroups until a difference appears or vanishes is the abuse the paradox warns against, not the cure.

A second objection: with enough data, won't this wash out? No. Simpson's paradox is not a small-sample artefact like the ones discussed in our note on why small classes produce unreliable scores. It can be arbitrarily large with arbitrarily large samples, because it is driven by the correlation structure of the groups, not by noise. More respondents make each subgroup mean more precise; they do nothing to remove the confounding that causes the reversal.

What honest reporting looks like

The defensible response is a reporting discipline, not a single magic statistic:

  1. Never compare raw aggregate means across heterogeneous units for any decision that matters. Report comparisons within course type, level, and mode.
  2. Show the disaggregation by default, not on request. If a unit's aggregate and its subgroup pattern disagree, that disagreement is the finding.
  3. Use models that pool partially. Multilevel models, discussed in our piece on variance partitioning in course evaluation, estimate course-level effects while accounting for the nesting of courses within programmes — a principled middle path between fully pooled (paradox-prone) and fully separate (noisy) estimates.
  4. State the confounders you adjusted for, so the comparison is auditable. This is simply the ENQA-aligned principle that evidence used in quality decisions should be transparent and fit for purpose.

How Koji reduces the exposure

Most of the damage from Simpson's paradox is done at the reporting layer, where a platform collapses rich, structured responses into a single headline mean and invites cross-unit ranking. Koji for Education is built to resist that collapse. Because every interview captures structured metadata — course, level, delivery mode, question type — alongside the response, the data arrives already disaggregable; programme- and institution-level reporting can present within-stratum comparisons rather than a single pooled number. Koji's question design spans six structured types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no), and its automatic thematic analysis of open-text feedback means a unit's score is never the only signal you have — you can see why a subgroup diverges, not merely that it does. None of this eliminates the paradox; nothing can, because it is a property of weighted averages. What good tooling does is make the disaggregated truth the path of least resistance instead of an analysis someone has to remember to run.

Teams running general user or customer research hit the identical trap when they average satisfaction across segments; the main Koji platform uses the same AI interview engine and the same disaggregate-by-default philosophy.

The takeaway

Simpson's paradox guarantees that somewhere in your institution, an aggregate course-evaluation comparison points the opposite way to the truth within its subgroups. You cannot collect your way out of it. You can only report your way around it — by deciding in advance what to hold constant, showing the disaggregation by default, and refusing to let a single pooled mean carry the weight of a fairness or personnel decision. Treat the headline average as a question, never an answer.


Koji for Education replaces the static, average-a-Likert-score report with AI-moderated conversational interviews and structured, disaggregable reporting at programme and institution level. See how Koji approaches course evaluation.