When Combining Sections Reverses the Result: Simpson's Paradox in Course-Evaluation Data
Aggregating course-evaluation scores across sections, cohorts, or years can reverse the very conclusion you are trying to draw. What Simpson's paradox is, how it appears in evaluation data, and how to report so the reversal cannot bite you.
Koji Education Team
Product
In brief. Simpson''s paradox is the unsettling fact that a pattern present in every subgroup of your data can vanish or reverse when you pool the subgroups together. In course evaluation it is not a curiosity — it is a live risk every time you average across sections, cohorts, disciplines, or years. An instructor can out-score a colleague in every course they both teach and still show a lower overall mean; a new questionnaire can look worse in aggregate while being better for every student group. The fix is not more data — it is refusing to trust an aggregate mean without checking the subgroups underneath it.
The paradox, in one sentence
The direction of an association between two variables can reverse when a third, lurking variable is ignored during aggregation. Combine the groups and the story flips.
What the research says
The canonical demonstration is Bickel, Hammel and O''Connell (1975), published in Science. Examining 1973 graduate admissions at the University of California, Berkeley, the aggregate figures looked like clear sex discrimination: about 44% of male applicants were admitted versus about 35% of female applicants. But when the data were broken down by department, the bias largely disappeared — and where it existed, it slightly favoured women. The explanation was a lurking variable: women disproportionately applied to competitive departments with low admission rates for everyone, while men applied to easier-to-enter departments. Aggregating across departments destroyed the information needed to interpret the numbers. As the authors put it, pooling the data made the university appear to do something it was not doing.
The modern methodological treatment most useful to evaluators is Kievit, Frankenhuis, Waldorp and Borsboom (2013) in Frontiers in Psychology, "Simpson''s paradox in psychological science: a practical guide." Their central finding matters directly for course evaluation: the paradox is most likely to occur when drawing inferences across different levels of analysis — from populations to subgroups, or subgroups to individuals — which is exactly what a programme office does when it rolls student-level ratings up to a section mean, section means up to an instructor mean, and instructor means up to a departmental figure. Kievit and colleagues provide statistical markers that flag when a pooled association may be an artefact, and an R toolbox for detecting it. The paradox is, in their assessment, "pretty common" and "typically results in incorrect interpretations with potentially harmful consequences."
Simpson''s paradox is a special, dramatic case of a broader problem statisticians call the ecological fallacy and aggregation bias: relationships that hold at the group level need not hold at the individual level, and vice versa. It is also mathematically why multilevel models exist — they keep the levels separate instead of collapsing them into one misleading average.
How it shows up in course evaluation
Four everyday scenarios put every quality office at risk.
1. Comparing two instructors across a shared set of courses. Instructor A teaches mostly small, elective, final-year seminars — contexts that reliably attract higher ratings. Instructor B teaches mostly large, required, first-year quantitative courses — contexts that reliably attract lower ratings. If B is actually the stronger teacher within each course type, B can still post a lower overall mean simply because of the course mix. Ranking them on the pooled average inverts the truth. The lurking variable is course type; class size, electivity, and level are its usual carriers.
2. Evaluating a questionnaire redesign. Roll out a new form and compare aggregate satisfaction to last year''s. If the new form was disproportionately used by demanding programmes (say it launched in engineering first), the pooled comparison can show a decline even if satisfaction rose in every individual programme. The lurking variable is which programmes are in each pool.
3. Trends across years with a changing cohort mix. A programme''s overall rating can drift down year over year while every constituent module improves, simply because enrolment shifted toward harder-rated modules. This is a cousin of regression-to-the-mean traps and equally seductive.
4. Response-rate imbalance across groups. If satisfied students respond more in some sections and dissatisfied students respond more in others, the pooled mean can point the opposite way to the within-section reality — non-response acting as the lurking variable.
In every case the aggregate is not wrong arithmetically; it is answering a different question than the one you think you asked.
A concrete illustration makes the mechanism vivid. Suppose Instructor A and Instructor B each teach two course types. In small seminars, A averages 4.6 and B averages 4.7; in large required lectures, A averages 3.8 and B averages 3.9. B is higher in both types. But if A teaches mostly seminars (say 80% of their load) and B teaches mostly large lectures (80% of theirs), A''s overall mean is roughly 4.44 while B''s is roughly 4.06 — so a league table ranks A above B, exactly reversing the within-type reality. Nothing in the arithmetic is mistaken; the pooled means simply weight each instructor toward the course type they happen to teach. The lurking variable is teaching load composition, and no amount of decimal precision on the overall figure will reveal it.
Why it matters for course evaluation in practice
The stakes are highest exactly where evaluation data carries weight: personnel comparisons, programme rankings, and change decisions. A dean who ranks instructors on pooled means, a committee that judges a curriculum reform on an aggregate movement, or a dashboard that shows a "declining programme" can all be reading a Simpson reversal. Because the pooled number looks authoritative — bigger n, tidy single figure — it is more persuasive precisely when it is most likely to mislead.
The defensive discipline is simple to state and easy to skip: never report a comparison on a pooled mean without also reporting it within the relevant subgroups, and treat any disagreement between the two as a signal to investigate the lurking variable, not to pick whichever number tells the story you prefer.
Limitations and honest caveats
- Not every aggregate is paradoxical. Simpson''s paradox is a possibility to check for, not a reason to distrust all averages. Most aggregates are fine; the point is that you cannot know which without looking underneath.
- Which grouping is "correct" is a substantive judgement, not a statistical one. Disaggregating by department resolved the Berkeley case — but there is no algorithm that tells you which third variable to condition on. That requires domain knowledge about what plausibly drives ratings (course type, level, size, modality). Condition on the wrong variable and you can manufacture a paradox as easily as hide one.
- Over-disaggregation destroys precision. Slice finely enough and every cell has three responses and a meaningless mean. The tension between aggregation bias and small-sample noise is real; this is why shrinkage estimators such as empirical-Bayes exist.
- Detection is probabilistic. The statistical markers in Kievit et al. flag risk; they do not prove that a given reversal is causally meaningful. Judgement remains.
The honest summary: Simpson''s paradox does not make course-evaluation data untrustworthy. It makes unexamined aggregation untrustworthy.
How Koji incorporates this
Koji''s reporting is built on the assumption that the aggregate mean is the end of an analysis, not the start — and that the subgroups underneath it must stay visible.
- Level-aware reporting, not collapsed averages. Koji keeps student-, section-, instructor-, and programme-level results distinct rather than flattening them into a single number, mirroring the multilevel logic that prevents the paradox from forming in the first place.
- Like-for-like comparison controls. When comparing instructors or cohorts, Koji is designed to compare within relevant strata — course type, level, size, modality — so a comparison is not silently driven by course mix. This is a mitigation, not a guarantee: the platform surfaces the lurking variables it can see; the evaluator still chooses which comparisons are fair.
- Response-mix transparency. Because non-response is a classic lurking variable here, Koji reports who responded in each group, so a pooled figure is never read as if the groups were balanced.
- Open-text context to explain reversals. When a subgroup breakdown disagrees with the aggregate, Koji''s AI-moderated interviews and automatic thematic analysis provide the why — often the substantive lurking variable in students'' own words — so an analyst can resolve the paradox rather than merely detect it.
Framed carefully: no tool can decide for you which grouping is the right one — that is a judgement about causes. Koji is designed to keep the subgroups in front of you so the decision is made deliberately, not hidden inside a single mean. Koji''s core research platform at koji.so applies the same level-aware analysis to product and customer segments, where "the metric went up overall but down for every cohort" is the identical trap.
Related resources
- Students Are Nested in Courses: Why Multilevel Models Beat Raw Averages
- Stop Comparing Raw Averages: Empirical-Bayes Shrinkage
- The 4.2 vs 4.4 Trap: Why Small Differences in Means Are Usually Noise
- Measurement Invariance: Can You Compare Scores Across Groups at All?
- Top-Box vs Mean: Reporting Scores Without Throwing Away Information
- Regression to the Mean in Course Evaluations
References
- Bickel, P. J., Hammel, E. A., & O''Connell, J. W. (1975). Sex Bias in Graduate Admissions: Data from Berkeley. Science, 187(4175), 398–404. https://doi.org/10.1126/science.187.4175.398
- Kievit, R. A., Frankenhuis, W. E., Waldorp, L. J., & Borsboom, D. (2013). Simpson''s paradox in psychological science: a practical guide. Frontiers in Psychology, 4, 513. https://doi.org/10.3389/fpsyg.2013.00513
- Blyth, C. R. (1972). On Simpson''s Paradox and the Sure-Thing Principle. Journal of the American Statistical Association, 67(338), 364–366. https://doi.org/10.1080/01621459.1972.10482387
- Robinson, W. S. (1950). Ecological Correlations and the Behavior of Individuals. American Sociological Review, 15(3), 351–357. https://doi.org/10.2307/2087176
Related articles
Measurement Invariance: Can You Compare Course-Evaluation Scores Across Groups at All?
Comparing average ratings across departments, languages or online vs paper assumes the questionnaire means the same thing to every group. Measurement invariance is the test that assumption usually fails — and almost nobody runs it.
The 4.2 vs 4.4 Trap: Why Small Differences in Course-Evaluation Means Are Usually Noise
Faculty and administrators routinely read meaning into tiny gaps between course-evaluation means. Boysen (2015) and Boysen et al. (2014) show this happens even when confidence intervals say the difference is nothing — and that warnings barely help. How to report uncertainty honestly.
Why a Low-Scoring Course Usually 'Improves' Next Year Even If Nothing Changed: Regression to the Mean in Course Evaluations
Year-over-year movements in course-evaluation scores are mostly statistical noise regressing toward an instructor's true average. Here is why regression to the mean fools quality-assurance committees, and how to read evaluation trends responsibly.
Stop Comparing Raw Averages: Empirical-Bayes Shrinkage for Course-Evaluation Scores
Why raw course-evaluation means mislead — especially for small classes — and how empirical-Bayes shrinkage (partial pooling), the method Kane and Staiger applied to teacher value-added, produces fairer, more accurate instructor estimates.