The Department Average Is Not the Student: The Ecological Fallacy in Course Evaluation
A correlation that holds between department averages need not hold — and can even reverse — for individual students. The ecological fallacy is inferring individual-level relationships from group-level data. Here is how it distorts course-evaluation analysis and how to reason at the level your decision is actually made.
Koji Education Team
Product
A department that averages 4.2 is not a department where the typical student rated the course 4.2, and a pattern you see between departmental averages need not exist — or may run the opposite way — inside any single classroom. The ecological fallacy is the error of reading individual-level conclusions off group-level data. In course evaluation it is everywhere: in benchmarking a student against a cohort mean, in correlating faculty-level aggregates to explain what makes an individual student satisfied, and in cross-national league tables. The fix is not to abandon aggregates; it is to match the level of your evidence to the level of your claim.
Answer box (BLUF): If you correlate department-level averages — "faculties with higher response rates have higher scores" — and then conclude something about individuals — "students who respond rate courses higher" — you have committed the ecological fallacy, and your conclusion may be false or even reversed. Group-level relationships and individual-level relationships are different quantities. Before acting, ask: is my decision about a group (a programme, a cohort) or an individual (a student, an instructor)? Analyse at that level, using individual records when the decision is individual, and never assume a correlation transfers across levels.
What the research says
The foundational demonstration is Robinson's (1950) "Ecological Correlations and the Behavior of Individuals" in the American Sociological Review — one of the most cited methodological papers in the social sciences. Robinson examined U.S. Census data on nativity and literacy. Across the 48 states, the ecological correlation between the percentage of foreign-born residents and the literacy rate was strongly positive (about +0.53): states with more immigrants had higher average literacy. Yet at the individual level the correlation was slightly negative (about −0.11): foreign-born individuals were, on average, less literate than the native-born. The sign flipped entirely. Immigrants had settled in states that already had high native literacy, so the state-level correlation reflected where people lived, not the relationship within people. Robinson's conclusion was blunt: ecological correlations cannot be used as substitutes for individual correlations.
The error has been rediscovered across fields. In epidemiology, Piantadosi, Byar, and Green (1988) showed formally how the ecological fallacy inflates and distorts exposure–disease relationships estimated from aggregate data. Freedman (2001) catalogued how ecological inference fails and why the assumptions needed to reverse it are rarely met. Gary King's (1997) A Solution to the Ecological Inference Problem proposed statistical machinery to recover individual relationships from aggregates — but even King is explicit that it works only under strong, checkable assumptions, not as a routine shortcut.
The constructive counterpart is multilevel (hierarchical) modelling, developed for education by Raudenbush and Bryk (2002). Because students are nested within courses within programmes, the right tool models both levels at once, keeping the within-group and between-group relationships distinct instead of collapsing them into a single misleading number.
Two neighbours are worth naming to avoid confusion. The ecological fallacy's mirror image is the atomistic fallacy — assuming individual-level relationships automatically hold at the group level. And it is not the same as Simpson's paradox, which is a specific reversal that appears when you condition on (or ignore) a grouping variable; the ecological fallacy is the broader problem of cross-level inference. (We cover Simpson's paradox separately.)
Why it matters for course evaluation in practice
University quality processes swim in aggregates — faculty dashboards, programme averages, national survey league tables — and constantly draw individual conclusions from them. Four common slips:
- Reading a mean as a person. A 4.2 average can come from unanimous 4.2s or from a bimodal split of delighted and alienated students. Deciding "students are broadly satisfied" from the mean alone erases the very minority a quality process exists to hear.
- Explaining individuals with faculty-level correlations. "Departments that use more continuous assessment score higher, so continuous assessment makes students happier" mixes levels. The department-level association could be driven by which disciplines favour continuous assessment, not by any student-level effect — and could reverse inside classrooms.
- Benchmarking individuals against aggregates. Flagging every instructor below the departmental mean treats a group statistic as an individual standard. Half of any group sits below its own mean by construction; the aggregate says nothing about whether a specific instructor's students were poorly served.
- Cross-national league tables. Correlating country-level response rates with country-level satisfaction, then inferring how an individual European student behaves, compounds the ecological fallacy with cross-cultural response-style differences (a separate documented problem).
The stakes are concrete: promotion cases, programme closures, and resource decisions are individual- or programme-level actions frequently justified with the wrong level of evidence. Matching level to claim is not statistical pedantry; it is the difference between a defensible decision and an arbitrary one.
Limitations and honest caveats
It would be an overcorrection to conclude that aggregate data is worthless — and a critical reader will press exactly there:
- Group-level data is valid for group-level questions. If your decision is about a programme ("does this programme meet its threshold?"), a programme-level statistic is the right evidence. The fallacy is only in cross-level inference, not in aggregates as such. Refusing to use group data when the question is a group question is the atomistic fallacy in reverse.
- Aggregation buys reliability. Averaging many noisy student responses yields a more stable estimate of a course's central tendency. The trade-off is real: aggregates are more reliable and less individually informative at the same time.
- The modifiable areal unit problem. How you draw the groups — by module, by instructor, by department — changes the ecological correlation, sometimes dramatically. There is no single "correct" aggregation, so ecological results are partly an artefact of the grouping choice.
- Individual data has its own limits. Individual-level analysis is not automatically superior; it can be underpowered, and it raises anonymity risks that aggregation is often introduced to protect. Small subgroups can re-identify students.
- Cross-level recovery is possible but fragile. Ecological-inference methods (King 1997) and multilevel models can, under assumptions, connect the levels — but those assumptions must be argued, not presumed.
The honest rule is narrow and durable: use the level of data that matches the level of your decision, and when you must move across levels, treat it as an inference requiring justification, not a free translation.
How Koji incorporates this
Ecological fallacies usually enter when a system only ever stored the aggregate, leaving analysts no choice but to reason from it. Koji is built to keep the individual level available so you can analyse at the level your decision actually occupies.
- Individual records are retained, not pre-averaged. Every conversational interview and structured response is preserved at the student level within its course and programme, so when a decision is about individuals you can analyse individuals — and when it is about a programme you can aggregate deliberately, knowing what you collapsed.
- Distributions and segments, not just means. Koji's reporting is designed to show the shape of responses — spread, bimodality, and segment breakdowns — so a mean is never mistaken for a description of the typical student, and a hidden split surfaces instead of averaging away.
- Structure that supports multilevel reasoning. Because responses stay nested (student → course → programme → faculty), the data is ready for the hierarchical analysis Raudenbush and Bryk prescribe, keeping within-group and between-group relationships distinct rather than conflating them.
- Heterogeneity the average hides. Koji's AI-moderated interviews probe why individual students answered as they did, exposing the within-cohort variation that a single ecological number erases — the delighted-and-alienated split behind a comfortable mean.
- Level-aware framing. Koji's bias-aware reporting is oriented toward flagging when a comparison mixes levels — for example, holding an individual instructor against a group aggregate — so the fallacy is caught before it becomes a decision.
As always, Koji is designed to support correct cross-level reasoning, not to certify it; the analyst still owns the inference. The same individual-level engine powers Koji's core research platform at koji.so, where segmenting customers rather than trusting a single NPS average is the same discipline applied to product research.
Frequently asked questions
What is the ecological fallacy in one sentence? It is the error of drawing conclusions about individuals from data measured on groups — assuming that a relationship between group averages holds, unchanged, for the individuals inside those groups.
How is it different from Simpson's paradox? Simpson's paradox is a specific reversal that appears when you aggregate over or condition on a grouping variable within one dataset. The ecological fallacy is the broader problem of inferring individual-level relationships from group-level data; Simpson's paradox is one way it can bite, but not the only one.
Does this mean I should never use department or programme averages? No. Aggregates are exactly the right evidence for aggregate questions — whether a programme meets a threshold, for instance. The fallacy is only in inferring individual behaviour from group statistics. Refusing to use group data for group questions is the opposite error, the atomistic fallacy.
Why did the sign flip in Robinson's data? Foreign-born residents had settled disproportionately in states that already had high native-born literacy. So at the state level, more immigrants coincided with higher literacy, even though immigrants themselves were individually less literate. The state correlation captured where people lived, not a relationship within people.
What is the right way to analyse student data across courses and programmes? Use multilevel (hierarchical) models that keep students nested within courses within programmes. These estimate within-group and between-group relationships separately, avoiding the collapse into a single cross-level number that causes the fallacy.
Can I ever recover individual relationships from aggregate course data? Sometimes, using ecological-inference methods, but only under strong assumptions that must be argued and checked. It is far safer to retain and analyse individual-level records when the decision is about individuals, which is why preserving that granularity matters.
Related resources
- Simpson's Paradox in Course-Evaluation Data
- Collider Bias and Berkson's Paradox in Course-Evaluation Selection
- Base-Rate Neglect in Reading Course Evaluations
- Denominator Neglect: Why Raw Counts and Percentages Mislead
- National Student Surveys Are Not Course Evaluations
- Funnel Plots for Fair Course-Evaluation Comparison
References
- Robinson, W. S. (1950). Ecological correlations and the behavior of individuals. American Sociological Review, 15(3), 351–357. https://doi.org/10.2307/2087176
- Piantadosi, S., Byar, D. P., & Green, S. B. (1988). The ecological fallacy. American Journal of Epidemiology, 127(5), 893–904. https://doi.org/10.1093/oxfordjournals.aje.a114892
- Freedman, D. A. (2001). Ecological inference and the ecological fallacy. International Encyclopedia of the Social & Behavioral Sciences, 6, 4027–4030. https://www.stat.berkeley.edu/~census/549.pdf
- King, G. (1997). A Solution to the Ecological Inference Problem. Princeton University Press. https://doi.org/10.1515/9781400849208
- Raudenbush, S. W., & Bryk, A. S. (2002). Hierarchical Linear Models: Applications and Data Analysis Methods (2nd ed.). Sage.
Related articles
When Combining Sections Reverses the Result: Simpson's Paradox in Course-Evaluation Data
Aggregating course-evaluation scores across sections, cohorts, or years can reverse the very conclusion you are trying to draw. What Simpson's paradox is, how it appears in evaluation data, and how to report so the reversal cannot bite you.
National Student Surveys Are Not Course Evaluations: What NSS and Studiebarometeret Can and Cannot Tell You
National student surveys sit at the wrong level of analysis to diagnose a course. Cheng and Marsh showed that most apparent difference between UK universities is not reliable variance at all — here is how to use national data alongside your own instrument instead of in place of it.
Which Instructor Scores Are Genuinely Unusual? Funnel Plots for Fair Course-Evaluation Comparison
Ranking instructors by mean evaluation score turns sampling noise into a league table. Funnel plots — Spiegelhalter's method for institutional comparison — show which scores are genuinely unusual and which are just small-sample wobble.
One Vivid Comment Is Not a Pattern: Base-Rate Neglect in Reading Course Evaluations
A single scathing open-text comment can outweigh forty neutral ones in a reviewer's mind. Bar-Hillel and Kahneman & Tversky showed why: people ignore base rates in favour of vivid, individuating detail. Here is how base-rate neglect distorts course-evaluation review — and how to design against it.