The Ecological Fallacy in Course Evaluation: Why a Department Mean Tells You Almost Nothing About an Individual Lecturer
A course evaluation dataset is nested — ratings within students within instructors within programmes — and a number that is true at one level is routinely false at another. Confusing the levels is a formal reasoning error with a name, and it drives some of the worst decisions made with student feedback.
Koji Education Team
Product · July 14, 2026
The short answer: Course evaluation data is nested — individual ratings sit inside students, students inside courses, courses inside instructors, instructors inside programmes. A statistic computed at one level frequently does not hold at another. Treating a department's mean score as a verdict on a single lecturer, or one vivid comment as a signal about the whole programme, is a cross-level inference error — the ecological fallacy (aggregate to individual) or its mirror, the atomistic fallacy (individual to aggregate). Neither is a rounding problem you can wave away. Both are structural, and both are avoidable with a multilevel discipline that most legacy evaluation systems never impose.
A 75-year-old warning that quality assurance keeps ignoring
In 1950 the sociologist W. S. Robinson published one of the most-cited cautionary results in the social sciences. Using 1930 US census data, he showed that the correlation between the percentage of a state's population that was foreign-born and the percentage that was literate was strongly positive at the state level, yet at the individual level the relationship ran the other way — immigrants were, individually, less likely to be literate than the native-born. The ecological correlation was around r = 0.77; the individual correlation was near r = 0.20 and pointed differently once you looked inside the aggregates (Robinson, 1950, summarised in Wikipedia's ecological fallacy entry). Robinson's blunt conclusion: ecological correlations cannot be used as substitutes for individual correlations.
That is exactly the substitution higher education makes every semester. A programme director sees that Department A averages 4.3 out of 5 and Department B averages 3.9, and concludes that A's teachers are better than B's. A promotion panel sees that Dr X's module scored below the faculty mean and concludes Dr X is a weaker teacher. Both inferences push a group-level number down onto an individual — the ecological fallacy in its purest form. As a 2009 methodological review revisiting Robinson put it, the individualistic and ecologic fallacies are two sides of the same coin: you cannot read individuals from aggregates, and you cannot read aggregates from individuals, without data at both levels.
Where the variance actually lives
The deeper problem is not just that the levels differ — it is where the variance in evaluation scores comes from. When researchers decompose the total variance in student ratings into components (the course, the teacher, the student, and the student-by-teacher interaction), the results are unkind to anyone hoping the teacher mean is a clean signal. A 2017 variance-components study in Assessment & Evaluation in Higher Education found that the interaction of students and teachers was the strongest single source of variance — larger, in that analysis, than the stable teacher component. In plain terms: a large share of the score reflects the particular fit between a given student and a given teacher, not a fixed property of the teacher that generalises across cohorts.
That finding is fatal to naive aggregation. If much of the signal is student-by-teacher fit, then two things follow. First, the same instructor will legitimately score differently across cohorts, so a single term's mean is a noisy draw, not a fixed grade. Second, a department mean averages over wildly different courses, class sizes, and student populations, so it is not a common yardstick you can hold every member up against. The number is real; the inference stacked on top of it is not.
The atomistic fallacy: reading the programme from one comment
The mirror error is just as common and even easier to commit, because open-text comments are vivid. A single searing paragraph — 'this module was a shambles' — gets read aloud in a committee and quietly becomes 'the programme has a problem'. That is the atomistic (or individualistic) fallacy: generalising from an individual observation to the group. One articulate respondent is one data point at the student level; it licenses a hypothesis about the programme, never a conclusion. The research on reading open-text feedback is clear that salience is not representativeness — the comment you remember is a function of its emotional charge, not its frequency.
What good practice looks like
The fix is not to stop measuring; it is to keep the levels straight and to model them explicitly.
- Report with uncertainty, not just point estimates. A mean of 3.9 from twelve respondents and a mean of 3.9 from two hundred are not the same claim. Confidence intervals and response counts belong next to every number. See our note on why averaging Likert scores misleads.
- Use multilevel models when you compare. Partial pooling — shrinking small, noisy course means toward the relevant benchmark — is the statistically honest way to compare instructors who teach different-sized classes. We cover this in multilevel models and variance partitioning for course evaluation.
- Benchmark within the right reference frame. Comparing a compulsory 300-person first-year statistics course to a 12-person final-year elective is a category error dressed as a ranking; align comparisons to discipline norms.
- Never let one comment stand in for a cohort. Treat qualitative feedback as themes across many voices, not as a highlight reel. Watch too for Simpson's paradox, where an aggregate trend reverses inside subgroups — the ecological fallacy's most notorious special case.
But doesn't every organisation aggregate? Isn't this just statistical purism?
The strongest counterargument is pragmatic: institutions have to summarise. Deans cannot read ten thousand comments; accreditation panels want programme-level evidence; a quality dashboard needs a headline number. All true. The ecological fallacy is not an argument against aggregation — it is an argument against unlicensed inference from aggregation. Aggregate all you like for description ('the programme's median satisfaction rose this year'). The error begins the moment a group statistic is used to make an individual-level decision it cannot support — flagging a named lecturer, denying a promotion, ranking colleagues — without individual-level evidence and appropriate uncertainty.
A second objection: if teacher variance is small relative to student-by-teacher interaction, does that mean evaluations are worthless? No. It means they are worthless for the purpose they are most often abused for — fine-grained, high-stakes ranking of individuals — and genuinely useful for the purpose they are best suited to: formative, course-level improvement, where the person acting on the data is the person who taught the course and can interpret the fit. The level at which a number is trustworthy is the level at which it should be used.
How Koji keeps the levels honest
Koji for Education is built around the recognition that a single averaged number is the least defensible artefact in the whole process. Instead of collapsing everything to a Likert mean, Koji runs AI-moderated conversational interviews that probe why a student holds a view, so a low rating comes attached to its reason rather than floating free of context. Its automatic thematic analysis treats open-text feedback as patterns across many respondents — surfacing how often a concern actually recurs, which is the direct antidote to the atomistic fallacy of amplifying one loud voice. Reporting is explicitly programme- and institution-level as well as course-level, so the aggregate and the individual views stay distinct rather than being silently conflated. And because the same standardized AI moderation is applied to every respondent, the student-by-teacher fit noise that dominates traditional scores is at least measured consistently rather than confounded with inconsistent human survey administration.
None of this eliminates the multilevel structure of educational data — nothing can. What Koji does is stop pretending the structure isn't there. The same conversational interview engine powers Koji's main platform for user and customer research — the discipline of not confusing one articulate interview with a market-wide truth is identical.
The takeaway
Robinson's warning is 75 years old and still routinely ignored in faculty meetings. A course evaluation number is only as trustworthy as the level it was computed at. Keep the levels straight — report uncertainty, model the nesting, benchmark within frame, and read comments as themes — and student feedback becomes a genuine instrument for improvement. Flatten the levels into a single decontextualised mean, and you get confident decisions built on a formal reasoning error.