Was It the Teacher or the Situation? The Fundamental Attribution Error in Reading Course Evaluations
Committees read a low evaluation score as evidence of a weak teacher. Social-psychology research on the fundamental attribution error shows why that inference is systematically biased — and how to read scores in situational context.
Koji Education Team
Product
In brief
When a committee reads a low course-evaluation score, the near-automatic inference is "this is a weak teacher." That is a dispositional judgment — it locates the cause inside the person. Decades of social-psychology research on the fundamental attribution error (Ross, 1977) and the correspondence bias (Gilbert & Malone, 1995) show that observers systematically under-weight the situation and over-attribute behaviour and outcomes to stable personal traits. In course evaluation this bias is amplified by a large, well-replicated literature showing that scores are pushed around by class size, course difficulty, discipline, required-versus-elective status, and delivery mode — situational forces the instructor did not choose. Reading a score without its situational context is not a neutral act; it is a bias waiting to fire.
What the research says
The fundamental attribution error (FAE) was named by Lee Ross in his 1977 chapter "The intuitive psychologist and his shortcomings." Ross defined it as the tendency of observers "to underestimate the impact of situational factors and to overestimate the role of dispositional factors in controlling behaviour." The empirical anchor came a decade earlier from Jones and Harris (1967), whose "attribution of attitudes" experiments found that participants inferred a person genuinely held a pro-Castro attitude even when told the person had been assigned to write a pro-Castro essay with no choice. The situational constraint was explicit, yet observers still reached for a dispositional explanation.
Gilbert and Malone (1995), in their authoritative Psychological Bulletin review "The correspondence bias," synthesised the field and proposed four mechanisms that produce the effect: (1) observers are often unaware of the situational forces acting on the person; (2) they hold unrealistic expectations about how behaviour should vary with the situation; (3) they inflate their categorisation of the behaviour itself; and (4) — crucially — even when they know the situation matters, they fail to adequately correct their initial dispositional inference. That last mechanism is the dangerous one for evaluation review: telling a committee "remember, this was a hard required course" is not enough, because the correction people apply is reliably too small.
Why does this matter specifically for course evaluation? Because the situational forces are real and measured. A substantial body of work documents that Student Evaluation of Teaching (SET) scores co-vary with factors outside the instructor''s pedagogical quality: larger classes tend to receive lower ratings; more difficult and heavier-workload courses are rated lower; quantitative and "hard" disciplines score below humanities; required courses score below electives; and grading leniency correlates positively with ratings. Each of these is, from the reader''s chair, a situational determinant of the number on the page. The FAE predicts that reviewers will nonetheless read a low number as a statement about the teacher.
Why it matters for course evaluation in practice
Course evaluations are consequential. They feed into contract renewal, promotion and tenure, teaching-award shortlists, and programme-level quality assurance. The FAE turns each of those decision points into a place where a situational artefact can be misread as a personal failing.
Consider three concrete failure modes:
- The hard-course penalty. An instructor who teaches the notoriously difficult second-year statistics course that everyone must pass receives a mean of 3.6, while a colleague teaching a popular final-year elective receives 4.4. A committee reading the two numbers side by side, with no context, concludes the first instructor is the weaker teacher. The correspondence bias makes this conclusion feel like an observation rather than an inference.
- The large-lecture penalty. The person assigned the 400-student introductory lecture is structurally disadvantaged relative to the colleague running 18-person seminars, yet both scores land in the same spreadsheet column.
- The new-preparation penalty. An instructor teaching a course for the first time, or one handed a course at short notice after a colleague''s departure, faces situational headwinds — an unfamiliar syllabus, no accumulated materials — that the score cannot distinguish from teaching skill.
The through-line is that the reader supplies the causal story, and the default story is dispositional. This is why interpretation guidance for SET (for example the responsible-use recommendations discussed in our companion pieces) repeatedly stresses that scores must never be read as context-free measures of instructor merit. The FAE explains why the warning is needed and why it is so easily ignored: the bias operates fast, feels like perception, and resists verbal correction.
Limitations and honest caveats
Intellectual honesty requires several caveats.
First, the FAE is one of the more contested constructs in social psychology. Critics (notably in the actor–observer literature) have argued that the effect is smaller and more moderated than the classic demonstrations suggest, and a well-known meta-analysis found the actor–observer asymmetry to be near zero on average. The point for evaluation review does not depend on the FAE being a universal law; it depends only on the weaker, well-supported claim that observers frequently under-correct for known situational constraints — which Gilbert and Malone''s mechanism (4) captures directly.
Second, "the situation caused the low score" can itself be over-applied. Not every low score is a situational artefact; some reflect genuine, addressable teaching problems. The corrective for the FAE is not to explain away all bad scores as circumstance — that would be the equal-and-opposite error — but to bring the situation into view so that person and situation can be weighed against each other rather than the person being assumed by default.
Third, the confound literature it leans on is correlational and heterogeneous. Effect sizes for class size, difficulty, and discipline vary across studies and institutions, and some apparent confounds partly reflect real differences in teaching effort or effectiveness. Situational context should therefore inform interpretation, not mechanically "adjust away" differences — the debate over statistically adjusting SET scores (see our IDEA-approach discussion) is exactly this tension.
Fourth, cultural generalisability is uncertain: much of the attribution research was conducted with Western, individualist samples, and the strength of dispositional bias may differ across cultures — relevant for pan-European institutions with international committees.
How Koji incorporates this
Koji is designed to put the situation in front of the reader so that the dispositional default has to compete with visible context, and to reduce the reader''s reliance on a single decontextualised number.
- Context-anchored reporting. Koji surfaces the situational metadata alongside every score — class size, course level, required-versus-elective status, discipline, and delivery mode — rather than presenting a bare mean. The design goal is Gilbert and Malone''s mechanism (1): reduce unawareness of the situation, because a committee that literally sees "core required module, 320 enrolled, first delivery" beside a 3.6 is less able to slide into a pure trait attribution.
- Bias-aware framing, not just a warning label. Because research shows verbal warnings under-correct, Koji is designed to present benchmarked context (how comparable courses of similar size and level score) so the comparison is built into the report rather than left to the reader''s memory. This reframes "3.6 is low" into "3.6 relative to comparable required first-deliveries" — a situational reference class instead of an absolute.
- Attribution-aware thematic analysis of open text. Koji''s automatic thematic analysis of open-ended comments is designed to separate feedback about situational features of the course (pace of a mandated syllabus, room, timetable, workload set by the programme) from feedback about the instructor''s own choices and conduct. Distinguishing "the situation" from "the person" in students'' own words gives a committee dispositional and situational signal side by side, rather than collapsing both into one rating.
- Probing beyond the number. Koji''s AI-moderated conversational interviews ask follow-up questions ("what specifically made the workload feel heavy — the material, or how it was taught?") that help attribute a complaint to its actual cause. A Likert number cannot tell a reader whether difficulty was intrinsic to a hard subject or created by poor scaffolding; a probed response can.
- Triangulation across cohorts and instruments. By collecting mid-cycle and end-of-cycle data and tracking the same instructor across different course situations, Koji helps a committee see whether a low score travels with the person (dispositional signal) or with the situation (a particular course), which is precisely the evidence needed to overcome a correspondence-bias default.
None of this eliminates the fundamental attribution error — no reporting tool can rewrite how human cognition assigns causes. The claim is narrower and honest: by making the situation unavoidable in the report and by giving reviewers attributed, probed evidence rather than a lone number, Koji is designed to mitigate the conditions under which the bias does the most damage.
Koji''s core research platform at koji.so applies the same AI-moderated interview engine to product and customer research, where the mirror-image bias — attributing a churned customer to "they just didn''t get it" rather than to a situational friction in onboarding — is equally costly.
Related resources
- Anchoring bias in evaluation review
- Contrast effects and narrow bracketing when reviewing reports
- The dilution effect in evaluation reports
- Class size effects on student evaluations
- Course difficulty and workload confounds
- Interpreting and reporting student ratings responsibly
References
- Jones, E. E., & Harris, V. A. (1967). The attribution of attitudes. Journal of Experimental Social Psychology, 3(1), 1–24. https://doi.org/10.1016/0022-1031(67)90034-0
- Ross, L. (1977). The intuitive psychologist and his shortcomings: Distortions in the attribution process. Advances in Experimental Social Psychology, 10, 173–220. https://doi.org/10.1016/S0065-2601(08)60357-3
- Gilbert, D. T., & Malone, P. S. (1995). The correspondence bias. Psychological Bulletin, 117(1), 21–38. https://doi.org/10.1037/0033-2909.117.1.21
- Spooren, P., Brockx, B., & Mortelmans, D. (2013). On the validity of student evaluation of teaching: The state of the art. Review of Educational Research, 83(4), 598–642. https://doi.org/10.3102/0034654313496870
Related articles
Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees
The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.
The Course Reviewed Straight After a Brilliant One Looks Worse: Contrast Effects and Narrow Bracketing in Evaluation Review
Course-evaluation bias research focuses on the student filling in the form. The evidence on sequential judgement says the committee reading twenty reports in an afternoon is biased too — by what it read immediately before.
Does the Number Anchor the Committee Before It Reads a Word? Anchoring Bias in Evaluation Review
When a review panel sees an instructor's 3.8 mean before reading the comments, that number quietly pulls every later judgement toward it. What the anchoring literature says, and how to sequence evaluation review to resist it.
How Padding a Report With Extra Data Weakens a Strong Signal: The Dilution Effect in Evaluation Review
A classic finding (Nisbett, Zukier & Lemley, 1981) shows that adding irrelevant, non-diagnostic information makes judgements less extreme. For evaluation committees, a padded report can dilute a genuinely strong signal.