The Fundamental Attribution Error in Course Evaluations: When Students Rate the Situation, Not the Teacher
A low course evaluation score feels like a verdict on the lecturer. But much of the variance in student ratings is driven by situational factors the instructor never controlled. Reading those numbers as character references is the fundamental attribution error, scaled to an institution.
Koji for Education
Research & Editorial Team · June 24, 2026
Bottom line up front: When a course scores 3.4 out of 5, the reflex in a promotion panel or a quality committee is to read it as a verdict on the lecturer. Decades of evidence say otherwise: a substantial share of the variance in student evaluation of teaching (SET) scores is driven by factors the instructor did not control — class size, whether the course is compulsory or elective, the discipline, the workload, the time of day, and students' prior interest in the subject. Treating a situational outcome as a statement about a person's ability is a textbook case of the fundamental attribution error, committed at institutional scale. The remedy is not to discard the student voice but to read it as situated evidence and to collect feedback rich enough to tell the situation apart from the teaching.
The bias hiding in the committee room
Social psychologists have a name for our tendency to explain other people's behaviour by their character while explaining our own by circumstance. Lee Ross coined "the fundamental attribution error" in 1977 to describe how observers systematically over-weight disposition and under-weight situation. We see a colleague's low score and infer something about them; we explain our own dip as "that was the 8am required statistics course with 240 students."
Course evaluation is unusually fertile ground for this error because the format strips away context. A single mean lands on a spreadsheet next to a column of other means, inviting a like-for-like comparison that the underlying data does not support. A 12-person final-year elective that students chose out of genuine interest is placed alongside a 300-person mandatory first-year service course — and the number is read as if the two instructors faced the same task.
What the evidence actually shows
The empirical literature on SET has converged on an uncomfortable point: ratings correlate with a long list of things that are not teaching quality.
- Course characteristics matter. Kenneth Feldman's syntheses and the IDEA Center reviews (Benton & Cashin, 2012) document reliable associations between ratings and electivity, course level, discipline, expected grade, and perceived workload. Elective, higher-level, smaller, and humanities courses tend to rate higher — independent of who teaches them.
- Ratings barely track learning. The most cited recent meta-analysis, Uttl, White & Gonzalez (2017) in Studies in Educational Evaluation, re-analysed the multisection studies and found that once you account for prior ability and small-sample bias, the SET–learning correlation is essentially zero (r ≈ −0.02). If the number does not track learning, it is even less defensible to treat it as a measure of the teacher.
- The label drives the number. In the well-known natural experiment by MacNell, Driscoll & Hunt (2015), online instructors who swapped only their perceived gender received significantly different ratings for identical teaching. Boring, Ottoboni & Stark (2016) reached the same conclusion: SET scores reflect the situation the student is in, not the instruction they received.
Put together, this is the attribution error in numbers. The variance is real, but a large slice of it belongs to the context, not the person.
Why this matters for high-stakes decisions
When a committee ranks instructors on raw means, it is implicitly asserting that the differences reflect teaching quality. The European Standards and Guidelines (ESG 2015, Standard 1.9) ask institutions to monitor and periodically review programmes using evidence — but evidence used badly is worse than no evidence, because it launders a bias into a decision. Ranking a 9am compulsory course below a popular elective, and then attaching tenure or contract-renewal consequences to that ranking, is the fundamental attribution error with a paper trail. (We cover the decision-stakes problem in more depth in the case against using ratings for tenure and promotion.)
"But doesn't this just excuse bad teaching?"
This is the strongest objection, and it deserves a direct answer. No — recognising situational variance is not the same as claiming the signal is empty.
First, situational confounds reduce the validity of between-course comparisons; they do not erase within-context information. A pattern that recurs across several cohorts of the same course, holding the situation roughly constant, carries genuine signal. The error is comparing across incomparable contexts, not listening to students at all.
Second, the alternative to naive numbers is not silence; it is better evidence. The honest move is to keep the student voice and add the context that lets you interpret it. Qualitative comments, when properly analysed, frequently name the situation — "the room was too small," "I only took this because it's required," "the prerequisites weren't enforced." That is exactly the information a mean throws away.
Third, dispositional reading cuts both ways: a charismatic lecturer in an elective can score highly while a course quietly fails to teach. Treating high numbers as proof of quality is the same error in its flattering form. The Dr Fox effect — fluent, confident delivery inflating ratings regardless of content — is the mirror image.
How to read student feedback without the error
- Compare like with like. Benchmark courses against their own discipline, level, and size band, never against a single institution-wide mean. Report distributions and confidence intervals, not just a point estimate — a 0.3 difference between two small classes is usually noise.
- Collect context-aware, qualitative evidence. Ask students why, not just how much. Open, specific prompts surface the situational drivers that numbers hide.
- Triangulate. Combine student feedback with peer observation, curriculum review, and learning evidence so no single situated signal carries a career. (More on triangulation here.)
- Partition the variance. Multilevel models can estimate how much of the spread sits at the instructor level versus the course or cohort level — often less than intuition assumes. (See our piece on variance partitioning.)
Where Koji fits
The fundamental attribution error thrives on thin data. A single Likert mean gives a committee nothing to attribute a result to except the person whose name is on the course. Koji for Education is built to put the situation back into the evidence.
Instead of a static survey, Koji runs AI-moderated conversational interviews that probe beyond the number — when a student rates a course low, the moderator asks what specifically, surfacing whether the experience reflects a mandatory 8am slot with unenforced prerequisites or a genuinely unclear explanation of the material. The moderation is standardized and bias-aware, so the probing is consistent across hundreds of students rather than dependent on which human happened to run a focus group. Automatic thematic analysis of the open-text responses clusters the situational factors — workload, scheduling, room, prerequisite gaps — separately from teaching-specific feedback, giving committees a defensible way to tell them apart. Programme- and institution-level reporting lets you benchmark within comparable contexts rather than across incomparable ones, and formative, mid-cycle collection means problems can be attributed and fixed while the course is still running.
Teams that also run general user and customer research use the same AI interview engine on the main Koji platform — the conversational depth that separates situation from disposition is not unique to education; it is how good qualitative measurement works everywhere.
Legacy SET tools — EvaSys, paper forms, generic Qualtrics surveys — were built to average a Likert score. That design quietly hard-codes the attribution error: it gives you the number and discards the context that would let you interpret it. Koji is built for the opposite: evidence rich enough that you can see the situation before you judge the teacher.
The takeaway
Student feedback is worth collecting and worth acting on. But a number without its context is an invitation to blame a person for their circumstances. Reading evaluations as situated evidence — comparing within context, probing for the why, and triangulating before any consequence — is how an institution avoids committing the fundamental attribution error in its own quality processes.
Ready to collect course feedback that tells the situation apart from the teaching? See how Koji for Education works.