Required Courses Get Worse Evaluations. That Is a Bias, Not a Verdict
Electives outscore required courses. Small classes outscore large ones. Quantitative modules sit below the average. None of it necessarily reflects teaching quality — yet it all lands on the same dashboard. Here is the evidence on enrollment-context bias, and how to compare fairly.
Koji Education Team
Product · July 9, 2026
Bottom line up front: Decades of evidence show that course characteristics students did not choose and instructors cannot change — whether a course is elective or required, its class size, its level, and whether it is quantitative — systematically shift student evaluation scores. The effects are individually modest but consistent, and they are irrelevant to teaching quality. When a dean compares a required first-year statistics module against a final-year elective seminar on the same 1-to-5 dashboard, the comparison is rigged before anyone reads a word of feedback. This is a fairness problem with a statistical fix, and a design problem with a better answer than "collect a number and rank it".
What the evidence actually says
The classic synthesis is Kenneth Feldman's body of work from the late 1970s onward, which examined how course characteristics correlate with student ratings across dozens of studies. The recurring findings, confirmed by later reviews and summarised by teaching centres such as the University of Michigan CRLT, are:
- Elective courses are rated higher than required courses. Students who chose to be there arrive with more interest and more favourable expectations, and expectation strongly colours evaluation.
- Smaller classes receive slightly higher ratings than large ones. The relationship is modest but reliable across the literature.
- Course level matters. Upper-level and graduate courses tend to be rated higher than large introductory service courses.
- Quantitative and "hard" disciplines score lower than humanities and social sciences, independent of instructor.
Each individual effect is small — typically a fraction of a scale point. But three things make them dangerous. They stack: a required, large, first-year, quantitative service course is disadvantaged on all four dimensions at once. They are non-random: the same courses are penalised every single term, so the noise never averages out — it is a persistent bias, not sampling error. And they are invisible on the dashboard: the number arrives stripped of the context that explains it.
Why this is not "teaching quality"
Set against a crucial baseline: the most rigorous recent evidence questions whether SET scores track teaching effectiveness at all. The meta-analysis by Uttl, White and Gonzalez (2017) found that, once small-sample and publication-bias artefacts are controlled, student ratings explain at most about 1% of the variability in student learning — a statistically non-significant relationship. If scores barely reflect learning, then a systematic gap between a required course and an elective almost certainly reflects the conditions of enrolment, not the quality of instruction inside the room.
Consider what this means concretely. A department chair sees that the compulsory Year-1 "Quantitative Methods" module sits at 3.6 while a Year-3 optional "History of Ideas" seminar sits at 4.5. The naive reading — "the methods lecturer is weaker" — is almost the textbook definition of the fundamental attribution error: attributing to the person what belongs to the situation. The methods lecturer is teaching a required, large, first-year, quantitative course to students who would rather be elsewhere. Four independent handicaps, none of them theirs.
The compounding injustice
Enrollment-context bias does not fall evenly. Required service teaching, large introductory cohorts, and quantitative "gateway" modules are disproportionately staffed by early-career academics, teaching-focused faculty, and contingent staff — precisely the people for whom a low evaluation score carries the highest stakes for contract renewal and progression. The structural bias in which courses score low therefore maps onto the structural precarity of who teaches them. A system that ranks raw means does not just mislead; it can systematically disadvantage the most vulnerable teachers for delivering the courses their institution most needs delivered.
Critics argue: "Required and large courses really are worse experiences"
The honest counterargument is that these differences are not pure artefact. Large lectures genuinely afford less interaction; required courses genuinely include disengaged students; gateway quantitative modules genuinely are harder and more anxiety-inducing. If students report a worse experience in those settings, isn't that real information a university should act on?
Yes — but that concedes the key point rather than refuting it. If the low score is driven by class size, compulsion, and subject difficulty, then the correct institutional response is to change those conditions (invest in tutorial support, redesign the gateway sequence, resource the large cohort) — not to conclude the instructor is underperforming and act on them personally. The counterargument actually strengthens the case: the score is telling you about the course design and resourcing, which is a management responsibility, not a verdict on the individual standing at the front. The failure is using a situational signal to make a personnel judgement.
The statistical fix, and its limits
Institutional research offices have a legitimate toolkit here. You can benchmark a course only against comparable courses — same level, similar size, similar discipline — rather than a single institution-wide threshold. You can model scores with the course characteristics as covariates, so a "3.6 in a large required quantitative module" is read against the expected score for that profile, not against 4.5. Empirical-Bayes shrinkage and multilevel models (covered elsewhere on this blog) formalise exactly this.
But statistical adjustment has a ceiling. It can tell you a required course scored as expected for its type — it cannot tell you what specifically to improve, because it is still working with a single collapsed number. Correcting the bias in the comparison is necessary but not sufficient. You also have to change what you collect.
Where Koji fits — evaluating the course, not the enrolment conditions
Koji for Education is designed to separate the signal a programme can act on from the enrolment context it cannot. Rather than reducing a required module to one comparable-or-not average, it collects structured, probeable feedback that stands on its own terms:
- Criterion-referenced, behavioural questions — via six structured question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) — ask whether specific teaching practices happened ("Were worked examples explained at a pace you could follow?") rather than a global "rate this course", so a compulsory module is judged against a standard, not against an unrelated elective.
- AI-moderated conversational probing distinguishes "I resent that this course is required" from "the sequencing of topics did not build logically" — two very different meanings behind the same low number, which a static form flattens into one.
- Automatic thematic analysis surfaces whether low scores cluster around fixable design and resourcing issues (cohort too large for the tutorial model, prerequisites assumed but not taught) rather than the instructor — turning a demoralising average into a management action list.
- Programme- and institution-level reporting lets QA teams see patterns across comparable modules, so a systematically disadvantaged course type becomes visible as a structural issue rather than a string of "weak" individuals.
Koji is explicit about scope: it mitigates and surfaces enrolment-context distortion; it does not claim to eliminate it. The same conversational interview engine also powers the main Koji platform for customer and user research, where separating situational noise from actionable signal is the entire job.
What to do on Monday
Three moves cost nothing and improve fairness immediately. First, stop comparing raw means across dissimilar course types — publish and use only like-for-like benchmarks. Second, annotate every course's score with its profile (required/elective, size, level, discipline) so no committee reads a number without its context. Third, treat a low score on a required, large, or quantitative module as a prompt to examine course design and resourcing before ever examining the instructor.
Required courses will keep scoring below electives for as long as we ask students to rate them on the same scale. The task is not to pretend the gap away — it is to stop mistaking it for a judgement on the person who volunteered to teach the course nobody chose.
Want evaluations that judge the teaching, not the timetable? Explore Koji for Education to see criterion-referenced, conversational course evaluation in practice.