Do Elective Courses Get Higher Ratings Than Required Ones? Course Characteristics and the Feldman Evidence
Course features the instructor cannot control — whether a course is elective or required, its level, size, subject, and timing — systematically shift student ratings. What Kenneth Feldman's synthesis established, and why comparing raw ratings across different course types is unfair.
Koji Education Team
Product
In brief: A substantial body of evidence, synthesised by Kenneth Feldman (1978), shows that characteristics of the course itself — not the teaching — predictably move student ratings. Elective courses and courses in a student's major are rated higher than required service courses; larger classes and certain disciplines tend to rate lower. These associations are mostly small to moderate, but they are systematic, which means comparing a raw rating from a compulsory 300-student statistics service module against an optional 20-student seminar is not a fair comparison. Quality assurance should benchmark within comparable course types, not across them.
What the research says
When an evaluation committee lines up instructors by their mean rating, it implicitly assumes the rating reflects the instructor. Decades of research say otherwise: part of every score is attributable to the course, independent of who teaches it.
The anchor synthesis: Feldman (1978)
Kenneth Feldman's review, "Course characteristics and college students' ratings of their teachers: what we know and what we don't" (Research in Higher Education, 1978), remains the definitive map of this terrain. Feldman organised the evidence around five course characteristics:
- Electivity — whether students chose the course or were required to take it, and how closely it related to their major or prior interest.
- Course level — introductory versus advanced/upper-division.
- Class size — number of enrolled students.
- Subject matter / discipline — the academic field.
- Time of day the course meets.
His central findings on the two most consequential: Electivity is reliably and positively associated with ratings — instructors are rated higher in courses students elected, or that connect to a pre-existing interest or major, than in compulsory service courses students take only to satisfy a requirement. Class size is generally inversely related to ratings, though several studies found a curvilinear (roughly U-shaped) pattern, with the very smallest and sometimes the very largest classes rating relatively higher than mid-sized ones. Course level and discipline showed weaker and more context-dependent effects, and time of day little consistent effect. Crucially, Feldman was careful: these are associations of modest magnitude, frequently confounded, and "what we don't know" was as much his message as "what we know".
Corroboration and refinement
Feldman's picture has held up. Herbert Marsh's extensive programme of research using the SEEQ instrument (Marsh, 1987) confirmed that prior subject interest is one of the strongest correlates of overall ratings, and that workload/difficulty and expected grade also carry small associations — while concluding that, taken together, these "biasing" characteristics explain a relatively small portion of total rating variance, leaving room for ratings to still reflect teaching. Benton and Cashin's synthesis for the IDEA Center (2012) likewise reported that student motivation and the elective/required distinction are among the course characteristics most consistently related to ratings, and recommended interpreting scores relative to comparison groups of similar courses. The discipline effect — quantitative and "hard" fields (mathematics, statistics, engineering) tending to receive lower ratings than humanities — has been documented repeatedly and is the subject of its own literature.
Why it matters for course evaluation in practice
The single most important practical consequence is about fairness of comparison:
- Raw cross-course rankings are biased. A league table that ranks all instructors in a faculty by mean rating penalises whoever teaches the compulsory, large, quantitative service course — regardless of how well they teach it. The course they were assigned depresses their score.
- Benchmark within comparable strata. Compare an instructor against the distribution of ratings for courses of similar electivity, level, size, and discipline — not against the faculty-wide average. A 3.9 in a required 250-person first-year statistics module may be stronger evidence of good teaching than a 4.4 in an optional final-year seminar.
- Adjust expectations, not just numbers. Even without formal statistical adjustment, committees should read scores in context: note the course type beside every mean and resist the temptation to treat a half-point gap as a teaching-quality gap when it tracks a course-type gap.
- Protect staff in structurally disadvantaged slots. Early-career and contingent staff are disproportionately assigned the large required service courses that rate lower. Ignoring course characteristics turns a timetabling artefact into a career penalty.
A worked example makes the fairness point concrete. Suppose Dr A teaches a compulsory first-year quantitative-methods course of 240 students and earns a mean of 3.8; Dr B teaches an optional final-year seminar of 18 students in her specialism and earns a 4.5. Read naively, Dr B "out-teaches" Dr A by 0.7 points. But every structural feature Feldman identified — high electivity, small class, advanced level, high prior interest, a self-selected cohort — works in Dr B's favour and against Dr A, before either of them says a word in front of a class. The raw gap is largely a property of the timetable, not of teaching skill. A committee that benchmarks each against the distribution for that kind of course may well conclude Dr A is the stronger teacher. This is why leading guidance (Benton & Cashin, 2012; Linse, 2017) insists on contextual comparison groups rather than a single institutional yardstick.
Limitations and honest caveats
- Most effects are small and confounded. Electivity is entangled with prior interest, motivation, and self-selection; class size is entangled with course level and pedagogy. Feldman himself stressed that isolating a clean "course-characteristic effect" from teaching is difficult, and that observational data cannot fully separate them.
- The evidence base is dated and largely North American. Feldman (1978) and much of the corroborating work predate online evaluation, mass higher education, and the European programme structures this audience works within. The direction of effects is robust; the magnitudes should not be imported uncritically into a Bologna-process context.
- "Bias" is the wrong word for some of it. If students genuinely learn and engage more in courses they chose, a higher rating may partly reflect a real difference in the educational experience, not a measurement artefact. The electivity effect is part artefact, part substance — and the two are hard to disentangle.
- Curvilinearity matters. The U-shaped class-size pattern means a simple linear "big classes rate lower" adjustment would itself be wrong. Any statistical correction must respect the actual functional form.
- Adjustment can over-correct. Statistically "controlling away" course characteristics risks erasing real teaching differences if the model is mis-specified. Transparent stratified benchmarking is usually safer than opaque regression adjustment.
A defensible reading is narrow: course characteristics move ratings enough that raw cross-course comparison is unfair, and scores should be interpreted within comparable course strata — without pretending the effects are large or cleanly causal.
How Koji incorporates this
Koji is designed to make course context a first-class part of the evaluation, so that scores are read in the light of the course type rather than ripped out of it. These mechanisms mitigate — they do not eliminate — course-characteristic bias.
- Context-aware, stratified reporting. Koji's reporting is built to compare a course against relevant peers — similar level, size, electivity, and discipline — rather than against a single institution-wide mean, directly addressing Feldman's fairness problem.
- Structured metadata on every study. Because each evaluation can carry structured attributes (course type, level, cohort), Koji can surface "this is a required first-year service course" alongside the score, prompting committees to interpret in context.
- AI-moderated conversational interviews that separate the course from the teaching. Rather than a single global rating that silently absorbs the electivity halo, Koji's interviewer can probe why a student rated as they did — distinguishing "I resented being required to take this" from "the instruction was weak". That qualitative layer is exactly what a raw mean hides.
- Question types that isolate constructs. Using
single_choiceandscaleitems on prior interest and motivation, plusopen_endedprobing, lets a programme measure the confounds (prior interest, electivity) explicitly instead of letting them leak into the overall score. - Triangulation and trend analysis let a programme follow the same course over time — a far fairer comparison than ranking unlike courses against each other in a single cycle.
Koji is designed to mitigate course-characteristic bias by foregrounding context and enriching the signal, not by pretending a single adjusted number is the truth. The core research platform at koji.so applies the same context-aware, AI-moderated approach to customer and product research, where comparing unlike segments on a single score is an equally common error.
Related Resources
- Class Size and Student Evaluations: What Bedard and Kuhn Found
- Does Course Difficulty and Workload Lower Student Evaluations?
- Why Quantitative Courses Get Lower Evaluations: The Discipline-Bias Problem
- Likability and Prior Subject Interest: The Hidden Confounds
- Selection Bias in Course Evaluations: What Goos and Salomons Found
- Measurement Invariance: Can You Compare Scores Across Groups at All?
References
- Feldman, K. A. (1978). Course characteristics and college students' ratings of their teachers: What we know and what we don't. Research in Higher Education, 9(3), 199–242. https://doi.org/10.1007/BF00976997
- Marsh, H. W. (1987). Students' evaluations of university teaching: Research findings, methodological issues, and directions for future research. International Journal of Educational Research, 11(3), 253–388. https://doi.org/10.1016/0883-0355(87)90001-2
- Benton, S. L., & Cashin, W. E. (2012). Student ratings of teaching: A summary of research and literature (IDEA Paper No. 50). The IDEA Center. https://www.ideaedu.org/idea_papers/student-ratings-of-teaching-a-summary-of-research-and-literature/
- Cohen, P. A. (1981). Student Ratings of Instruction and Student Achievement: A Meta-analysis of Multisection Validity Studies. Review of Educational Research, 51(3), 281–309. https://doi.org/10.3102/00346543051003281
Related articles
Measurement Invariance: Can You Compare Course-Evaluation Scores Across Groups at All?
Comparing average ratings across departments, languages or online vs paper assumes the questionnaire means the same thing to every group. Measurement invariance is the test that assumption usually fails — and almost nobody runs it.
Does Course Difficulty and Workload Lower Student Evaluations? What Centra Found
A research-grounded look at whether harder, heavier courses are punished in student evaluations of teaching. Centra (2003) and Marsh & Roche (2000) show the relationship is non-linear and weaker than faculty fear — with practical implications for fair course evaluation.
Why Quantitative Courses Get Lower Evaluations: The Discipline-Bias Problem
Uttl & Smibert (2017) show that instructors of quantitative courses receive systematically lower student evaluations than those teaching qualitative subjects — a bias with real career consequences. What the evidence says and how to compare ratings fairly across disciplines.
Class Size and Student Evaluations: What Bedard and Kuhn Found
Does class size bias student evaluations of teaching? Bedard and Kuhn (2008) found a large, non-linear negative effect of enrolment on instructor ratings even after controlling for instructor and course. This article synthesises the evidence and explains how to stop class size from contaminating cross-instructor comparisons.