Likability and Prior Subject Interest: The Hidden Confounds in Student Evaluations
Two things a lecturer cannot control — how likeable students find them and how interested students already were in the subject — quietly move evaluation scores. The evidence on which is a real bias, and which is not.
Koji Education Team
Product
In short: Two variables outside an instructor's control shape course-evaluation scores: how likeable students find the lecturer, and how much prior interest students already had in the subject. Feistauer and Richter (2018) found likability exerts a substantial bias with no conceptual link to teaching quality, while prior subject interest exerts only a weak one. Marsh's classic work shows student and course background variables together explain roughly 12–14% of rating variance. The practical rule: likability is a genuine validity threat to guard against; prior interest is better understood as a course characteristic than a teacher failing — and neither is visible in a raw average.
Two confounds, two very different verdicts
A great deal of bias research treats every non-teaching influence on ratings as equally damning. The more careful finding is that they are not equal. One of these confounds — likability — is a real threat to the validity of an evaluation as a measure of teaching. The other — prior subject interest — is mostly a feature of the course and its students, and only becomes a "bias" when scores are misused. Telling them apart is what separates a defensible interpretation from a lazy one.
What the research says
The sharpest recent evidence is Feistauer and Richter (2018) in Studies in Educational Evaluation. In a German university, 260 students evaluated psychology courses across a semester, yielding 517 data points, and the researchers measured both likability and prior subject interest at two times — the beginning of the course and the moment of evaluation. Using cross-classified multilevel models (which correctly separate variance attributable to students, courses and instructors), they found:
- Likability exerts a substantial bias on evaluations at both measurement times. Critically, the bias is overestimated when likability is measured at the time of evaluation — because by then the teaching experience has fed back into how likeable the lecturer seems — and is better gauged from a baseline measure. Since likability bears no conceptual relationship to teaching quality, its influence directly compromises the validity of the rating.
- Prior subject interest exerts only a weak bias. Students who arrive already interested rate a bit more favourably, but the effect is small relative to likability.
Their conclusion is pointed: because likability is doing work that has nothing to do with whether teaching was good, evaluations used as if they were pure measures of teaching quality are, to that extent, invalid.
This refines rather than overturns the foundational study, Marsh (1980) in the American Educational Research Journal. Marsh examined 16 background characteristics across 511 undergraduate courses taught by 221 instructors. Each single background variable generally explained less than 5% of the variance in any one rating, but the whole set explained about 12–14%. In stepwise analyses the most influential variables were expected grades, workload/difficulty, prior subject interest, and taking the course for general interest. So prior subject interest is a long-recognised, reliable correlate of ratings — but a modest one, and embedded in a cluster of course-level factors.
Marsh's own interpretation, developed with Dunkin, is that prior subject interest does not constitute a bias in the damning sense — except where ratings are used summatively — because it largely reflects the course (its appeal, its audience) rather than anything the teacher did or failed to do. That is the crucial distinction: a variable can systematically move scores without being a "bias" in the validity sense, if it reflects something legitimately part of what is being evaluated. Likability fails that test; prior interest mostly passes it.
Why it matters for course evaluation in practice
For a quality office, the two confounds call for opposite responses.
Likability should be actively guarded against, because it is a validity threat masquerading as a teaching signal. A warm, charismatic, well-liked lecturer can earn high scores while a demanding, less personable, but genuinely effective one earns lower ones — and the raw average will not tell you which is happening. This is the same family of effect as the Dr Fox phenomenon and the halo effect, and the defence is the same: do not let a single global rating, or a single likeability-laden item, stand in for teaching quality, and triangulate against evidence that is harder to charm.
Prior subject interest should be contextualised, not "corrected away". Because it reflects the course and its students, the right move is to compare like with like — not to rank a compulsory quantitative methods module against a self-selected elective and pretend the gap is about teaching. Marsh's framing is the operative principle: prior interest only becomes a problem when summative decisions ignore it. Reporting that presents a module's scores alongside its enrolment context is far more honest than a context-free league table.
The unifying lesson from Marsh is one of proportion: student and course background factors together account for only around an eighth of rating variance, so evaluations are not dominated by confounds — but that eighth is more than enough to flip a close comparison or a borderline promotion case, which is exactly where the stakes are highest.
Limitations and honest caveats
A critical reader should weigh several limitations:
- Single-institution, single-discipline samples. Feistauer and Richter studied psychology students at one German university; likeability and interest effects may differ by discipline, culture and course level.
- Measuring "likability" is itself slippery. It is partly endogenous to teaching — a good teacher is often more likeable — so disentangling the bias portion from legitimate rapport is genuinely hard, and estimates depend on when likability is measured.
- "Variance explained" is not "bias". Marsh's 12–14% includes variables (like workload) whose status as bias is contested; the figure bounds the problem rather than proving misuse.
- Correlational designs limit causal claims. Most of this evidence is observational; experimental manipulations of likability are artificial, and field studies cannot fully rule out reverse causation.
The honest summary: likability is a well-supported validity threat of moderate size; prior subject interest is a small, mostly legitimate course effect; and both are invisible in the number a committee actually reads.
How Koji incorporates this
Koji is an AI-native course-evaluation platform whose design is aimed at separating teaching signal from rapport and predisposition — framed, deliberately, as designed to mitigate rather than eliminate these confounds:
- Behaviourally specific questions instead of global affect. Rather than leaning on a single "overall, how good was this lecturer?" item that soaks up likeability, Koji's
scaleandopen_endedquestions target concrete teaching behaviours (clarity of explanation, quality of feedback, structure), which are harder to answer from pure liking. - AI-moderated probing for evidence. When a student gives a glowing or harsh rating, the moderator asks why and for an example — surfacing whether the judgement rests on a teaching behaviour or on "they were nice / I never liked this subject", information a Likert number hides.
- Capturing context, not just scores. Koji can collect prior-interest and course-context signals so that reporting compares like with like, exactly as Marsh's distinction requires, instead of producing context-free rankings.
- Bias-aware, triangulated reporting. Results are presented to be read against other evidence and within enrolment context, resisting the summative misuse that turns a legitimate course effect into an unfair penalty.
The same AI-moderated interview engine runs Koji's core research platform at koji.so, where separating how much customers like a brand from how well a product actually performs is the identical methodological challenge; in education it is tuned to protect the validity of teaching judgements.
Rapport is not the same as pandering
A fair objection: is likability really a problem, when a warm, trusted teacher genuinely helps students learn? The distinction matters. Rapport that supports learning — being approachable, responsive, fair — is part of effective teaching and should show up in ratings. The validity threat is the residual liking that has no instructional content: the lecturer students enjoy being around but who explains poorly, gives thin feedback, or sets unclear expectations. Feistauer and Richter's finding that likability bias is overestimated when measured at evaluation time is the practical clue here — by the end of term, legitimate teaching quality and raw affect have fused, which is exactly why a single global "how was the lecturer?" item cannot separate them. The remedy is not to suppress warmth but to decompose the judgement: ask about specific teaching behaviours that a merely likeable-but-ineffective teacher would score poorly on, and read those against the global score. A large gap — high overall warmth, weak specific behaviours — is the signature of likability bias doing the work, and it is precisely what a behaviourally specific, probing instrument is built to expose.
Related resources
- Does Grading Leniency Inflate Student Evaluations?
- Does Course Difficulty and Workload Lower Student Evaluations? What Centra Found
- What Do Student Evaluations Actually Measure? Marsh and the SEEQ
- The Halo Effect in Course Evaluations
- The Dr. Fox Effect: Does Lecturer Charisma Inflate Course Evaluations?
- Do Student Evaluations Measure Learning? The Uttl Meta-Analysis Revisited
References
- Feistauer, D., & Richter, T. (2018). Validity of students' evaluations of teaching: Biasing effects of likability and prior subject interest. Studies in Educational Evaluation, 59, 168–178. https://doi.org/10.1016/j.stueduc.2018.07.009
- Marsh, H. W. (1980). The influence of student, course, and instructor characteristics in evaluations of university teaching. American Educational Research Journal, 17(2), 219–237. https://doi.org/10.3102/00028312017002219
- Feistauer, D., & Richter, T. (2018). The role of clarity about study programme contents and interest in student evaluations of teaching. Psychology Learning & Teaching, 17(3), 332–351. https://doi.org/10.1177/1475725718779727
Related articles
What Do Student Evaluations Actually Measure? Marsh, the SEEQ, and the Case for Multidimensional Feedback
Herbert Marsh spent decades showing that well-built student evaluations are multidimensional, reliable, and stable — and that a single global score throws away most of what they can tell you. What his work establishes, where critics push back, and how to design feedback that is actually usable.
Does Course Difficulty and Workload Lower Student Evaluations? What Centra Found
A research-grounded look at whether harder, heavier courses are punished in student evaluations of teaching. Centra (2003) and Marsh & Roche (2000) show the relationship is non-linear and weaker than faculty fear — with practical implications for fair course evaluation.
The Halo Effect in Course Evaluations: When One Impression Colours Every Rating
When students like an instructor, that single global impression bleeds into their ratings of unrelated specifics. From Nisbett & Wilson (1977) to Feistauer & Richter (2018) and Cannon & Cipriani (2021), here is how the halo effect distorts itemised course evaluations — and what to do about it.
Do Student Evaluations Measure Learning? The Uttl Meta-Analysis Revisited
A research-grounded reading of the Uttl, White & Gonzalez (2017) meta-analysis on the SET–learning relationship, with implications for European course evaluation and quality-assurance policy.