Do Online and Hybrid Courses Get Lower Student Evaluations? The Evidence on Modality Bias
Online and hybrid courses often score lower on student evaluations than face-to-face equivalents — but the gap is smaller, more conditional, and more confounded than most committees assume. Here is what the controlled evidence actually shows, and how to evaluate teaching fairly across modalities.
Koji for Education
Research & Editorial Team ·
Bottom line up front: Across the literature, online and hybrid courses tend to receive slightly lower student evaluations of teaching (SET) than face-to-face equivalents, but the effect is small, inconsistent, and heavily confounded by response rates and who self-selects into each modality. Treating a lower online score as evidence of worse teaching is not defensible. The honest position is that delivery mode is one more variable that contaminates the raw SET number — which is exactly why a single average is the wrong instrument for comparing teaching across formats.
Why this question matters now
Since 2020, almost every European university teaches a blend of in-person, hybrid, and fully online courses, often taught by the same staff to overlapping cohorts. When promotion panels, programme reviews, and accreditation self-assessments line up SET scores side by side, they implicitly assume the number means the same thing in each format. It does not. If modality systematically shifts the score independent of teaching quality, then ranking a lecturer's online section against a colleague's seminar is comparing two differently-calibrated instruments and calling it a measurement.
What the controlled evidence shows
The cleanest design holds the instructor and the course constant and varies only the delivery mode. In one such study, Marzano and Allen (2016) found that courses taught by the same instructor with the same content were rated lower in the online modality than face-to-face. That within-instructor comparison is the strongest argument that something about the format — not the teaching — moves the rating.
But the picture is genuinely mixed. Other analyses comparing online and in-person sections find no meaningful difference once you look closely. A text-analytic comparison of open-ended comments across hundreds of students and dozens of sections found no significant difference in the proportion of positive appraisal segments by delivery method — suggesting that when students explain themselves rather than circle a number, the apparent "bias" can shrink or vanish. Reviews also note that partially online (hybrid) courses are often rated slightly more favourably than fully online ones, hinting that the effect is about the loss of in-person contact and immediacy rather than digital delivery per se.
So the defensible summary is: a real but small and conditional tendency for online SET to run lower, not a robust, uniform penalty.
The confounds that make raw comparisons meaningless
Three confounds matter more than most evaluation committees acknowledge:
1. Response rates collapse online. When evaluations move from in-class paper to online administration, participation drops sharply — from roughly 70–80% on paper to around 50–60% online in many institutions, a gap of over 20 percentage points in review after review. Lower response rates do not just add noise; they change who answers. Students with strong opinions — often negative — are disproportionately likely to respond when participation is voluntary and unsupervised, which can drag online averages down for reasons that have nothing to do with the teaching.
2. Self-selection into modality. Students who choose (or are forced into) online sections differ systematically — in work commitments, commuting distance, motivation, and prior achievement. If the online cohort is more stressed or more instrumental about a credential, their ratings reflect their circumstances, not the instructor's effectiveness.
3. Technology friction is attributed to the teacher. A flaky video platform, a clunky LMS, or a poorly designed institutional portal degrades the student experience, and students routinely fold that frustration into their rating of the instructor. The lecturer is then penalised for infrastructure they did not build and cannot fix.
But doesn't a lower score still tell us something?
A fair critic will object: even if it is confounded, isn't a consistently lower online score a signal worth heeding? Yes — but a signal about the experience, not a verdict on the teacher. The error is not in noticing the gap; it is in attributing it. If online sections systematically frustrate students, that is actionable intelligence about course design, platform choice, and support — institutional problems, mostly. The methodological sin is converting "students found the online experience harder" into "this lecturer teaches worse," then attaching consequences to it.
This connects to a deeper, well-documented problem. The validity of SET as a measure of teaching effectiveness is weak to begin with: the landmark meta-analysis by Uttl, White and Gonzalez (2017) found that SET ratings are essentially unrelated to how much students actually learn. Layering a modality confound on top of an already shaky instrument compounds the problem rather than creating a new one.
What a fair evaluation across modalities looks like
If you cannot compare raw averages, what can you do?
- Never benchmark online scores against face-to-face norms. If you must use numbers, compare like-with-like — online against online, within discipline and level.
- Report response rates and treat low-response results as low-confidence. A 35% online response rate is not a representative sample; label it accordingly.
- Separate the platform from the pedagogy. Ask explicitly about the technology and the institutional tools so that infrastructure frustration does not bleed into judgements of the teacher.
- Privilege the qualitative. When students explain why, modality artefacts become legible — "the recordings kept buffering" is obviously not a teaching deficit. Numbers hide that; narratives surface it.
This is where the instrument itself, not just its interpretation, needs to change. Koji for Education replaces the static Likert form with AI-moderated conversational interviews that probe beyond a single number. Because the AI follow-up is standardized rather than dependent on an inconsistent human moderator, every student in every modality is asked to elaborate with the same rigour — so a low score in an online section gets interrogated ("What specifically made this harder?") instead of recorded and ranked. Koji's automatic thematic analysis then separates platform complaints from pedagogy at scale, surfacing whether an online cohort struggled with the teaching or the tooling. It will not magically equate modalities — nothing can fully de-confound self-selection — but it makes the underlying causes visible rather than collapsing them into one misleading average. (The same conversational interview engine powers general user and customer research on the main Koji platform for teams that evaluate experiences beyond the classroom.)
Modality bias is real enough to take seriously and small enough to resist panic. The right response is not to defend or attack online teaching, but to stop pretending a single cross-format average measures it.
The effect is not uniform across course types
One reason the literature looks contradictory is that "online" is not a single thing, and neither is the teaching it replaces. A large lecture that was already a one-to-many broadcast loses relatively little in translation to video; a lab, studio, clinical placement, or small discussion seminar loses a great deal of what made it work in person — the hands-on apparatus, the embodied feedback, the spontaneous back-and-forth. We should expect — and the partial-versus-fully-online pattern noted above is consistent with — a larger modality penalty where the in-person format carried more pedagogical weight that digital delivery cannot reproduce.
That has a direct fairness implication: the modality "bias" is partly a real degradation of a specific learning experience, not a pure artefact. A chemistry practical delivered as a video is genuinely a lesser version of the course, and students rating it lower are not being irrational. The error is still attributing that to the instructor, who often had no say in the decision to move the course online. Aggregating a forced-online lab and a natively-online lecture into one "online" benchmark mixes these cases and produces a number that means nothing for anyone.
Key takeaways
- Online and hybrid courses tend to score slightly lower on SET, but the effect is small, inconsistent, and confounded by response rates, self-selection, and technology friction.
- Within-instructor studies (same teacher, same content) show a modest online penalty; open-ended and hybrid comparisons often show little or none.
- Online administration cuts response rates by 20+ points, skewing who responds and dragging averages down for non-pedagogical reasons.
- Never benchmark online scores against face-to-face norms; report response rates; separate platform from pedagogy; and privilege qualitative feedback.
- Conversational, AI-moderated evaluation surfaces why an online cohort rated as it did — distinguishing teaching from tooling instead of hiding both in one number.