Do Online Courses Get Lower Evaluations? Course Modality as a Confound
Students often rate the same instructor teaching the same content lower when the course is delivered online. This guide separates course modality (how the course is taught) from administration mode (how the survey is run), reviews the evidence, and explains why modality must be controlled before comparing scores.
Koji Education Team
Product
Answer first: Yes — there is credible evidence that the same instructor teaching the same content tends to receive lower course evaluations when the course is delivered online rather than face-to-face. Crucially, this is not the same question as whether online survey administration lowers response rates; it is about the delivery mode of the course itself acting as a confound. Because modality is correlated with both ratings and with who enrols, comparing online and in-person scores on a single league table compares formats as much as teaching — and modality should be controlled, not ignored.
Two different "online" questions people keep conflating
There are two distinct "online" effects in course evaluation, and mixing them up produces bad decisions.
- Administration mode — whether the questionnaire is delivered on paper in class or online afterward. This mainly affects response rates and is covered in online vs paper course evaluations.
- Course modality — whether the course was taught in person, fully online, or in a hybrid/blended format. This is a property of the teaching being evaluated, not of the survey, and it is the subject of this article.
The pandemic forced a natural, if messy, experiment: millions of courses moved online and then back, often taught by the same instructors. The question for quality assurance is whether a lower evaluation for an online section reflects worse teaching, or simply the modality.
What the research says
The cleanest evidence isolates modality by holding the instructor and content constant. Marzano & Allen (2016), in the Online Journal of Distance Learning Administration, compared course evaluations for courses taught by the same instructor using the same content in both online and face-to-face formats. Their finding was direct: the online sections were rated lower than the face-to-face sections. Because the instructor and material were held constant, the gap is attributable to modality (and everything that travels with it — reduced presence, more self-direction, technical friction) rather than to the teacher being worse.
That result must be read against a second, well-known body of evidence about learning outcomes, which points the other way. The U.S. Department of Education meta-analysis (Means, Toyama, Murphy, Bakia & Jones, 2009/2010) found that, on average, students in online or blended conditions performed as well as or modestly better than those in purely face-to-face conditions. The tension is instructive: satisfaction ratings and learning outcomes can diverge. Students may learn at least as much online yet rate the experience lower — a pattern consistent with the wider finding, documented elsewhere on this site, that the feeling of learning and actual learning diverge. A course evaluation captures the former far more readily than the latter.
A third strand adds nuance and caution. Reviews of online-versus-face-to-face teaching consistently conclude that modality per se is a weak predictor once design quality is accounted for — well-designed online courses can match in-person satisfaction, and "emergency remote teaching" (hastily improvised pandemic delivery) is not the same construct as deliberately designed online learning. Several comparative studies of student evaluations across modalities report mixed or context-dependent gaps, with differences shrinking when course design, instructor experience with the format, and student preparation are controlled. In other words, the raw modality gap is real but is partly a proxy for design and circumstance.
Synthesising: (1) the same teaching often scores lower online, so modality is a genuine confound; (2) lower satisfaction does not imply less learning; and (3) much of the gap is mediated by course design rather than by modality being intrinsically worse.
Why it matters for course evaluation in practice
Never compare modalities on a single, unadjusted scale. If an institution ranks all sections on one league table while some ran online and some in person, online instructors carry a structural handicap that has nothing to do with their effectiveness. This is the same class of error as comparing across disciplines or elective versus required courses without adjustment.
Treat modality as a reported context variable. At minimum, evaluation reports should label each section's modality and compare like with like. Where numbers allow, modality can enter a model as a covariate so teaching signal is read net of format.
Question whether your items even mean the same thing across modalities. An item like "The classroom environment supported learning" does not have the same referent online. This is a measurement-invariance problem: if the items function differently across formats, comparing their means is comparing different constructs.
Distinguish the two "online" effects in your own data. Falling scores after a shift to online delivery may be a course-modality effect, an administration-mode effect on who responds, or both — and the remedies differ.
What to do when a programme moves online mid-cycle
The post-pandemic period left many programmes with mixed cohorts — some sections in person, some online, occasionally the same course in both within a single term. Three practical steps follow from the evidence. First, tag every section's modality at the point of collection, because retrofitting it later is error-prone and incomplete. Second, when reporting trends across years that span a modality shift, annotate the discontinuity rather than drawing an uninterrupted line: a dip that coincides with emergency remote teaching is a format artefact, not proof of declining teaching quality. Third, resist using a modality-mixed average to judge individuals, because the format gap is large enough relative to genuine teaching differences that it can dominate a ranking. Where the goal is improvement rather than comparison, lean on the open-text themes specific to each modality — the technical, social, and pedagogical complaints rarely point in the same direction, and only the last of them is about teaching.
Limitations and honest caveats
The evidence base has real weaknesses a critical reader should weigh. Same-instructor/same-content designs like Marzano & Allen control the obvious confounds but cannot randomise who chooses online sections; students self-select into modalities by schedule, motivation, and circumstance, so part of any gap is a selection effect, not a pure modality effect. Much of the recent data comes from pandemic-era emergency remote teaching, which conflates "online" with "stressful, involuntary, and improvised," and likely overstates the gap for well-designed online courses. The Department of Education meta-analysis has been critiqued for heterogeneity in what counted as "online" and for confounding format with additional learning time and resources. Effect sizes in this area are generally small and highly context-dependent, and publication and institutional-reporting practices vary. The honest position is that modality is a confound worth controlling, but its magnitude is unstable across settings and partly reducible to design quality and self-selection.
How Koji incorporates this
Koji for Education is built to keep modality visible and to keep the survey-mode question separate from the course-delivery question.
- Modality as first-class metadata. Koji captures each course's delivery mode (in person, online, hybrid) alongside size, level, and discipline, so reports can compare like-for-like and avoid one undifferentiated league table that penalises online sections.
- Probing the source of dissatisfaction. Where a flat survey records only a lower number for an online section, Koji's AI-moderated conversational interview asks why — distinguishing complaints about technical friction or isolation from complaints about the teaching itself. Automatic thematic analysis then separates modality-driven themes ("the recordings kept dropping") from pedagogy-driven ones, so a low score is diagnosed rather than merely recorded.
- Mode-aware item framing. Because an item about the physical classroom is meaningless online, Koji supports tailoring structured questions to the modality, reducing the measurement-invariance risk of asking the same literal item across formats.
- Separating administration mode from course modality. Koji is itself an online, conversational instrument, so it treats survey-delivery and course-delivery as distinct dimensions in analysis rather than collapsing "online" into a single label.
These are designed-to-mitigate mechanisms: they help interpret and contextualise a modality gap, not erase the self-selection and design confounds that produce it. Koji's core research platform at koji.so applies the same AI-moderated interview engine to product and customer research, where channel and context act as analogous confounds on satisfaction scores.
Related Resources
- Online vs. Paper Course Evaluations: Do Lower Response Rates Mean Worse Data?
- Students Rate the Classes Where They Learn the Most Lower: The Feeling-of-Learning Gap
- Measurement Invariance: Can You Compare Course-Evaluation Scores Across Groups at All?
- Why Quantitative Courses Get Lower Evaluations: The Discipline-Bias Problem
- Do Elective Courses Get Higher Ratings Than Required Ones?
- Does a Conversational Course Evaluation Make Students Less Honest?
References
- Marzano, R., & Allen, R. (2016). Online vs. Face-to-Face Course Evaluations: Considerations for Administrators and Faculty. Online Journal of Distance Learning Administration, 19(4). https://ojdla.com/archive/winter194/marzano_allen194.pdf
- Means, B., Toyama, Y., Murphy, R., Bakia, M., & Jones, K. (2009/2010). Evaluation of Evidence-Based Practices in Online Learning: A Meta-Analysis and Review of Online Learning Studies. U.S. Department of Education. https://eric.ed.gov/?id=ED505824
- Lowenthal, P. R., Bauer, C., & Chen, K.-Z. (2015). Student Perceptions of Online Learning: An Analysis of Online Course Evaluations. American Journal of Distance Education, 29(2), 85–97. https://doi.org/10.1080/08923647.2015.1023621
- Deslauriers, L., McCarty, L. S., Miller, K., Callaghan, K., & Kestin, G. (2019). Measuring actual learning versus feeling of learning in response to being actively engaged in the classroom. PNAS, 116(39), 19251–19257. https://doi.org/10.1073/pnas.1821936116
Related articles
Measurement Invariance: Can You Compare Course-Evaluation Scores Across Groups at All?
Comparing average ratings across departments, languages or online vs paper assumes the questionnaire means the same thing to every group. Measurement invariance is the test that assumption usually fails — and almost nobody runs it.
Why Quantitative Courses Get Lower Evaluations: The Discipline-Bias Problem
Uttl & Smibert (2017) show that instructors of quantitative courses receive systematically lower student evaluations than those teaching qualitative subjects — a bias with real career consequences. What the evidence says and how to compare ratings fairly across disciplines.
Does a Conversational Course Evaluation Make Students Less Honest? Mode Effects and Social Desirability
Survey mode shapes honesty: interviewer-administered surveys invite more social-desirability bias than self-administered ones. What Tourangeau and Yan (2007) and the mode-effects literature mean for anonymous, AI-moderated conversational course evaluations.
Online vs. Paper Course Evaluations: Do Lower Response Rates Mean Worse Data?
Online course evaluations consistently draw lower response rates than in-class paper forms, but the research shows the resulting scores are largely equivalent. Here is what Dommeyer, Nulty, and Stowell actually found, and what an adequate response rate really requires.