Did COVID Lower Course Evaluations? Emergency Remote Teaching as a Natural Experiment in Confounding
What student-evaluation data from the 2020 emergency shift to remote teaching reveals about how much SET scores reflect factors outside an instructor's control — and why pandemic-era ratings need a giant asterisk in any personnel decision.
Koji Education Team
Product
In brief. When universities moved online in weeks in spring 2020, student-evaluation scores did not simply fall — they moved in different directions for different questions, disciplines, and cohorts. Studies tracking the same programmes before, during, and after the pandemic find that ratings of teacher support and empathy often rose, while ratings of interactivity and hands-on practice fell, largely independent of individual teaching skill. That divergence is the clearest natural-experiment evidence we have that a substantial part of a Student Evaluation of Teaching (SET) score measures the conditions of instruction, not the instructor. Pandemic-era scores should never be read as a clean signal of teaching quality.
Why the pandemic is a useful accident
Course evaluation researchers rarely get a controlled shock: the same instructors, the same courses, the same students, and then one variable — the mode of delivery — changes abruptly for everyone at once. The emergency remote teaching (ERT) period of 2020 was exactly that. It is not the same as planned online education (which is designed, resourced, and chosen); it was improvised under stress. That makes it a powerful lens on a long-standing question: how much of a course-evaluation score is about the teacher, and how much is about everything else?
What the research says
Studies that compare SET across pre-pandemic, pandemic, and post-pandemic periods within the same institutions tell a consistent, textured story rather than a simple "scores dropped".
A longitudinal analysis published in Frontiers in Education (2025) tracking student evaluation of teaching prior to, during, and after COVID-19 found that students gave higher ratings on items about the teacher''s professional, motivating, and supportive attitude during the pandemic than before or after — while giving lower ratings on the interactivity of practice-based courses during the same period. In other words, the same instructors were rated more warmly on empathy and more harshly on hands-on engagement, simultaneously, because the situation changed — not their competence.
A separate Frontiers in Education (2022) study of an institution''s abrupt educational-model transition found the effect was discipline-dependent: in the traditional model, the School of Engineering and Sciences showed an improvement in SET during the shift, while the School of Medicine and Health Sciences showed a decline. Practice- and clinical-heavy programmes, whose core value was hardest to deliver remotely, took the biggest hit — again, a property of the subject and its delivery constraints, not of any individual teacher.
These findings sit on top of a mature literature warning that SET scores track things other than learning. Uttl, White and Gonzalez (2017), in a re-analysis of decades of multisection studies, found essentially no meaningful correlation between SET ratings and actual learning once older studies'' small-sample artefacts were controlled — and Uttl has argued specifically (2021, 2023) that pandemic disruption compounds the existing case against using SET results to judge teaching effectiveness. The pandemic data also intersect with grade effects: analyses in this period reported the familiar moderate association between the grades students expect and the ratings they give, a confound that ERT''s grading disruptions only muddied further.
The mechanism is well understood. Ratings during ERT absorbed variance from: technology access (bandwidth, devices, quiet space), student stress and isolation, the loss of laboratory, studio, and placement experiences, delayed or degraded feedback because remote assessment was harder, and shifted expectations about what a course could deliver. None of these is a teaching-skill variable. Yet all of them landed inside the SET score.
Why it matters for course evaluation in practice
1. Pandemic-era scores are contaminated evidence for personnel decisions. An instructor whose interactivity rating fell in 2020 was, in many cases, teaching a subject that could not be made interactive over Zoom that term. Using such a score in a tenure, promotion, or contract-renewal case imports a confound the instructor could not control. Where SET already misclassifies instructors even in normal times, ERT amplifies the error.
2. Year-over-year trend lines break at 2020. Any dashboard that plots a course''s mean rating across years will show a discontinuity in 2020–21. Reading that as "the teaching changed" is a mistake; the measurement conditions changed. This is exactly the situation an interrupted time series analysis is built to handle — model the level shift explicitly rather than pretend the series is continuous.
3. It is a general lesson, not a historical footnote. The pandemic made visible a permanent property of SET: modality, resourcing, and student circumstances move scores. The same logic applies to a course forced online by a strike, a room double-booked all term, or a cohort hit by a local crisis. The modality confound did not end in 2022.
5. It exposed an equity dimension inside the score. The pandemic did not distribute its confounds evenly. Students with poor connectivity, no quiet study space, caring responsibilities, or disabilities that remote delivery served badly experienced a materially different course than their better-resourced peers — and rated it accordingly. A pooled SET mean silently averages those experiences together, so a low score can encode a resourcing and access failure the institution owns rather than a teaching failure the instructor owns. Reading pandemic (and post-pandemic online) ratings without disaggregating by student circumstance risks mistaking a digital-divide effect for a quality signal, and penalising the very instructors who taught the most disadvantaged cohorts.
4. Warmth and interactivity are separable. The pandemic''s split result — support up, interactivity down — is direct evidence for multidimensional reporting. A single "overall" number would have hidden both movements and averaged them into noise. This reinforces the case for measuring distinct dimensions rather than one global score.
Limitations and honest caveats
- Confounded natural experiment. ERT changed many things at once — mode, stress, assessment, expectations. You cannot cleanly attribute the score movement to "mode" alone. The pandemic is a strong existence proof that context matters, but a weak instrument for estimating the size of any single effect.
- Generalisability. Much of the published work is single-institution, single-country, and discipline-skewed toward medicine, health, and engineering, where the studies were fastest to appear. Effects differ by system, by digital-infrastructure baseline, and by how a country handled lockdown.
- Selection and non-response shifted too. Response rates and who responded changed during ERT (students under stress, with patchy connectivity). Some of the observed score movement is non-response bias, not a real change in perceived quality.
- The counterfactual is unobservable. We never see how the same instructor''s "normal" 2020 would have looked. Pre/post comparisons assume the rest of the world held still, which it did not.
- Recency. Post-pandemic "recovery" of scores may reflect adaptation, changed expectations, or grade normalisation as much as any return to a baseline.
The correct posture is humility: pandemic SET data is best used to understand the limits of SET, not to rank the people who taught through it.
How Koji incorporates this
The pandemic''s lesson is that a single Likert number, read without context, silently blames instructors for their circumstances. Koji is designed to separate the teacher from the conditions.
- AI-moderated conversational interviews probe the "why". When a student rates interactivity low, Koji''s interview engine can ask why — surfacing "the lab moved online and we couldn''t use the equipment" rather than leaving a bare 2/5 that looks like a teaching failure. That distinction is exactly what a personnel committee needs and a number cannot give.
- Structured questions that separate dimensions. Because Koji collects distinct
scaleitems for support, clarity, interactivity, and assessment — plusopen_endedcontext — the "warmth up, interactivity down" pattern shows up as two honest signals, not one averaged one. - Bias-aware, context-tagged reporting. Koji''s reporting is designed to flag when a cohort''s conditions (modality, disruption) differ from a comparison group, so a committee sees the asterisk rather than a naked mean. It is designed to mitigate misattribution, not to certify a score as clean.
- Automatic thematic analysis at scale. Across a disrupted term, thousands of open comments name the real constraints — connectivity, isolation, missing placements. Koji clusters these so the institution can act on the causes (resourcing, support) instead of penalising the symptom (a dip in one score).
Framed honestly: Koji cannot remove the confounds the pandemic exposed, but it is built to make them visible and to keep circumstance-driven scores out of decisions they should never drive. Koji''s core research platform at koji.so applies the same context-probing interview engine to customer and product research, where "did the product get worse, or did the user''s situation change?" is the identical inference problem.
Related resources
- Do Online Courses Get Lower Evaluations? Course Modality as a Confound
- Evaluating Online and Blended Teaching: The Community of Inquiry Framework
- Did the Teaching Change Actually Work? Interrupted Time Series
- Course Evaluation Evidence for Online/Distance Learning QA (EADTU)
- Do Student Evaluations Measure Learning? The Uttl Meta-Analysis
- Even Fair Evaluations Can Rank Unfairly: Misclassifying Instructors
References
- Author collective (2025). Student evaluation of teaching prior, during and after the COVID-19 pandemic. Frontiers in Education, 10, 1593000. https://doi.org/10.3389/feduc.2025.1593000
- Author collective (2022). Educational model transition: Student evaluation of teaching amid the COVID-19 pandemic. Frontiers in Education, 7, 991654. https://doi.org/10.3389/feduc.2022.991654
- Uttl, B., White, C. A., & Gonzalez, D. W. (2017). Meta-analysis of faculty''s teaching effectiveness: Student evaluation of teaching ratings and student learning are not related. Studies in Educational Evaluation, 54, 22–42. https://doi.org/10.1016/j.stueduc.2016.08.007
- Hodges, C., Moore, S., Lockee, B., Trust, T., & Bond, A. (2020). The Difference Between Emergency Remote Teaching and Online Learning. EDUCAUSE Review. https://er.educause.edu/articles/2020/3/the-difference-between-emergency-remote-teaching-and-online-learning
Related articles
Course Evaluation Evidence for E-xcellence & Online/Distance Learning QA (EADTU)
How to turn student feedback into quality-assurance evidence for online, open and flexible education under the EADTU E-xcellence benchmarks. Maps E-xcellence quality criteria to concrete, standardized course-evaluation outputs for blended and online programmes.
Even Fair Evaluations Can Rank Unfairly: Why Course-Evaluation Scores Misclassify Instructors
A simulation-based finding that even unbiased, reliable, and valid course-evaluation scores produce high misclassification rates when used to rank or compare instructors — and what it means for tenure, promotion, and merit decisions.
Did the Teaching Change Actually Work? Interrupted Time Series for Course-Evaluation Trends
Comparing this year''s evaluation mean to last year''s cannot tell you whether a curriculum redesign worked. Interrupted time series with segmented regression can — here is how to apply it, and where it breaks.
Do Student Evaluations Measure Learning? The Uttl Meta-Analysis Revisited
A research-grounded reading of the Uttl, White & Gonzalez (2017) meta-analysis on the SET–learning relationship, with implications for European course evaluation and quality-assurance policy.