New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
analysis-reporting9 min read

Did COVID Lower Course Evaluations? Emergency Remote Teaching as a Natural Experiment in Confounding

What student-evaluation data from the 2020 emergency shift to remote teaching reveals about how much SET scores reflect factors outside an instructor's control — and why pandemic-era ratings need a giant asterisk in any personnel decision.

Koji Education Team

Product

In brief. When universities moved online in weeks in spring 2020, student-evaluation scores did not simply fall — they moved in different directions for different questions, disciplines, and cohorts. Studies tracking the same programmes before, during, and after the pandemic find that ratings of teacher support and empathy often rose, while ratings of interactivity and hands-on practice fell, largely independent of individual teaching skill. That divergence is the clearest natural-experiment evidence we have that a substantial part of a Student Evaluation of Teaching (SET) score measures the conditions of instruction, not the instructor. Pandemic-era scores should never be read as a clean signal of teaching quality.

Why the pandemic is a useful accident

Course evaluation researchers rarely get a controlled shock: the same instructors, the same courses, the same students, and then one variable — the mode of delivery — changes abruptly for everyone at once. The emergency remote teaching (ERT) period of 2020 was exactly that. It is not the same as planned online education (which is designed, resourced, and chosen); it was improvised under stress. That makes it a powerful lens on a long-standing question: how much of a course-evaluation score is about the teacher, and how much is about everything else?

What the research says

Studies that compare SET across pre-pandemic, pandemic, and post-pandemic periods within the same institutions tell a consistent, textured story rather than a simple "scores dropped".

A longitudinal analysis published in Frontiers in Education (2025) tracking student evaluation of teaching prior to, during, and after COVID-19 found that students gave higher ratings on items about the teacher''s professional, motivating, and supportive attitude during the pandemic than before or after — while giving lower ratings on the interactivity of practice-based courses during the same period. In other words, the same instructors were rated more warmly on empathy and more harshly on hands-on engagement, simultaneously, because the situation changed — not their competence.

A separate Frontiers in Education (2022) study of an institution''s abrupt educational-model transition found the effect was discipline-dependent: in the traditional model, the School of Engineering and Sciences showed an improvement in SET during the shift, while the School of Medicine and Health Sciences showed a decline. Practice- and clinical-heavy programmes, whose core value was hardest to deliver remotely, took the biggest hit — again, a property of the subject and its delivery constraints, not of any individual teacher.

These findings sit on top of a mature literature warning that SET scores track things other than learning. Uttl, White and Gonzalez (2017), in a re-analysis of decades of multisection studies, found essentially no meaningful correlation between SET ratings and actual learning once older studies'' small-sample artefacts were controlled — and Uttl has argued specifically (2021, 2023) that pandemic disruption compounds the existing case against using SET results to judge teaching effectiveness. The pandemic data also intersect with grade effects: analyses in this period reported the familiar moderate association between the grades students expect and the ratings they give, a confound that ERT''s grading disruptions only muddied further.

The mechanism is well understood. Ratings during ERT absorbed variance from: technology access (bandwidth, devices, quiet space), student stress and isolation, the loss of laboratory, studio, and placement experiences, delayed or degraded feedback because remote assessment was harder, and shifted expectations about what a course could deliver. None of these is a teaching-skill variable. Yet all of them landed inside the SET score.

Why it matters for course evaluation in practice

1. Pandemic-era scores are contaminated evidence for personnel decisions. An instructor whose interactivity rating fell in 2020 was, in many cases, teaching a subject that could not be made interactive over Zoom that term. Using such a score in a tenure, promotion, or contract-renewal case imports a confound the instructor could not control. Where SET already misclassifies instructors even in normal times, ERT amplifies the error.

2. Year-over-year trend lines break at 2020. Any dashboard that plots a course''s mean rating across years will show a discontinuity in 2020–21. Reading that as "the teaching changed" is a mistake; the measurement conditions changed. This is exactly the situation an interrupted time series analysis is built to handle — model the level shift explicitly rather than pretend the series is continuous.

3. It is a general lesson, not a historical footnote. The pandemic made visible a permanent property of SET: modality, resourcing, and student circumstances move scores. The same logic applies to a course forced online by a strike, a room double-booked all term, or a cohort hit by a local crisis. The modality confound did not end in 2022.

5. It exposed an equity dimension inside the score. The pandemic did not distribute its confounds evenly. Students with poor connectivity, no quiet study space, caring responsibilities, or disabilities that remote delivery served badly experienced a materially different course than their better-resourced peers — and rated it accordingly. A pooled SET mean silently averages those experiences together, so a low score can encode a resourcing and access failure the institution owns rather than a teaching failure the instructor owns. Reading pandemic (and post-pandemic online) ratings without disaggregating by student circumstance risks mistaking a digital-divide effect for a quality signal, and penalising the very instructors who taught the most disadvantaged cohorts.

4. Warmth and interactivity are separable. The pandemic''s split result — support up, interactivity down — is direct evidence for multidimensional reporting. A single "overall" number would have hidden both movements and averaged them into noise. This reinforces the case for measuring distinct dimensions rather than one global score.

Limitations and honest caveats

  • Confounded natural experiment. ERT changed many things at once — mode, stress, assessment, expectations. You cannot cleanly attribute the score movement to "mode" alone. The pandemic is a strong existence proof that context matters, but a weak instrument for estimating the size of any single effect.
  • Generalisability. Much of the published work is single-institution, single-country, and discipline-skewed toward medicine, health, and engineering, where the studies were fastest to appear. Effects differ by system, by digital-infrastructure baseline, and by how a country handled lockdown.
  • Selection and non-response shifted too. Response rates and who responded changed during ERT (students under stress, with patchy connectivity). Some of the observed score movement is non-response bias, not a real change in perceived quality.
  • The counterfactual is unobservable. We never see how the same instructor''s "normal" 2020 would have looked. Pre/post comparisons assume the rest of the world held still, which it did not.
  • Recency. Post-pandemic "recovery" of scores may reflect adaptation, changed expectations, or grade normalisation as much as any return to a baseline.

The correct posture is humility: pandemic SET data is best used to understand the limits of SET, not to rank the people who taught through it.

How Koji incorporates this

The pandemic''s lesson is that a single Likert number, read without context, silently blames instructors for their circumstances. Koji is designed to separate the teacher from the conditions.

  • AI-moderated conversational interviews probe the "why". When a student rates interactivity low, Koji''s interview engine can ask why — surfacing "the lab moved online and we couldn''t use the equipment" rather than leaving a bare 2/5 that looks like a teaching failure. That distinction is exactly what a personnel committee needs and a number cannot give.
  • Structured questions that separate dimensions. Because Koji collects distinct scale items for support, clarity, interactivity, and assessment — plus open_ended context — the "warmth up, interactivity down" pattern shows up as two honest signals, not one averaged one.
  • Bias-aware, context-tagged reporting. Koji''s reporting is designed to flag when a cohort''s conditions (modality, disruption) differ from a comparison group, so a committee sees the asterisk rather than a naked mean. It is designed to mitigate misattribution, not to certify a score as clean.
  • Automatic thematic analysis at scale. Across a disrupted term, thousands of open comments name the real constraints — connectivity, isolation, missing placements. Koji clusters these so the institution can act on the causes (resourcing, support) instead of penalising the symptom (a dip in one score).

Framed honestly: Koji cannot remove the confounds the pandemic exposed, but it is built to make them visible and to keep circumstance-driven scores out of decisions they should never drive. Koji''s core research platform at koji.so applies the same context-probing interview engine to customer and product research, where "did the product get worse, or did the user''s situation change?" is the identical inference problem.

Related resources

References