Online vs. Paper Course Evaluations: Do Lower Response Rates Mean Worse Data?
Online course evaluations consistently draw lower response rates than in-class paper forms, but the research shows the resulting scores are largely equivalent. Here is what Dommeyer, Nulty, and Stowell actually found, and what an adequate response rate really requires.
Koji Education Team
Product
Answer
Moving course evaluations online almost always lowers the response rate — typically from around 70–80% on paper to 40–55% online — but a generation of randomized and quasi-experimental studies finds that the mean ratings themselves are statistically equivalent across the two modes. The practical risk of online administration is therefore not biased scores but thin scores: small classes can fall below the number of respondents needed for a trustworthy class average. The defensible strategy is to keep the convenience and richer text-capture of online administration while actively engineering the response rate upward and reporting how representative each result is.
What the research says
The foundational comparison is Dommeyer, Baum, Hanna and Chapman (2004), who collected teaching evaluations for matched course sections using either in-class paper forms or online forms. The in-class sections returned roughly 75% of students; the online sections returned roughly 43%. Crucially, when they compared the evaluation scores themselves, the online and in-class means did not differ significantly — and that held even across several incentive conditions designed to lift online participation. The mode changed who bothered to respond at the margin, but it did not move the average rating of the instructor.
Stowell, Addison and Smith (2012) strengthened this conclusion with a stronger design. Across 32 instructors and 2,057 students, faculty teaching two or more sections of the same course had at least one section randomly assigned to online evaluation and at least one to paper. Online response rates were again significantly lower, yet there were no significant differences in the overall ratings, in the number of written comments, or in the valence (positive/neutral/negative) of those comments. They also found no significant differences between online and paper respondents in sex, class standing, or expected grade — evidence that, despite lower turnout, the online respondents were not a skewed subset of the class.
The third pillar is Nulty (2008), whose review reframed the question from "is online worse?" to "what response rate is enough?" Nulty synthesized response-rate data and applied sampling logic to estimate the proportion of a class you must hear from before the result is adequate for accountability versus improvement purposes. His central, widely-cited point is that the required response rate rises sharply as class size falls: a large lecture can be well represented by a modest percentage, while a seminar of fifteen needs a much higher fraction of students to respond before the average is trustworthy. Nulty also catalogued the practical levers — timing, reminders, in-class time set aside, instructor endorsement, and clear communication of how results are used — that reliably raise participation.
Taken together, the literature supports a precise claim: the online/paper choice is a response-rate problem, not a validity problem. The mode does not systematically inflate or deflate scores; it changes how many voices you collect, which matters most in small classes.
Why it matters for course evaluation in practice
For a quality-assurance office, three implications follow.
1. Do not "correct" for mode. Because online and paper means are equivalent, there is no methodological basis for adjusting online scores upward or treating a year of online data as non-comparable to an earlier paper baseline. Trend lines can cross the paper-to-online transition without a discontinuity caused by the mode itself.
2. Treat small classes differently from large ones. A 38% response rate is fine for a 300-student lecture and dangerous for a 12-student seminar. The same percentage carries completely different evidential weight depending on the denominator. QA dashboards that flag results purely by a fixed percentage threshold (e.g. "below 50% = unreliable") will simultaneously over-trust small classes and under-trust large ones. The right gate is a minimum number of respondents, informed by the reliability evidence (see our companion piece on how many responses you need).
3. Invest in participation, not in mode-switching. Since the data quality problem is turnout, the return on effort comes from the Nulty levers — protected in-class time to complete the survey, instructor encouragement framed around "closing the loop," well-timed reminders before the exam period, and transparent communication that past feedback led to real changes. These move response rates more than any choice of platform.
Limitations and honest caveats
A critical reader should hold these findings with appropriate caution.
- Equivalence is about averages, not every item. The studies show no significant mean difference, but a non-significant result is not proof of identity, and online administration can still change the texture of open-text comments (students typing at home may write longer or blunter comments than students scribbling in the last two minutes of class). The "no difference in comment valence" finding in Stowell et al. is reassuring but specific to their instrument and population.
- Generalizability. Most of this evidence is North American and pre-dates the near-universal shift to smartphone completion. Response-rate norms have drifted downward sector-wide since 2010, so the absolute rates in Dommeyer (2004) should not be read as today's benchmarks.
- Non-response can still bias results even when observed covariates match. Stowell et al. found no difference on sex, standing, or expected grade — but unobserved factors (how strongly a student feels, positively or negatively) can still drive who responds. Equivalence of means in these studies is encouraging evidence against strong non-response bias, not a guarantee it is absent in your context. This is why selection effects deserve separate scrutiny; see selection bias in course evaluations.
- Incentives have side effects. Grade-based incentives raise response rates but raise ethical and validity questions of their own (coercion, and the possibility that incentive structures attract a different respondent mix). Dommeyer tested incentives and still found equivalent scores, but institutions should weigh the governance cost.
How Koji incorporates this
Koji for Education is built for online, conversational administration — and is designed to mitigate the one real weakness that the research identifies in that mode: thin participation in smaller cohorts.
- Response-aware reporting. Rather than printing a class mean regardless of turnout, Koji reports results alongside the number of respondents and the share of the cohort reached, so a programme director can see at a glance whether a seminar result rests on four voices or fourteen. This operationalizes Nulty's core warning that the adequate response rate depends on class size.
- Engineered participation. Koji's reminder cadence, mobile-first completion, and the ability to run evaluation as a short AI-moderated conversation are aimed squarely at the turnout problem rather than the (non-existent) mode-bias problem. Because the interview adapts to each student, it tends to sustain engagement better than a static form left open in a portal.
- Richer-than-Likert capture, online. One genuine advantage of online administration is that it removes the time pressure of the last minutes of class. Koji uses that headroom: its AI-moderated interview probes beyond a single Likert number with structured follow-ups (open_ended, scale, single_choice, multiple_choice, ranking, yes_no), and applies automatic thematic analysis to the resulting text. This is designed to convert the online mode from a liability (lower counts) into an asset (deeper, codable responses per student).
- Triangulation across cohorts and cycles. Where a single small class is under-powered, Koji supports triangulation across sections and across mid-cycle and end-of-cycle collection, so a programme is not forced to over-interpret one thin result.
These mechanisms are designed to mitigate low-turnout risk; they do not eliminate non-response bias, and Koji reports the data needed for a reviewer to judge representativeness themselves. Koji's core research platform at koji.so applies the same AI-moderated interview engine to product and customer research, where online response quality is an equally central concern.
Putting it into practice: a response-rate playbook
For institutions transitioning to — or already running — online evaluation, the research converts into a concrete operating playbook. None of these levers depend on switching back to paper; they all attack the turnout problem the evidence isolates.
- Protect time, even online. The single biggest historical advantage of paper was the captive in-class moment. Recreate it: ask instructors to set aside five minutes for students to complete the evaluation on their phones during a session, then leave the room. This recovers much of the response-rate gap without reintroducing paper logistics.
- Sequence reminders before the exam crush. Open the window early and send timed reminders before assessment deadlines consume students'' attention. Nulty''s levers are about salience and timing, not nagging volume.
- Close the loop visibly. Students respond more when they believe it matters. Publish a short "you said, we did" each cycle so the next cohort sees that feedback produced change — the most durable response-rate intervention there is.
- Report the denominator, always. Pair every result with the respondent count and cohort share so small classes are read with appropriate caution and large classes get the trust their numbers earn.
- Avoid coercive incentives. Grade-linked incentives lift rates but raise governance and validity questions; favour intrinsic framing (impact, transparency) over compulsion.
Adopted together, these moves let an institution keep the analytical and text-capture advantages of online administration while neutralising its one evidenced weakness — thin counts in the smallest, highest-touch classes.
Related Resources
- How Many Responses Do You Need for a Reliable Course Evaluation?
- Selection Bias in Course Evaluations: What Goos and Salomons Found
- Class Size and Student Evaluations: What Bedard and Kuhn Found
- How Many Scale Points Should a Course-Evaluation Question Have?
- What Can Open-Text Student Comments Tell You That Likert Scores Cannot?
- Turning Student Feedback into ESG / ENQA Accreditation Evidence
References
- Dommeyer, C. J., Baum, P., Hanna, R. W., & Chapman, K. S. (2004). Gathering faculty teaching evaluations by in-class and online surveys: their effects on response rates and evaluations. Assessment & Evaluation in Higher Education, 29(6), 611–623. https://doi.org/10.1080/0260293042000227245
- Nulty, D. D. (2008). The adequacy of response rates to online and paper surveys: what can be done? Assessment & Evaluation in Higher Education, 33(3), 301–314. https://doi.org/10.1080/02602930701293231
- Stowell, J. R., Addison, W. E., & Smith, J. L. (2012). Comparison of online and classroom-based student evaluations of instruction. Assessment & Evaluation in Higher Education, 37(4), 465–473. https://doi.org/10.1080/02602938.2010.545869
Related articles
Turning Student Feedback into ESG / ENQA Accreditation Evidence
A buyer's guide mapping the ESG 2015 internal quality assurance standards to concrete, accreditation-ready evidence you can generate from student feedback — and how AI-moderated evaluation closes the loop.
How Many Responses Do You Need for a Reliable Course Evaluation?
Reliability of a class-average rating rises steeply with the number of students who respond. Generalizability theory — from Marsh''s syntheses to Dzakadzie (2026) — shows raters matter more than items, and a class of 12 cannot be evaluated like a class of 120.
What Can Open-Text Student Comments Tell You That Likert Scores Cannot?
A research-grounded guide to open-ended student comments in course evaluation: what Alhija and Fresko found about who writes them and what they contain, how thematic analysis surfaces issues numbers miss (Stupans et al.), how rare abusive comments actually are (Tucker), and how to turn free text into reliable evidence.
How Many Scale Points Should a Course-Evaluation Question Have?
What the measurement literature — Preston & Colman (2000), Weng (2004), Dawes (2008) — says about the optimal number of response categories on rating scales, and why the answer for course evaluation is not just a number but a question about what a Likert item can and cannot capture.