What Actually Raises Course-Evaluation Response Rates? The Experimental Evidence
Online course evaluations chronically under-perform paper. We review the experimental evidence — Dommeyer''s grade-incentive trials and Nulty''s adequacy thresholds — on what genuinely lifts response rates, what it costs in data quality, and how to hit a defensible rate without coercion.
Koji Education Team
Product
In brief: Online course evaluations typically return far lower response rates than paper (Dommeyer et al. 2004 found ~43% online vs ~75% paper), and low rates threaten representativeness. The experimental evidence is clear about what works: a small grade incentive closes almost the entire gap (raising online rates to ~87%), while reminders, protected in-class time, and visible follow-through give reliable but smaller lifts. Nulty (2008) provides defensible thresholds. The catch: grade incentives raise ethical and validity questions, so the goal is a representative sample obtained without coercion — not a rate maximised at any cost.
The question this answers
Two facts sit in tension. First, low response rates are the most common technical complaint about online course evaluations, because a small, self-selected sample may not represent the class — see our analysis of non-response bias. Second, the institution that wants more responses has a menu of interventions of wildly different effectiveness, cost, and ethical weight. Which ones actually move the rate, by how much, and at what price to data quality? This article reviews the experimental evidence and turns it into defensible practice.
What the research says
The anchor is Dommeyer, Baum, Hanna & Chapman (2004), Gathering Faculty Teaching Evaluations by In-Class and Online Surveys: Their Effects on Response Rates and Evaluations, Assessment & Evaluation in Higher Education, 29(5), 611–623. The authors ran a controlled comparison across course sections, contrasting paper and online administration and testing several incentives. Key findings:
- Online lagged paper substantially. Overall, online sections returned about a 43% response rate against roughly 75% for paper — a gap large enough to worry about representativeness.
- A tiny grade incentive nearly erased the gap. When students were offered a very small grade incentive (a quarter of one percent added to the course grade), online response rates jumped to around 87% — statistically indistinguishable from paper (~87%) and far above the other treatments and the control.
- Non-grade incentives helped less. Other nudges (e.g. assurances, informational appeals) produced smaller, less reliable lifts than the grade incentive.
- The evaluation scores themselves were largely unaffected by mode, suggesting the ratings were comparable; the problem with online was participation, not a shift in what respondents said.
The second anchor sets the target. Nulty (2008), The Adequacy of Response Rates to Online and Paper Surveys: What Can Be Done?, Assessment & Evaluation in Higher Education, 33(3), 301–314, derives, from sampling theory, how high a response rate must be for the data to be adequate for accountability versus improvement, and how this depends on class size. Nulty's central, counter-intuitive message: small classes need a much higher response rate than large ones to be representative — a 20-student class may need 50–80% while a 500-student lecture is adequately represented by a far smaller percentage. He also catalogues practical levers: explaining the purpose, protecting time, reminders, assurances of anonymity, and crucially, closing the loop so students see that past feedback changed something.
Corroborating work fills in the mechanisms. Reviews of online evaluation response rates (e.g. Nulty 2008; subsequent systematic reviews) consistently rank protected in-class time to complete the survey, multiple reminders, instructor endorsement, and demonstrated use of results among the most reliable non-coercive levers — while confirming that explicit grade incentives are the single most powerful and the most ethically fraught.
Why it matters for course evaluation in practice
The evidence reframes the response-rate problem in three useful ways:
- The mode switch, not student apathy, is the main culprit. Online administration is the variable that dropped rates; the fix is therefore administrative design, not lecturing students about civic duty. Restoring something paper had for free — protected, in-context time — recovers much of the loss.
- Set the target by class size, not a blanket percentage. Nulty shows a flat "we need 70% everywhere" rule is wrong: it is unattainable and unnecessary in large lectures and insufficient in small seminars. Adequacy thresholds should scale with enrolment. See how many responses you actually need.
- Incentives trade rate for validity. A grade incentive maximises participation but couples the evaluation to the grade — the very thing that distorts ratings elsewhere — and raises a coercion concern: a student pressured to respond may satisfice or straight-line, adding bodies without adding signal. The defensible aim is a representative sample, not the highest possible number.
The strongest single lever that carries no ethical cost is visible follow-through. When students have seen that last year's feedback changed the course, they respond more — which ties response-rate strategy directly to closing the feedback loop.
Limitations and honest caveats
A rigorous reader should hold several reservations:
- Generalisability and era. Dommeyer et al. (2004) predates ubiquitous smartphones and modern survey UX; absolute rates today differ, and the paper-versus-online gap may be narrower where mobile completion is frictionless. The relative ordering of interventions has held up better than the exact percentages.
- Grade incentives are confounds, not free wins. Tying participation to grades can pressure reluctant students into the sample and may itself nudge responses, and many institutions and ethics frameworks prohibit it. A higher rate bought this way is not automatically a better sample.
- Higher rate is not the same as lower bias. If the non-respondents differ from respondents (e.g. disengaged students), raising the rate reduces bias only if the newly recruited respondents resemble the missing ones. An incentive that recruits the already-engaged adds little. Response rate is a proxy for representativeness, not a guarantee of it.
- Reminders have diminishing returns and fatigue costs. Past a point, more reminders annoy students and feed the survey fatigue that depresses rates across an over-surveyed programme.
The honest conclusion: protected time, well-timed reminders, instructor endorsement, and demonstrated use of results are the evidence-based, ethically clean levers; grade incentives are powerful but should be approached with caution and governance, and no intervention substitutes for checking that the achieved sample is actually representative.
How Koji incorporates this
Koji for Education is designed to lift meaningful participation — recovering what online administration lost — without leaning on coercive grade incentives.
- Low-friction conversational completion on any device. Koji's AI-moderated interview runs in the browser on a phone, so the practical barrier that depressed early online rates is minimised. Protected in-class time — the most effective ethical lever — works cleanly because students can complete the conversation in a few minutes where they sit.
- Built-in, configurable reminders. Koji automates multiple, well-timed reminders to non-respondents and stops once a student has participated, applying the evidence on reminders while avoiding the over-messaging that drives fatigue.
- Class-size-aware adequacy reporting. Reflecting Nulty (2008), Koji reports response rates against a target that scales with enrolment and flags when a small cohort has not yet reached a representative threshold — so a quality office reads adequacy correctly instead of applying a misleading flat percentage.
- Closing-the-loop tools that raise future rates. Koji's action-tracking and "you said, we did" reporting make prior follow-through visible to students, directly exploiting the most durable non-coercive driver of participation. Higher rates next cycle come from demonstrated use, not pressure.
- Representativeness checks, not just counts. Because rate is only a proxy, Koji's reporting is designed to surface who responded and to caution against over-reading a sample that may not represent the class — supporting the principle that the goal is a representative, honest sample rather than a maximised number.
Koji is designed to mitigate low online participation through frictionless completion, smart reminders, and visible follow-through; it deliberately does not rely on grade coercion, and it frames response rate as evidence of representativeness to be checked rather than a target to be maximised. The same engine powers general user and customer research at koji.so, where recruiting a representative sample without incentive-induced bias is an equally central concern.
The practical takeaway for a quality office
A defensible response-rate strategy is a sequence, not a single lever. Start with the cheapest, cleanest move: protect a few minutes of in-context time so completion does not compete with everything else in a student's week. Layer on two or three well-timed reminders that stop the moment a student responds, and have the instructor visibly endorse the survey. Set the adequacy target by enrolment using Nulty's logic — high for seminars, lower for large lectures — rather than a flat institutional percentage. Reserve grade incentives for last, if at all, and only under explicit governance, because they buy participation at the cost of coupling the evaluation to the grade. Finally, audit who responded: if the achieved sample skews toward the already-engaged, a high headline rate is still a biased one. The objective throughout is a representative sample obtained honestly, not a leaderboard number.
Related resources
- Online vs. Paper Course Evaluations: Do Lower Response Rates Mean Worse Data?
- How Many Responses Do You Need for a Reliable Course Evaluation?
- Selection Bias in Course Evaluations
- Survey Fatigue: Why Over-Surveying Students Wrecks Response Rates
- Does Closing the Feedback Loop Actually Matter?
References
- Dommeyer, C. J., Baum, P., Hanna, R. W., & Chapman, K. S. (2004). Gathering faculty teaching evaluations by in-class and online surveys: Their effects on response rates and evaluations. Assessment & Evaluation in Higher Education, 29(5), 611–623. https://doi.org/10.1080/02602930410001689171
- Nulty, D. D. (2008). The adequacy of response rates to online and paper surveys: What can be done? Assessment & Evaluation in Higher Education, 33(3), 301–314. https://doi.org/10.1080/02602930701293231
- Adams, M. J. D., & Umbach, P. D. (2012). Nonresponse and online student evaluations of teaching: Understanding the influence of salience, fatigue, and academic environments. Research in Higher Education, 53(5), 576–591. https://doi.org/10.1007/s11162-011-9240-5
Related articles
How Many Responses Do You Need for a Reliable Course Evaluation?
Reliability of a class-average rating rises steeply with the number of students who respond. Generalizability theory — from Marsh''s syntheses to Dzakadzie (2026) — shows raters matter more than items, and a class of 12 cannot be evaluated like a class of 120.
Survey Fatigue: Why Over-Surveying Students Quietly Wrecks Your Response Rates
Porter, Whitcomb & Weitzer (2004) showed that administering multiple surveys in one year suppresses later response rates. A research-grounded guide to survey fatigue in course evaluation — what causes it, what the evidence shows, and how to design around it.
Does Closing the Feedback Loop Actually Matter? The Evidence on Acting on Student Evaluations
Universities are good at collecting student feedback and bad at acting on it visibly. The research — Watson (2003), Leckey & Neill (2001), Shah et al. (2017) — shows that failing to close the loop drives the scepticism and declining response rates that quietly destroy your evaluation data.
Online vs. Paper Course Evaluations: Do Lower Response Rates Mean Worse Data?
Online course evaluations consistently draw lower response rates than in-class paper forms, but the research shows the resulting scores are largely equivalent. Here is what Dommeyer, Nulty, and Stowell actually found, and what an adequate response rate really requires.