What Actually Raises Course-Evaluation Response Rates: The Evidence, Not the Folklore
Most response-rate advice is folklore. The controlled evidence points to a small number of interventions that reliably work — and one that matters more than any nudge. Here is what the research supports.
Koji Education Team
Product · July 4, 2026
When online course evaluations replaced paper, response rates fell — and universities have been chasing them ever since with a grab-bag of tactics, most of which have never been tested. The good news is that the research literature is clearer than the folklore suggests. A handful of interventions reliably move response rates; a few popular ones do not; and the single most powerful lever is not a nudge at all. This piece separates what the controlled evidence supports from what merely sounds plausible — because chasing response rates with the wrong tools wastes effort and, worse, can corrode the trust the whole exercise depends on.
First, how high does the rate actually need to be?
Before optimising response rates, know your target — and it is not a flat 70%. The foundational reference is Nulty (2008), The adequacy of response rates to online and paper surveys, which showed that the response rate needed for a defensible result depends heavily on class size. Small classes need a much higher proportion of students to respond to reach the same sampling confidence as large ones: a class of 20 may need well over a third — on Nulty's liberal criterion, roughly half — of its students to respond, while a large lecture of several hundred can be adequately represented by a far smaller fraction. Online surveys, Nulty showed, generally require higher response rates than paper to yield equivalent confidence, precisely because they tend to attract fewer responses.
This reframes the whole problem. A blanket "we need 70% everywhere" target is both unattainable in large classes and insufficient in small ones. The right target is class-size-dependent, and — as we argue in generalizability theory and the dependability of course evaluation — ideally derived from the dependability your instrument actually requires. Chasing an arbitrary universal number is the first folklore to drop.
What the controlled evidence says works
1. Reminders — the most reliable lever. Across studies, structured reminder sequences are the workhorse. Institutional data shows that layered announcement-and-reminder strategies can double online response rates — reported jumps from around 35% to around 70% are documented — and that multiple automated reminders lift rates substantially. Reminders are cheap, automatable, and evidence-backed. The main caveat is diminishing returns and irritation: there is a point past which more reminders harvest fatigue, not responses, a tension we examine in the response-rate crisis universities built for themselves.
2. Explaining how the feedback is used — the closing-the-loop effect. Telling students what changed because of previous cohorts' feedback is one of the better-supported interventions. Students who believe their input is read and acted on respond at higher rates; students who suspect it vanishes into an archive do not. This is not just a response-rate tactic — it is the same action gap that undermines the credibility of the whole system. Response rates are partly a referendum on whether the institution listens.
3. In-class protected time. Setting aside time during a session for students to complete the evaluation on their devices consistently outperforms leaving it to students' own time. It works because it removes the friction and forgetting that kill voluntary online completion — the mechanism behind the online-versus-paper gap in the first place.
4. Personalisation. Personal emails from the instructor or department, rather than generic system-generated requests, modestly but reliably improve participation. The signal that a real person wants this student's view matters.
What the evidence is lukewarm or divided on
Incentives are genuinely mixed. The randomised-controlled-trial evidence — much of it from adjacent survey fields like physician and health-professional surveys — is inconsistent. Some trials find small unconditional incentives (a token gift) lift response; others find no statistically significant effect from an incentive at all. Within course evaluation, class-wide incentives (for example, releasing results or a small reward once the class hits a response threshold) tend to perform about as well as individual point-based ones, without the pedagogical awkwardness of grading students for compliance. The honest reading: incentives sometimes help, the effect is small and context-dependent, and they carry a real risk of coercion.
Grade-withholding and mandatory completion raise the sharpest concerns. They can mechanically lift rates, but at a cost to data quality and ethics we examine in should course evaluations be mandatory? — coerced responses import straightlining and resentment, trading a better quantity for a worse quality.
The strongest counterargument — "You are optimising the wrong number"
Here is the objection a methodologist will raise, and it is correct: A higher response rate is not automatically a more representative one. If reminders and incentives disproportionately recruit the already-engaged — or the already-aggrieved — you can raise the percentage while making non-response bias worse.
This is the crucial caveat, and it reframes everything above. Response rate is a proxy for what we actually care about: response representativeness. The two usually move together — higher rates generally shrink the room for non-response bias, which is why non-response bias is the real threat behind low response rates — but not always. An intervention that lifts the rate by mobilising only the most satisfied or most furious students can worsen the very bias it appears to fix.
Two implications follow. First, prefer interventions that reduce friction for everyone (in-class time, reminders, personalisation) over interventions that selectively motivate a subgroup (issue-driven appeals, or incentives that appeal only to some). Friction-reducers pull in the ambivalent middle — precisely the students most likely to be missing. Second, always check who responded, not just how many. Compare respondent demographics and grade distributions against the enrolled population. A 75% response rate that over-represents high-achievers is worse than a well-characterised 50% you can adjust for. Optimising the headline number while ignoring representativeness is folklore dressed as rigour.
What to do with this
- Set class-size-dependent targets from Nulty's logic and your own dependability requirements, not a universal 70%.
- Automate a measured reminder sequence — enough to lift rates, not so many that you manufacture fatigue.
- Close the loop visibly. Show students what previous feedback changed. This raises rates and fixes the underlying problem.
- Protect in-class time for completion where feasible; it is among the most effective friction-reducers.
- Treat incentives cautiously and prefer low-coercion class-wide forms; do not rely on them as a primary lever.
- Audit representativeness every cycle. Response rate is the proxy; representativeness is the target.
Where Koji fits
Two of the best-supported levers — closing the loop and reducing friction — are exactly where Koji for Education is designed to help. Koji's closing-the-loop action tracking makes visible what changed as a result of feedback, turning the most durable response-rate driver into a built-in feature rather than a manual campaign. Its AI-moderated conversational format reduces the friction and monotony of a long Likert grid — a shorter, adaptive, genuinely engaging interview is easier to start and finish, which lifts completion without coercion. And because Koji captures richer information per response through probing, it lessens the pressure to chase raw response counts: a smaller number of deep, well-characterised conversations can carry more dependable signal than a larger pile of thin ticks, which is the point of measuring representativeness rather than headline rate. Koji mitigates the response-rate problem; it does not claim to abolish non-response bias, which remains a design responsibility for every institution.
The same conversational engine powers the main Koji platform for customer and user research, where completion and representativeness are the identical twin challenges.
The bottom line
Response-rate advice is mostly folklore, but the evidence underneath it is not. Reminders, closing the loop, protected in-class time and personalisation reliably help; incentives are a weak and risky lever; and mandatory completion buys quantity at the price of quality. Above all, remember what the number is for: a high response rate that misrepresents the class is worse than a modest one you understand. Chase representativeness, not just the percentage — and make the institution visibly worth responding to.
Want a feedback system that earns responses by closing the loop, not by coercing them? Explore Koji for Education.