Response Rates and Non-Response Bias in Course Evaluations
When response rates fall, the question is not just "is the sample big enough?" but "who stopped answering?" Non-response bias can distort a course evaluation more than any single survey item. Here is what the evidence says and how to fix the root cause.
Koji for Education
Research & Editorial Team · June 1, 2026
The short answer: Falling response rates are not only a sample-size problem — they are a bias problem. The students who choose not to respond are systematically different from those who do, so a low-response evaluation can be skewed in ways no margin-of-error figure reveals. The move to online evaluation made the problem worse, and chasing the rate with reminders treats the symptom. The durable fix is to make giving feedback worth the student's time, so participation rises because the instrument is better — not because they were nagged.
Why response rate is really about non-response bias
A response rate is easy to put on a dashboard, which is exactly why it gets misread. The number that matters is not how many students answered but whether the ones who didn't answer would have said something different. If non-responders are a random slice of the class, a smaller sample is merely less precise. If non-responders differ systematically — and they usually do — the result is biased, and collecting more of the same kind of response will not fix it.
In course evaluation the non-response mechanism is rarely random. Students with strong feelings — either delighted or aggrieved — are more motivated to respond, while the broad middle quietly opts out. Disengaged students, the very group whose experience a programme most needs to understand, are also the least likely to complete a voluntary end-of-term form. The result is a polarised, unrepresentative picture that can make a solid course look divisive or a struggling one look acceptable.
The online shift made it worse
When institutions moved from in-class paper forms to online evaluation, response rates fell. The canonical reference here is Nulty (2008), whose analysis of online versus paper-based course and teaching evaluations documented that online surveys typically achieve lower response rates than paper administered in class, and who set out practical guidance on how high a rate needs to be before the data can responsibly support accountability and improvement decisions (Nulty, 2008, Assessment & Evaluation in Higher Education 33(3):301–314).
Nulty's contribution is not a single magic threshold but a sliding scale: the required response rate depends on class size and on how stringent your inferential standard is. Crucially, smaller classes need a much higher proportion of respondents to be trustworthy. A 60% response rate in a 200-student lecture and a 60% rate in a 12-student seminar are not equivalent pieces of evidence — yet legacy systems report both as the same green tick.
Why "just send more reminders" fails
The instinctive response to a low rate is administrative pressure: automated reminders, nudge emails, withholding grade access until the form is done. These tactics can lift the number, but they often do nothing for the bias — and can worsen it. Coerced responses tend to be rushed, low-effort, and box-ticked, which degrades quality while flattering the rate. And mandatory completion gates raise legitimate ethical and data-protection concerns when feedback is supposed to be voluntary.
The reframing that matters: a low response rate is usually a signal that the instrument is not worth the student's time. A long, repetitive Likert form that students suspect no one reads earns exactly the engagement it deserves. Fix the why, and participation follows for the right reasons.
The strongest counterargument — taken seriously
"Surely a low response rate is fine if the responses are representative — and we can check that against demographics?" This is the most serious objection, and partly correct. If you can show responders and non-responders match on the variables that matter, low response is less damaging, and weighting can help. Representativeness, not raw rate, is the real target.
But two cautions keep this from being a free pass. First, you can only check representativeness on variables you happen to have — gender, programme, prior grades. Non-response bias often operates on variables you cannot observe, above all the student's underlying satisfaction, which is the very thing you are trying to measure. Matching on demographics does not guarantee matching on sentiment. Second, weighting corrects for known imbalances only and adds variance of its own. So the honest version of the counterargument is: representativeness matters more than rate, but you can rarely verify representativeness on the dimension that counts — which is why raising genuine engagement remains the safer strategy than defending a low rate after the fact.
How to raise participation for the right reasons
- Make it short and relevant. Adaptive questioning that skips what does not apply respects students' time and signals that the instrument is thoughtfully designed.
- Close the loop visibly. Students participate when they have seen feedback lead to change. "You said, we did" is the single most powerful response-rate intervention, and it is about credibility, not coercion.
- Collect at the right moment. Mid-cycle (formative) collection, while the course is still running, both improves the current cohort's experience and demonstrates that feedback is acted on — which lifts later participation.
- Prefer depth over volume of items. A short conversation that surfaces real reasoning beats a 30-item grid that students abandon halfway.
- Monitor representativeness, not just the rate. Track who is responding and treat a skewed responder profile as the warning it is.
Where Koji fits
Koji for Education attacks the root cause: it makes giving feedback genuinely worth a student's time. Instead of a static form, Koji runs AI-moderated conversational interviews — adaptive, responsive, and quick — that feel like being listened to rather than processed. Higher-quality engagement tends to draw in students who would have ignored a traditional form, which directly addresses the non-response skew toward only the strongly-opinionated.
Because Koji supports formative, mid-cycle collection and closing-the-loop action tracking, institutions can show students that feedback changes things — the most durable driver of voluntary participation. Its quality scoring flags low-effort responses so a high rate is never mistaken for high-quality evidence, and programme- and institution-level reporting lets you monitor responder representativeness, not just the headline percentage. To be precise about the claim: better engagement mitigates and surfaces non-response bias and gives you a more representative picture; it does not eliminate the fact that participation is voluntary, nor should it — coerced feedback is worse evidence, not better.
Teams that also run customer or user research outside the classroom use the same engagement-first interview engine in the main Koji platform.
What a healthy participation signal looks like
If the goal is representative engagement rather than a high number, your monitoring should change accordingly. Start by comparing the responder profile against the class composition on every variable you can observe — programme, year, gender, and broad attainment band. A response rate of 70% that draws disproportionately from high-achieving students is weaker evidence than a 45% rate that mirrors the cohort. Treat a skewed responder profile as a red flag that demands caution in interpretation, not a green tick to be celebrated.
Next, watch the shape, not just the size. A bimodal pattern — clusters at the extremes with a hollow middle — is the classic fingerprint of motivated-only response, where the contented majority stayed silent. Pair the quantitative split with the qualitative reasoning to distinguish a genuinely divisive course from a sampling artefact. Then segment by engagement where you can: the feedback of students who attended regularly carries different weight from that of students who had largely disengaged, and both are worth hearing for different reasons.
Finally, set context-appropriate expectations. Following Nulty's logic, a 12-student seminar and a 300-student lecture should not be held to the same numeric standard, and reporting that pretends otherwise misleads decision-makers. The healthiest signal of all is a trend: rising voluntary participation over successive cycles, which indicates students increasingly believe their feedback is read and acted upon. That trajectory, not any single percentage, is the metric worth managing.
The bottom line
Stop watching the response-rate gauge as if a green number meant valid data. The real risk is non-response bias — the systematic silence of the students you most need to hear — and it gets worse, not better, when you chase the rate with reminders. Build feedback that students want to give, collect it while it can still help them, and prove it changes things. Participation that rises because the instrument earned it is the only response rate worth trusting.