New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Methodology9 min read

Do Incentives Fix Course-Evaluation Response Rates? What the Evidence Actually Shows

Lotteries, prize draws and bonus marks reliably lift how many students respond - but a landmark study found lottery incentives "neither alleviate nor increase" the bias in who responds. Raising quantity is not the same as fixing representativeness. Here is how to tell them apart.

Koji Education Team

Product ยท August 18, 2026

Every quality office facing a 30% response rate reaches for the same lever: an incentive. A prize draw, a voucher, a couple of bonus marks. The intuition is sound - more responses feel safer. But the evidence draws a sharp and often-missed distinction between two different problems, and incentives solve only one of them.

The short answer (BLUF): Incentives - especially prize draws and grade-linked rewards - reliably increase the number of students who respond. What they generally do not do is fix who responds. In the most-cited study on the topic, lottery incentives raised response rates but "neither alleviate nor increase" the demographic bias between respondents and the population. That matters because a course-evaluation dataset can be simultaneously larger and no more representative - and non-response bias, not sample size, is usually the real threat to validity. Grade-linked incentives push participation highest of all, but raise a genuine coercion problem. The most effective non-coercive lever is not an incentive at all: giving students protected time to respond.

Two problems that look like one

When people say "our response rate is too low," they usually mean two things at once:

  1. Precision - a small sample gives noisy, unstable estimates, especially for small classes. More responses genuinely help here.
  2. Representativeness - the students who bother to respond may differ systematically from those who do not. This is non-response bias, and it is the one that can invert your conclusions.

Incentives target the first problem. They are far weaker on the second - and the second is the one that should keep a quality officer awake. A course could jump from 30% to 70% response and still over-represent the most satisfied or the most aggrieved students. When the students who stay silent differ from those who speak in ways related to their rating, the data are what statisticians call missing not at random - the hardest kind of missingness to fix after the fact.

What the evidence says, incentive by incentive

Lottery / prize-draw incentives. These are the most common because they are cheap: one voucher covers the whole cohort. The foundational study here is Porter and Whitcomb's randomised work on lottery incentives in student surveys, which found that lotteries produced a modest lift in response - but, critically, that respondents still differed from the population on background characteristics, and the lottery "neither alleviate[d] nor increase[d]" that bias (Porter & Whitcomb, 2003, Research in Higher Education). In other words: you buy more responses, not a more representative sample. Larger surveys of monetary incentives echo the participation lift - a systematic review and meta-analysis of 46 randomised trials confirms monetary incentives increase survey involvement (Zhang et al., 2023, review of 46 RCTs) - but the representativeness question is the one course evaluation lives or dies on.

Grade-linked incentives. These are the most effective at driving raw participation. When completion is tied to a small number of bonus points or to releasing grades, reported response rates climb into the 90-100% range - a grade-incentive study in nursing education reported rates at that level once participation was rewarded (Int. Journal of Nursing Education Scholarship, 2018). But there is a serious catch. Rewarding or gating on completion raises a coercion problem: several institutions explicitly advise against grade-altering incentives because students may reasonably perceive them as compulsion, which corrodes both the ethics and the honesty of the exercise. It also collides with the argument for voluntary rather than mandatory evaluation: a student pressured into responding may satisfice, straightline or answer resentfully, degrading data quality even as the count rises.

Protected class time. The least glamorous lever is often the strongest and the cleanest. Setting aside a few minutes of class for students to complete the evaluation on their own devices consistently produces some of the largest gains of any single intervention - substantially outperforming prize draws - without the coercion of grade-linking. It works because it removes a friction (forgetting) rather than dangling a reward, and it does not selectively attract only the reward-motivated.

Nudges and messaging. Reminders, personalised appeals and framing effects also raise rates and are complementary to everything above; we cover them separately in do nudges increase response rates?. The evidence there is that well-designed communication helps at the margin, but rarely rescues a badly-timed or badly-designed instrument on its own.

The trap: a bigger sample that is just as biased

Here is the failure mode to guard against. Suppose incentives lift your response rate from 35% to 65%. The dashboard turns green, the mean tightens, everyone relaxes. But if the extra 30 percentage points came disproportionately from, say, the reward-motivated or the already-engaged, your estimate of the course's quality can be exactly as wrong as before - now with a false sense of security because n is larger. Sample size shrinks the confidence interval around a biased point estimate; it does not move the estimate toward the truth. This is why serious practice pairs any response-rate push with a non-response analysis: compare respondents to the known population on the variables you do have (grade distribution, attendance, demographics, programme), and, where the gap is material, apply weighting or post-stratification rather than pretending the raw mean is representative.

"But critics argue..." - the strongest objections

"Any response is better than none - a low rate is the real scandal." Partly true. A 15% response rate is genuinely uninterpretable, and getting off the floor matters. But "more is always better" quietly assumes the additional respondents look like the missing ones. When they do not, you have traded an obviously-weak dataset for a deceptively-strong-looking one. The goal is not maximum n; it is representative n.

"Grade incentives get us to 95% - surely that solves representativeness by brute force." Near-census response does dramatically reduce non-response bias, and that is a real argument in its favour. The problem is the quality of the coerced responses and the ethics of compelling them. A student who clicks through to unlock a grade may not be giving you signal at all - and survey fatigue means every forced, low-effort completion trains students to disengage from the next one. High participation bought with careless answers is not a win.

"So we should not use incentives at all?" No. Use the cheap, non-coercive ones - protected class time, good reminders, transparency about how feedback is used - and treat prize draws as a mild supplement, not a fix. Just do not let a higher number persuade you that the representativeness problem has gone away. It usually has not.

Reframing the problem: quality per response, not responses per class

The deepest point is that the entire response-rate panic is partly an artefact of the instrument. A five-minute Likert form yields so little information per respondent that you need a large fraction of the class just to say anything. If each response carried more signal, the marginal value of the fiftieth respondent would fall, and the pressure to inflate counts by any means necessary would ease.

This is where Koji for Education changes the arithmetic. Because Koji runs an AI-moderated conversational interview rather than a static form, each participating student contributes far richer, probed, thematically-analysable evidence - so the information yield per response is higher, and a representative sample becomes more valuable than an unrepresentative census. Koji supports formative, mid-cycle collection (feedback while there is still time to act, which is itself a non-coercive reason to respond), closing-the-loop action tracking so students can see their input mattered - the single most durable driver of voluntary participation - and built-in reporting that surfaces who responded so you can run the non-response checks this article argues for. The same conversational engine underpins general customer and user research on the main Koji platform. None of this eliminates non-response bias - nothing does - but it attacks the root cause (thin data per response) instead of papering over it with a prize draw.

Chasing the response-rate number is easy. Chasing a representative, honest, information-rich sample is the harder and more defensible goal - and the two are not the same.

Related reading