New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Methodology8 min read

Should Course Evaluations Be Mandatory? The Evidence on Coercing Response Rates

Faced with falling response rates, more institutions are tempted to make course evaluations compulsory — even withholding grades until students comply. Does coercion fix the data, or quietly corrupt it? An evidence-led look at response-rate thresholds, non-response bias, satisficing, and consent.

Koji Education Team

Product · June 9, 2026

Bottom line up front: Making course evaluations mandatory — especially by withholding grades or results until a student submits — reliably raises the response rate, but it trades one data-quality problem (non-response bias) for two others (satisficing and coerced, low-engagement answers) while raising real consent and ethics concerns. The better goal is not a high completion number but representative, thoughtful feedback, and that is earned by making feedback worth giving, not by forcing it.

Why institutions are tempted

Online course evaluations have a well-documented response-rate problem. When evaluations moved from paper-in-class to online-on-your-own-time, average response rates fell sharply; the literature typically puts them in the 30–60% range, and often lower for large cohorts (CBE review of student-rating strategies). Low rates are not just embarrassing; they threaten validity, because the students who bother to respond may not represent the class.

So administrators reach for the lever that works fastest: compulsion. Approaches range from soft ("carrot") incentives to hard ("stick") measures such as withholding final grades or transcripts until the evaluation is submitted. The stick works on the metric — completion climbs toward 100%. The question is what it does to everything else.

What "enough responses" actually requires

Before deciding whether to coerce, it helps to know the target. Nulty (2008) derived the response rates needed for online evaluations to be reasonably representative, and — importantly — the threshold depends on class size. As a rule of thumb from that work, you need roughly 58% for a class under 20, but only around 35% for a class over 50, because larger absolute numbers stabilise the estimate. The practical implication is liberating: large classes do not need anything like 100% to be trustworthy. The case for blunt, universal compulsion weakens considerably once you accept that a representative two-thirds or one-third, depending on size, is often statistically sufficient.

The hidden cost of compulsion: satisficing

Here is the part proponents of mandatory evaluation tend to skip. A response rate is a measure of quantity, not quality. When students are forced to submit to unlock a grade, a predictable share will do the minimum to get through: straight-lining every scale to the same point, leaving open boxes blank or filling them with a single word, and clicking through without engaging. Survey methodologists call this satisficing — giving a merely satisfactory rather than optimal answer — and it is exactly what coercion incentivises. You can drive the response rate to 95% and simultaneously dilute the signal with low-effort responses that look like data but are not.

So compulsion does not eliminate bias; it relocates it. Voluntary evaluations risk non-response bias (the wrong people answer). Mandatory evaluations risk satisficing bias (the right people answer badly). Neither is automatically worse, but pretending the mandatory route is bias-free is the central error.

And then there is consent

In the European context, the ethics are not optional flavour — they are legal architecture. Open-text comments and identifiable responses are personal data; in many designs the lawful basis leans on the voluntariness of participation. Conditioning a student's grade on the surrender of their feedback sits uneasily with the idea of freely given engagement, and several institutions and commentators flag mandatory completion as coercion precisely because of the power imbalance between student and institution (CBE review). Coercion is also pedagogically self-defeating: it teaches students that feedback is a toll to be paid, not a voice that is heard.

But low response rates make the data useless — isn't mandatory the lesser evil?

This is the strongest counterargument, and it has a real premise: a 12% response rate genuinely can be unrepresentative and unusable. But the conclusion does not follow. Three points answer it.

First, as Nulty shows, the threshold for usefulness is lower than people assume, especially for large classes — so the "useless" framing often overstates the problem.

Second, the choice is not binary between "voluntary and ignored" and "mandatory and coerced." The evidence on raising response rates points to a third path: students respond when they believe it matters. The single most effective lever in the literature is closing the loop — visibly acting on previous feedback and telling students what changed ("you said, we did"). Appealing to students' sense of their own role in shaping their education raises participation without coercion (Nulty, 2008). Convenience, timely reminders, in-class time set aside, and short, engaging instruments all help.

Third, if you must nudge hard, prefer soft mandatory designs — for example, requiring a student to open the evaluation and explicitly choose to participate or decline — over withholding grades. You preserve the prompt without conditioning an academic outcome on compliance.

A better target than 100%

Reframe the metric. The goal is a representative, thoughtful sample, not a maximal one. That means:

  • Set size-appropriate response-rate targets (Nulty), not a blanket 100%.
  • Reduce the friction and the fatigue, rather than raising the penalty.
  • Make participation feel consequential by closing the loop every cycle.
  • Measure response quality, not just response rate.

What about incentives — are carrots any better than sticks?

If compulsion is the wrong lever, are incentives the right one? Partly. The literature distinguishes "stick" approaches (withholding grades) from "carrot" approaches (prize draws, early access to results, small course credit), and the carrots are ethically and methodologically gentler — they do not condition an academic outcome on compliance. But incentives carry their own quiet risk: pay students to respond and you may buy the same satisficing you were trying to avoid, plus a population that responds for the reward rather than because they have something to say.

The most defensible incentives are intrinsic and informational rather than transactional. Early release of the class's aggregate results, a visible summary of what the previous cohort's feedback changed, or simply framing the evaluation as a genuine consultation all reward participation with relevance rather than with a prize. That keeps the motivation aligned with the data quality you actually want: students who answer because they believe it matters give better answers than students who answer to win a voucher. The rough hierarchy is: close the loop first, reduce friction second, offer informational incentives third, and reach for compulsion only in the soft "open-and-choose" form, if at all.

A note on culture and comparability

One more reason to be wary of chasing maximal response rates by force: response styles are not uniform, and neither are the consequences of compulsion. In the diverse, international cohorts that are now the norm across European higher education, coercing reluctant respondents can systematically over-sample students who comply readily and under-represent those most alienated from the process — the very voices an evaluation most needs to hear. A high response rate can thus look representative while hiding a quieter selection problem. The honest metric pairs response rate with response quality and with some check on who is and is not answering. None of this means response rates do not matter — they do, and a chronically low one is a real warning sign. It means the number is a means, not the end. The end is a trustworthy, representative account of the student experience, and that is occasionally served, but never guaranteed, by making everyone click submit.

Where Koji fits

Koji for Education is built around earning engagement rather than compelling it. Instead of a static form students rush through, Koji runs a conversational, AI-moderated interview that feels like being listened to, which lifts both participation and the depth of what students say — without withholding anyone's grade. Quality scoring lets you distinguish thoughtful responses from satisficing, so you are measuring signal, not just completion. Formative, mid-cycle collection gives students feedback that visibly shapes the course they are still taking — the most powerful loop-closing nudge there is — and closing-the-loop action tracking lets you show the "you said, we did" that drives next term's response rate honestly. Because participation is genuinely voluntary and the AI moderator is consistent and bias-aware, the approach stays comfortably on the right side of European consent norms.

The same engine supports general user and customer research on the main Koji platform, for teams whose response-rate battles extend beyond the classroom.

You do not have to choose between coercion and silence. Build feedback students want to give, prove you act on it, and the response rate takes care of itself. Explore Koji for Education to see what voluntary-but-high-engagement evaluation looks like.