New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
best-practices9 min read

Should You Force Students to Complete Course Evaluations? Coercion, Ethics, and Data Quality

Withholding grades to force evaluation completion pushes response toward 100% but injects careless responding — the evidence, the ethics, and better routes to a representative sample.

Koji Education Team

Product

In brief

Forcing students to complete evaluations — by withholding grades, blocking registration, or making the survey a required gate — reliably pushes response rates toward 100%. But the best direct evidence shows coercion buys the rate at the cost of data quality: when Bahous and colleagues (2018) made clerkship evaluations compulsory, response jumped from 56% to 100%, yet about a third of compelled students endorsed a deliberately nonsensical "sham" item, revealing widespread inattentive answering. A high response rate obtained under coercion is not the same as a representative, honest one, and the ethics of compelling participation deserve as much scrutiny as the statistics.

What the research says

The instinct behind mandatory completion is sound in one narrow sense: low response rates genuinely threaten representativeness, and missing-not-at-random nonresponse can bias results if the students who skip differ systematically from those who answer. So administrators reach for the strongest lever available — making completion a condition of seeing grades or of re-enrolling — and the response rate duly climbs. The question the evidence answers is what else changes when it does.

The cleanest study is Bahous, Salameh, Salloum, Salameh, Park and Tekian (2018), who compared voluntary and compulsory administration of the same clerkship evaluations at a medical school. Making completion compulsory raised the response rate from 56% to 100% (p < 0.001) — exactly the win administrators hope for. But the authors had embedded a "sham item": a deliberately irrelevant, nonsensical question that an attentive respondent should never endorse. Under compulsion, roughly 33% of respondents gave a positive rating to this meaningless item, direct evidence that a large minority were clicking through without reading. Ratings were mostly similar between conditions, though one clerkship scored significantly lower under compulsion; reliability was adequate and comparable across the two. The authors' conclusion is measured and important: using authority to force participation raised the quantity of responses but not their quality or validity, and it introduced a fresh threat — inattentive, careless responding — that voluntary administration did not have to the same degree.

This sits inside a wider literature on what response rates actually mean. Nulty (2008) argues that the relevant question is not "is the rate high?" but "is the rate high enough to make the sample adequately representative for the decision at hand" — and that adequacy depends on class size and stakes, not on a universal threshold. Dommeyer, Baum, Hanna and Chapman (2004) tested a range of incentives for online evaluations and found that raising rates without coercion is possible but hard, and that different collection methods trade off rate against conditions of completion. Adams and Umbach (2012), analysing roughly 135,000 online evaluations, show that nonresponse is patterned — driven by salience, fatigue and the academic environment — which is why simply inflating the count does not automatically fix representativeness if the marginal responders are answering carelessly. The through-line is consistent: a response rate is a means to representativeness and honesty, and a rate manufactured by compulsion can deliver the number while undermining the thing the number was supposed to guarantee.

Why it matters for course evaluation in practice

Three practical consequences follow for any institution weighing a mandatory-completion policy.

Coercion can degrade the very data it is meant to protect. The point of chasing response rate is to reduce nonresponse bias. But if compulsion injects a third of respondents who straightline, endorse sham items, and click through — the behaviour Bahous et al. measured — then the "complete" dataset now carries a measurement problem in place of the coverage problem. You have traded a known, estimable bias (who is missing) for a hidden one (who is present but not really answering), and the second is harder to detect and correct. A dataset that is 100% complete and 33% inattentive is not obviously better than one that is 56% complete and largely sincere.

It changes the meaning of consent and can breed resentment. Withholding grades or blocking registration until a student evaluates is coercive in a straightforward sense: the student cannot decline without penalty. Beyond the ethics, coerced participants who resent the gate are more likely to satisfice or retaliate, and the feedback that results is contaminated by the manner of its collection — an interaction between how honest students are and whether they trust the process. Mandatory policies also intensify survey fatigue when every course gates on completion, so students race through a stack of required surveys at term's end.

The high number can mislead decision-makers. A committee that sees "100% response" naturally trusts the result more than "56%." If that 100% is partly careless, the apparent precision is false, and personnel or programme decisions rest on a firmer-looking but weaker foundation. Honest reporting of a coerced dataset should disclose the completion conditions and screen aggressively for careless responding, not present the rate as an unqualified strength.

Limitations and honest caveats

The case against coercion should be stated fairly, without overreach.

The strongest single study is in one context. Bahous et al. (2018) is a medical-school clerkship setting with its own culture and stakes; the sham-item result is striking but should be read as strong indicative evidence, not a universal constant. The exact proportion of careless responding under compulsion will vary by discipline, survey length, and how the mandate is framed. The finding that matters — that compulsion can manufacture rate without quality — is robust in principle even where the 33% figure is not the local number.

Not all "mandatory" policies are equally coercive. There is a spectrum from a hard grade-withholding gate, through registration holds, to a soft default where the survey opens automatically but can be dismissed, to persistent reminders with no penalty. The evidence against coercion is strongest for the hard-gate end; softer nudges that raise salience without penalty are closer to the legitimate rate-raising Dommeyer and Nulty discuss and may not carry the same quality cost.

Low response rates are a genuine problem, not a non-issue. Rejecting coercion is not the same as tolerating a 20% response rate. Nonresponse bias is real, and a programme still needs a defensible rate. The argument is that the route to that rate matters: representativeness and honesty are the goals, and compulsion optimises a proxy (raw count) at their expense. The constructive response is to raise rates through salience, trust and good timing, and to measure and disclose representativeness rather than assume a high number settles it.

Institutional and legal context varies. In some systems, tying academic records to survey completion may raise data-protection or fairness concerns of its own. Policies should be checked against local regulation and student-charter commitments, not adopted purely on response-rate grounds.

How Koji incorporates this

Koji's design philosophy treats a voluntary, high-quality response as the goal and a coerced, careless one as a failure mode to detect. Rather than relying on grade-withholding to drive completion, Koji is built to raise genuine engagement through the format itself: an AI-moderated conversational evaluation is shorter-feeling and more responsive than a long grid, which addresses the salience-and-fatigue drivers of nonresponse that Adams and Umbach identify — the lever Nulty and Dommeyer point to, pulled through experience rather than compulsion.

Where an institution's policy does mandate completion, Koji is designed to protect the resulting data. Its data-quality tooling screens for the exact behaviour Bahous et al. measured — straightlining, implausibly fast completion, and inattentive patterns — so a compelled-but-careless response can be flagged rather than silently counted, and reports can distinguish attentive from likely-careless submissions. Because the moderator asks follow-up questions and expects substantive answers, pure click-through is harder to sustain than in a static form, and a student who tries to race through is gently prompted to elaborate. Koji's reporting is designed to present response representativeness and quality alongside the raw rate, so a decision-maker sees a "100% but screen-flagged" dataset for what it is rather than trusting the headline number. None of this is framed as eliminating the problem — a determined student can always satisfice — but as designed to mitigate the quality cost of coerced participation and to keep the focus on honest, usable feedback. The same conversational engine and careless-response screening run on Koji's core research platform at koji.so, where incentivised customer-research panels present the identical risk of paid-but-inattentive answers.

Frequently asked questions

Does making evaluations mandatory improve the data?

It reliably improves the response rate but not necessarily the quality. In the clearest study, Bahous et al. (2018) found compulsion lifted response from 56% to 100%, yet about a third of compelled respondents endorsed a deliberately nonsensical sham item — evidence of widespread inattentive answering. You trade a coverage problem (who is missing) for a measurement problem (who is present but not really responding), and the second is harder to detect and correct.

Is withholding grades to force evaluation completion ethical?

It is coercive: the student cannot decline without a real penalty, which strains the idea of voluntary participation. Beyond ethics, coerced and resentful respondents are more prone to satisficing or retaliation, contaminating the feedback with the manner of its collection. Many institutions also face data-protection or fairness constraints on tying academic records to survey completion, so hard grade-withholding gates should be adopted cautiously, if at all, and checked against local regulation.

Isn't a low response rate a serious problem, though?

Yes — nonresponse bias is real, and a 20% rate can genuinely misrepresent a cohort. The argument is not that response rate is unimportant but that how you raise it matters. Compulsion optimises the raw count while potentially degrading honesty and attention, which are the actual goals. Better routes are raising salience and trust, improving timing, and shortening the experience — then measuring representativeness rather than assuming a high number settles it.

What is a "sham item" and why does it matter?

A sham item is a deliberately meaningless question that an attentive respondent should never endorse, embedded to detect careless answering. In Bahous et al. (2018), roughly 33% of compelled students gave it a positive rating, a direct measure of how many were clicking through without reading. Sham or "bogus" items are a practical way to audit whether a high-completion dataset is actually being answered attentively.

How high does my response rate actually need to be?

There is no universal threshold. Nulty (2008) argues adequacy depends on class size and the stakes of the decision: a small class needs a higher proportion than a large one to represent it reliably, and high-stakes personnel uses demand more than formative feedback. The right question is "is this sample representative enough for this decision?" — answered by checking who responded against the cohort, not by hitting a fixed number.

What should we do instead of coercion?

Raise rates through legitimate means and then verify quality. Improve salience and timing, communicate that feedback is genuinely used, keep the instrument short, and use conversational rather than grid formats to reduce fatigue. Then screen the responses you get for careless answering, and report representativeness alongside the rate. This produces feedback that is both reasonably complete and actually honest, rather than complete-but-hollow.

References

  • Bahous, S. A., Salameh, P., Salloum, A., Salameh, W., Park, Y. S., & Tekian, A. (2018). Voluntary or compulsory student evaluation of clerkships? Effect on validity and potential bias. BMC Medical Education, 18, 9. https://doi.org/10.1186/s12909-017-1116-8
  • Nulty, D. D. (2008). The adequacy of response rates to online and paper surveys: What can be done? Assessment & Evaluation in Higher Education, 33(3), 301–314. https://doi.org/10.1080/02602930701293231
  • Dommeyer, C. J., Baum, P., Hanna, R. W., & Chapman, K. S. (2004). Gathering faculty teaching evaluations by in-class and online surveys: Their effects on response rates and evaluations. Assessment & Evaluation in Higher Education, 29(5), 611–623. https://www.csuchico.edu/ir/_assets/documents/set/dommeyer-2004.pdf
  • Adams, M. J. D., & Umbach, P. D. (2012). Nonresponse and online student evaluations of teaching: Understanding the influence of salience, fatigue, and academic environments. Research in Higher Education, 53(5), 576–591. https://doi.org/10.1007/s11162-011-9240-5

Related resources