New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods9 min read

Should Course Evaluations Offer a "Don't Know" or "Not Applicable" Option?

The survey-methodology evidence on no-opinion and "not applicable" response options in course evaluations, anchored to Krosnick et al. (2002) and Krosnick (1991), with practical guidance on when an N/A option helps and when it invites satisficing.

Koji Education Team

Product

Quick answer: Offer a genuine "Not Applicable" option only when an item can truly fail to apply to a respondent (for example, "the lab sessions were useful" for a student who skipped labs). Do not add a generic "Don't Know" or "No Opinion" escape hatch to ordinary teaching-quality items. The landmark survey-methodology evidence (Krosnick et al., 2002, nine experiments) shows that no-opinion filters do not improve data quality; they mainly invite satisficing — they let unmotivated respondents skip the cognitive work of forming a judgement, removing meaningful opinions rather than removing noise.

The design question every evaluation form faces

Almost every course-evaluation instrument has to decide, item by item, what to do with a respondent who is unsure. Should the scale run cleanly from "strongly disagree" to "strongly agree," or should it include a "Don't Know" / "No Opinion" / "Not Applicable" exit? The intuition behind adding one is humane and statistical at once: surely it is better to let students who genuinely have no view say so, rather than forcing them to pick a number at random and pollute the data. That intuition is widely held — and the best evidence says it is mostly wrong.

What the research says

Krosnick et al. (2002), The Impact of "No Opinion" Response Options on Data Quality, is the definitive treatment. Across nine experiments embedded in three household surveys, the authors compared items that offered an explicit no-opinion option against matched items that omitted it and gently encouraged respondents to answer. The core finding: the quality of the attitude reports — judged by over-time consistency and by responsiveness to legitimate question manipulations — was not compromised by omitting the no-opinion option. In other words, the students and respondents who were "rescued" by a Don't Know option were not mostly people without real opinions; many had usable views they simply declined to exert effort to report. Offering the filter did not purify the data; it suppressed meaningful signal. The authors interpret no-opinion options as an "invitation to satisfice."

Krosnick (1991), Response strategies for coping with the cognitive demands of attitude measures in surveys, supplies the theory. Answering a survey item properly requires four cognitive steps — interpreting the question, retrieving relevant information, integrating it into a judgement, and mapping that judgement onto the response options. Satisficing occurs when a respondent, low on motivation or ability or facing a difficult item, shortcuts these steps. Selecting "Don't Know" is one of the canonical satisficing shortcuts, alongside straight-lining, choosing the first plausible option, and gravitating to the midpoint. The more an item offers a low-effort escape, the more satisficing it absorbs — which is precisely why a blanket no-opinion option is risky.

A third strand keeps the recommendation honest. Work in the same tradition (for example, debates summarised in survey-methods reviews and the broader Krosnick corpus) distinguishes sharply between a no-opinion filter on an attitude item and a substantively necessary "Not Applicable" branch. The latter is not satisficing — it is correct routing. A student who never attended the optional seminars cannot meaningfully rate them; forcing a number there manufactures noise. The methodological consensus is therefore conditional, not absolute: remove generic Don't Know escapes from judgement items, but preserve true N/A where non-applicability is real.

Why it matters for course evaluation in practice

  1. Generic Don't Know options quietly shrink your usable sample. Every student who satisfices into "No Opinion" on "the assessment was fair" is a lost data point that you actually could have collected. On already low-response evaluations, this compounds non-response bias with item-level non-response.

  2. N/A and Don't Know are not interchangeable, and conflating them corrupts denominators. If your platform treats a genuine "I didn't take the labs" the same as "I can't be bothered to judge the lectures," your percentages and means are computed over the wrong base. Reporting "% favourable" requires knowing whether a blank means inapplicable or unmotivated.

  3. Mid-points and Don't Knows interact. A scale that has both a neutral midpoint and a Don't Know gives satisficers two doors out. The cleaner design for most teaching-quality items is a fully labelled scale with no separate no-opinion column, paired with the freedom to skip — so that a true non-response is recorded as missing rather than masquerading as a neutral judgement.

Limitations and honest caveats

Three honest qualifications. First, Krosnick et al. (2002) studied political and social attitudes in household surveys, not course evaluations specifically; the cognitive mechanism (satisficing) generalises well, but the exact magnitudes in a 19-year-old rating a statistics module may differ. Second, the recommendation to drop no-opinion options assumes the items themselves are well-targeted to what students can actually observe — asking students to rate things they cannot judge (e.g., the instructor's command of the literature) and then removing their escape hatch simply forces guesswork; the fix there is a better item, not a Don't Know box. Third, cross-cultural response styles matter: in some contexts a missing midpoint or no-opinion option pushes acquiescent responders toward agreement, so the optimal design interacts with the population — a reason to test, not to assume. None of this overturns the core finding; it bounds its application.

How Koji incorporates this

Koji is built around conversational, AI-moderated interviews rather than static forms, which changes the no-opinion problem at its root.

  • Probing instead of an escape hatch. When a student hesitates or gives a thin answer, a Koji interview can ask a brief, neutral follow-up ("What made that hard to judge?") instead of offering a one-click "Don't Know." This is designed to recover the genuine signal that a no-opinion option would otherwise have suppressed — addressing exactly the loss Krosnick et al. (2002) identified — while still letting a student who truly cannot answer say so explicitly.

  • Distinguishing true N/A from disengagement. Because the interview is adaptive, Koji can route around items that do not apply (a student who reports not attending the labs is not asked to rate them) rather than presenting a blanket N/A that gets misused. The result is cleaner denominators: "not applicable" is captured as genuine non-applicability, separate from skipped or unmotivated responses.

  • Structured question types used deliberately. For items where a forced judgement is appropriate, Koji uses fully labelled scale and single_choice questions without a generic no-opinion column; for items where non-applicability is real, yes_no gating ("Did you attend the seminars?") branches the respondent appropriately. Open_ended, multiple_choice, and ranking questions capture nuance that a Don't Know option would have flattened.

  • Quality scoring and bias-aware reporting. Koji flags low-effort or satisficed responses through response-quality signals and reports missingness transparently, so a quality-assurance office can see whether a low item count reflects genuine non-applicability or disengagement — rather than silently averaging over an ambiguous base.

The same engine powers Koji's core research platform at koji.so, where the no-opinion-versus-satisficing trade-off is just as live in product and customer surveys as it is in the classroom.

Frequently asked questions

Should we ever include a "Don't Know" option on a teaching-quality item? Generally no. Krosnick et al. (2002) found that removing it does not harm data quality and that including it mainly invites satisficing. Prefer a fully labelled scale plus the ability to skip, which records a true non-response as missing.

What about "Not Applicable"? Keep it — but only where an item can genuinely fail to apply (optional labs, field trips, group work a student opted out of). True N/A is correct routing, not satisficing.

Isn't forcing an answer worse than letting students say "Don't Know"? The evidence says the people who take the Don't Know exit are not mostly opinionless; many hold usable views and simply avoid the effort. Better item design and gentle probing recover more valid data than a blanket escape hatch.

Does a neutral midpoint solve the same problem? No — a midpoint and a Don't Know are different. Offering both gives satisficers two low-effort exits. Decide each deliberately rather than stacking them.

How does a conversational format change this? It lets you probe a hesitant respondent for the real reason instead of offering a single dismissive option, and it can branch around items that do not apply — recovering signal a static form would lose.

References

  • Krosnick, J. A., Holbrook, A. L., Berent, M. K., Carson, R. T., Hanemann, W. M., Kopp, R. J., Mitchell, R. C., Presser, S., Ruud, P. A., Smith, V. K., Moody, W. R., Green, M. C., & Conaway, M. (2002). The impact of "no opinion" response options on data quality: Non-attitude reduction or an invitation to satisfice? Public Opinion Quarterly, 66(3), 371-403. https://doi.org/10.1086/341394
  • Krosnick, J. A. (1991). Response strategies for coping with the cognitive demands of attitude measures in surveys. Applied Cognitive Psychology, 5(3), 213-236. https://doi.org/10.1002/acp.2350050305

Related resources

Related articles

research-methods

Why Students Click Straight Down the Middle: Satisficing in Course Evaluations

A large share of students do not answer evaluation questions carefully — they straightline, speed through, and pick the easy option. Barge and Gehlbach (2012) showed this "satisficing" doesn't just add noise; it inflates your reliability and validity statistics, making bad data look good.

research-methods

Response Styles and Likert Scales: Why Cross-Cultural Evaluation Needs More Than Numbers

Acquiescence and extreme response styles vary systematically by culture (Harzing, 2006; Baumgartner & Steenkamp, 2001), which means raw Likert averages are not directly comparable across nationalities in Europe's multinational classrooms. This article explains the evidence and how to evaluate fairly across diverse cohorts.

research-methods

What Can Open-Text Student Comments Tell You That Likert Scores Cannot?

A research-grounded guide to open-ended student comments in course evaluation: what Alhija and Fresko found about who writes them and what they contain, how thematic analysis surfaces issues numbers miss (Stupans et al.), how rare abusive comments actually are (Tucker), and how to turn free text into reliable evidence.

research-methods

How Many Scale Points Should a Course-Evaluation Question Have?

What the measurement literature — Preston & Colman (2000), Weng (2004), Dawes (2008) — says about the optimal number of response categories on rating scales, and why the answer for course evaluation is not just a number but a question about what a Likert item can and cannot capture.