New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
evaluation-design9 min read

How to Design Open-Ended Course Evaluation Questions That Get Useful Answers

Open-ended evaluation questions usually fail not because students have nothing to say, but because the prompt asks for too little. The survey-methodology evidence on answer-box size, verbal instructions, and specific framing — and what it means for the comment box.

Koji Education Team

Product

In brief: The quality of open-text course feedback is largely a property of the prompt, not the student. Experimental survey research (Smyth, Dillman, Christian & McBride, 2009) shows that small design choices — a larger answer box, a brief instruction signalling that a thoughtful answer matters, and a specific rather than generic question — measurably increase the length, number of distinct themes, and usefulness of open-ended responses. A blank box labelled "Any other comments?" is the weakest possible design; targeted prompts with the right cues reliably do better.

Most course evaluations end with a generic comment box, and most of those boxes come back empty or filled with a one-line "good course." Institutions often conclude that students will not write detailed feedback. The survey-methodology evidence says the opposite: students will write, but only if the question is engineered to invite it. The comment box is not a neutral container into which thoughts flow; it is an instrument whose design determines how much and how usefully people respond.

What the research says

The key experiment is Smyth, Dillman, Christian and McBride (2009), "Open-ended questions in web surveys: Can increasing the size of answer boxes and providing extra verbal instructions improve response quality?", published in Public Opinion Quarterly. The authors manipulated two features of open-ended web-survey questions: the size of the answer box and the presence of extra verbal instructions in the question stem (cues such as noting that the question is important and asking respondents to take their time). The manipulations worked. Larger answer boxes elicited longer responses and more distinct themes; verbal instructions emphasising importance and effort increased both the time respondents spent and the substance of what they wrote. In other words, two cheap, content-neutral design changes raised response quality without changing the underlying question.

This sits inside a robust strand of survey methodology associated with Dillman's tailored-design tradition, which holds that visual and verbal cues are read by respondents as signals of what kind of answer is expected. A tiny single-line box says "a few words will do"; a generous box says "I expect a paragraph." A bare prompt says "this is optional"; a short, motivating instruction says "your answer matters and will be read." Related work in the same tradition shows that the specificity of the question stem drives the specificity of the answer: a generic "Do you have any comments?" invites generic pleasantries, while a focused prompt ("What is one thing that would have helped you learn more in this course?") channels the respondent toward concrete, actionable content. Specific framing guides without leading — it signals the topic and the desired granularity, not the desired verdict.

A second well-established finding from this literature concerns item nonresponse: open-ended questions are abandoned far more often than closed ones, and the abandonment rate is sensitive to placement, burden, and motivation. Asking many open questions, or placing them after the respondent is already fatigued, depresses both completion and quality — a concern that connects directly to the evidence on survey fatigue and questionnaire length. The methodological consensus is that a few well-designed, well-placed, specific open questions outperform a long tail of generic ones.

Why it matters for course evaluation in practice

Open text is the part of an evaluation that quality teams and instructors say they value most — it is where the why behind the numbers lives, and where actionable suggestions appear. But its value collapses if the prompt is badly designed. The research translates into concrete practice.

First, replace the generic comment box with specific prompts. "Any other comments?" is the single least productive question on most forms. Splitting it into a small number of targeted prompts — one on what helped learning, one on what hindered it, one on a concrete change for next time — produces richer, codable answers. This complements our guidance on writing better course evaluation questions and on the distinctive value of student-written comments.

Second, signal that the answer matters. A one-line instruction ("Your specific examples directly shape how this course is taught next year — please take a moment") is not filler; the experimental evidence shows such cues lengthen and deepen responses. Students ration effort according to perceived payoff, and most forms give them no reason to believe anyone will read the box.

Third, give the answer room. On the web and on mobile, the visible size of the response field is a cue. A cramped box suppresses detail; a clearly expandable field invites it. Because much evaluation is now completed on phones (see device effects), the field must look invitingly answerable on a small screen, not like a tweet limit.

Fourth, budget open questions deliberately. Given nonresponse and fatigue, resist the urge to add an open box to every section. A few specific, well-placed prompts yield more usable text than a dozen generic ones — and they make downstream thematic analysis tractable, because specific prompts return on-topic answers that code cleanly.

Crucially, prompts should be pretested. What reads as specific to a committee may be ambiguous to a first-year student. Cognitive interviewing and pretesting catch prompts that students interpret differently from designers — a frequent, invisible cause of off-topic open-text.

Limitations and honest caveats

The evidence base, while solid, has boundaries. Smyth et al. and most of the supporting work were conducted on general web surveys, not course evaluations specifically; students completing a familiar end-of-term form may respond to box size and instructions somewhat differently from a one-off survey panel. The direction of the effects is well replicated, but the magnitude in a particular course-evaluation context should be confirmed locally, not assumed.

There is also a quantity-versus-quality subtlety. Larger boxes and motivating instructions reliably increase response length and theme count; length is a reasonable proxy for substance, but it is not identical to it. A longer answer can be longer venting rather than more insight. The design changes raise the ceiling on useful feedback; they do not guarantee every extra word is useful, and they can, at the margin, increase the volume of negative or abusive content that then needs handling.

Specific prompts carry their own risk: over-specificity can narrow the lens. A prompt that asks only about "lectures" may suppress feedback about assessment or supervision. The remedy is balance — a small set of prompts covering the main dimensions plus one genuinely open catch-all — rather than either a single generic box or an over-engineered checklist of micro-questions. Finally, these are design effects, not motivation cures. No prompt design rescues an evaluation that students believe is ignored; the closing-the-loop evidence shows perceived impact is itself a powerful driver of whether students bother to write at all.

How Koji incorporates this

Koji's central design choice — an AI-moderated conversational interview rather than a static form — is, in effect, the prompt-design research taken to its logical conclusion: instead of optimising a single fixed box, it asks specific, adaptive questions and follows up.

  • Specific, adaptive prompting instead of a generic box. Rather than "Any other comments?", Koji's conversational engine asks targeted questions and then probes — "you mentioned the labs were rushed; what specifically would have helped?" — the dynamic equivalent of the specific-framing and follow-up cues the research shows lift response quality. This turns a one-line "good course" into a substantive, on-topic exchange.
  • Built-in "your answer matters" signalling. A conversation inherently communicates that responses are being read and used, supplying the motivational cue Smyth et al. created with verbal instructions — without the student having to trust a static promise.
  • Structured open-ended items where a form is preferred. For teams using classic forms, Koji supports a true open_ended question type that can be paired with specific stems and appropriate field affordances, plus scale, ranking, single_choice, multiple_choice, and yes_no for the closed items, so designers can deploy a few well-targeted open prompts rather than one catch-all.
  • Fatigue-aware design. Because the conversational format adapts and does not bolt an open box onto every section, it manages the nonresponse-and-fatigue trade-off the literature warns about, asking deeper questions only where the respondent has something to say.
  • Downstream thematic analysis and quality scoring. Because specific prompts return on-topic text, Koji's automatic thematic analysis and per-response quality scoring work on cleaner input, and low-substance or abusive responses can be flagged for appropriate handling.

Koji does not pretend that better prompting guarantees insight — length is not the same as substance, and no design overcomes the belief that feedback is ignored. What it does is apply the well-evidenced levers (specificity, follow-up, perceived importance, appropriate burden) systematically, where a static form applies them once or not at all. The same conversational interview engine powers Koji's core research platform at koji.so, where eliciting specific, codable open-ended responses is the central craft of product and customer research.

Related Resources

References

  1. Smyth, J. D., Dillman, D. A., Christian, L. M., & McBride, M. (2009). Open-ended questions in web surveys: Can increasing the size of answer boxes and providing extra verbal instructions improve response quality? Public Opinion Quarterly, 73(2), 325–337. https://academic.oup.com/poq/article-abstract/73/2/325/1938934
  2. Dillman, D. A., Smyth, J. D., & Christian, L. M. (2014). Internet, Phone, Mail, and Mixed-Mode Surveys: The Tailored Design Method (4th ed.). Wiley. https://www.wiley.com/en-us/Internet%2C+Phone%2C+Mail%2C+and+Mixed+Mode+Surveys%3A+The+Tailored+Design+Method%2C+4th+Edition-p-9781118456149
  3. Christian, L. M., Dillman, D. A., & Smyth, J. D. (2007). Helping respondents get it right the first time: The influence of words, symbols, and graphics in web surveys. Public Opinion Quarterly, 71(1), 113–125. https://doi.org/10.1093/poq/nfl039