How to Design Open-Ended Course Evaluation Questions That Get Useful Answers
Open-ended evaluation questions usually fail not because students have nothing to say, but because the prompt asks for too little. The survey-methodology evidence on answer-box size, verbal instructions, and specific framing — and what it means for the comment box.
Koji Education Team
Product
In brief: The quality of open-text course feedback is largely a property of the prompt, not the student. Experimental survey research (Smyth, Dillman, Christian & McBride, 2009) shows that small design choices — a larger answer box, a brief instruction signalling that a thoughtful answer matters, and a specific rather than generic question — measurably increase the length, number of distinct themes, and usefulness of open-ended responses. A blank box labelled "Any other comments?" is the weakest possible design; targeted prompts with the right cues reliably do better.
Most course evaluations end with a generic comment box, and most of those boxes come back empty or filled with a one-line "good course." Institutions often conclude that students will not write detailed feedback. The survey-methodology evidence says the opposite: students will write, but only if the question is engineered to invite it. The comment box is not a neutral container into which thoughts flow; it is an instrument whose design determines how much and how usefully people respond.
What the research says
The key experiment is Smyth, Dillman, Christian and McBride (2009), "Open-ended questions in web surveys: Can increasing the size of answer boxes and providing extra verbal instructions improve response quality?", published in Public Opinion Quarterly. The authors manipulated two features of open-ended web-survey questions: the size of the answer box and the presence of extra verbal instructions in the question stem (cues such as noting that the question is important and asking respondents to take their time). The manipulations worked. Larger answer boxes elicited longer responses and more distinct themes; verbal instructions emphasising importance and effort increased both the time respondents spent and the substance of what they wrote. In other words, two cheap, content-neutral design changes raised response quality without changing the underlying question.
This sits inside a robust strand of survey methodology associated with Dillman's tailored-design tradition, which holds that visual and verbal cues are read by respondents as signals of what kind of answer is expected. A tiny single-line box says "a few words will do"; a generous box says "I expect a paragraph." A bare prompt says "this is optional"; a short, motivating instruction says "your answer matters and will be read." Related work in the same tradition shows that the specificity of the question stem drives the specificity of the answer: a generic "Do you have any comments?" invites generic pleasantries, while a focused prompt ("What is one thing that would have helped you learn more in this course?") channels the respondent toward concrete, actionable content. Specific framing guides without leading — it signals the topic and the desired granularity, not the desired verdict.
A second well-established finding from this literature concerns item nonresponse: open-ended questions are abandoned far more often than closed ones, and the abandonment rate is sensitive to placement, burden, and motivation. Asking many open questions, or placing them after the respondent is already fatigued, depresses both completion and quality — a concern that connects directly to the evidence on survey fatigue and questionnaire length. The methodological consensus is that a few well-designed, well-placed, specific open questions outperform a long tail of generic ones.
Why it matters for course evaluation in practice
Open text is the part of an evaluation that quality teams and instructors say they value most — it is where the why behind the numbers lives, and where actionable suggestions appear. But its value collapses if the prompt is badly designed. The research translates into concrete practice.
First, replace the generic comment box with specific prompts. "Any other comments?" is the single least productive question on most forms. Splitting it into a small number of targeted prompts — one on what helped learning, one on what hindered it, one on a concrete change for next time — produces richer, codable answers. This complements our guidance on writing better course evaluation questions and on the distinctive value of student-written comments.
Second, signal that the answer matters. A one-line instruction ("Your specific examples directly shape how this course is taught next year — please take a moment") is not filler; the experimental evidence shows such cues lengthen and deepen responses. Students ration effort according to perceived payoff, and most forms give them no reason to believe anyone will read the box.
Third, give the answer room. On the web and on mobile, the visible size of the response field is a cue. A cramped box suppresses detail; a clearly expandable field invites it. Because much evaluation is now completed on phones (see device effects), the field must look invitingly answerable on a small screen, not like a tweet limit.
Fourth, budget open questions deliberately. Given nonresponse and fatigue, resist the urge to add an open box to every section. A few specific, well-placed prompts yield more usable text than a dozen generic ones — and they make downstream thematic analysis tractable, because specific prompts return on-topic answers that code cleanly.
Crucially, prompts should be pretested. What reads as specific to a committee may be ambiguous to a first-year student. Cognitive interviewing and pretesting catch prompts that students interpret differently from designers — a frequent, invisible cause of off-topic open-text.
Limitations and honest caveats
The evidence base, while solid, has boundaries. Smyth et al. and most of the supporting work were conducted on general web surveys, not course evaluations specifically; students completing a familiar end-of-term form may respond to box size and instructions somewhat differently from a one-off survey panel. The direction of the effects is well replicated, but the magnitude in a particular course-evaluation context should be confirmed locally, not assumed.
There is also a quantity-versus-quality subtlety. Larger boxes and motivating instructions reliably increase response length and theme count; length is a reasonable proxy for substance, but it is not identical to it. A longer answer can be longer venting rather than more insight. The design changes raise the ceiling on useful feedback; they do not guarantee every extra word is useful, and they can, at the margin, increase the volume of negative or abusive content that then needs handling.
Specific prompts carry their own risk: over-specificity can narrow the lens. A prompt that asks only about "lectures" may suppress feedback about assessment or supervision. The remedy is balance — a small set of prompts covering the main dimensions plus one genuinely open catch-all — rather than either a single generic box or an over-engineered checklist of micro-questions. Finally, these are design effects, not motivation cures. No prompt design rescues an evaluation that students believe is ignored; the closing-the-loop evidence shows perceived impact is itself a powerful driver of whether students bother to write at all.
How Koji incorporates this
Koji's central design choice — an AI-moderated conversational interview rather than a static form — is, in effect, the prompt-design research taken to its logical conclusion: instead of optimising a single fixed box, it asks specific, adaptive questions and follows up.
- Specific, adaptive prompting instead of a generic box. Rather than "Any other comments?", Koji's conversational engine asks targeted questions and then probes — "you mentioned the labs were rushed; what specifically would have helped?" — the dynamic equivalent of the specific-framing and follow-up cues the research shows lift response quality. This turns a one-line "good course" into a substantive, on-topic exchange.
- Built-in "your answer matters" signalling. A conversation inherently communicates that responses are being read and used, supplying the motivational cue Smyth et al. created with verbal instructions — without the student having to trust a static promise.
- Structured open-ended items where a form is preferred. For teams using classic forms, Koji supports a true
open_endedquestion type that can be paired with specific stems and appropriate field affordances, plusscale,ranking,single_choice,multiple_choice, andyes_nofor the closed items, so designers can deploy a few well-targeted open prompts rather than one catch-all. - Fatigue-aware design. Because the conversational format adapts and does not bolt an open box onto every section, it manages the nonresponse-and-fatigue trade-off the literature warns about, asking deeper questions only where the respondent has something to say.
- Downstream thematic analysis and quality scoring. Because specific prompts return on-topic text, Koji's automatic thematic analysis and per-response quality scoring work on cleaner input, and low-substance or abusive responses can be flagged for appropriate handling.
Koji does not pretend that better prompting guarantees insight — length is not the same as substance, and no design overcomes the belief that feedback is ignored. What it does is apply the well-evidenced levers (specificity, follow-up, perceived importance, appropriate burden) systematically, where a static form applies them once or not at all. The same conversational interview engine powers Koji's core research platform at koji.so, where eliciting specific, codable open-ended responses is the central craft of product and customer research.
Related Resources
- How to Write Better Course Evaluation Questions
- The Value of Student-Written Comments in Course Evaluation
- Text Analytics for Open-Ended Student Comments
- Cognitive Interviewing and Pretesting Course Evaluation Questions
- Survey Fatigue and Course Evaluation Response Rates
- Can AI Analyze Open-Text Student Feedback? LLM Thematic Analysis
References
- Smyth, J. D., Dillman, D. A., Christian, L. M., & McBride, M. (2009). Open-ended questions in web surveys: Can increasing the size of answer boxes and providing extra verbal instructions improve response quality? Public Opinion Quarterly, 73(2), 325–337. https://academic.oup.com/poq/article-abstract/73/2/325/1938934
- Dillman, D. A., Smyth, J. D., & Christian, L. M. (2014). Internet, Phone, Mail, and Mixed-Mode Surveys: The Tailored Design Method (4th ed.). Wiley. https://www.wiley.com/en-us/Internet%2C+Phone%2C+Mail%2C+and+Mixed+Mode+Surveys%3A+The+Tailored+Design+Method%2C+4th+Edition-p-9781118456149
- Christian, L. M., Dillman, D. A., & Smyth, J. D. (2007). Helping respondents get it right the first time: The influence of words, symbols, and graphics in web surveys. Public Opinion Quarterly, 71(1), 113–125. https://doi.org/10.1093/poq/nfl039
Related articles
Text Analytics for Open-Ended Student Comments: What NLP Can and Cannot Tell You
Natural-language processing can turn thousands of free-text course-evaluation comments into themes and sentiment at scale — but the research (Cunningham-Nelson 2019; Sunar & Khalid 2023) shows where automated analysis is reliable and where human judgement is still required.
Survey Fatigue: Why Over-Surveying Students Quietly Wrecks Your Response Rates
Porter, Whitcomb & Weitzer (2004) showed that administering multiple surveys in one year suppresses later response rates. A research-grounded guide to survey fatigue in course evaluation — what causes it, what the evidence shows, and how to design around it.
What Can Open-Text Student Comments Tell You That Likert Scores Cannot?
A research-grounded guide to open-ended student comments in course evaluation: what Alhija and Fresko found about who writes them and what they contain, how thematic analysis surfaces issues numbers miss (Stupans et al.), how rare abusive comments actually are (Tucker), and how to turn free text into reliable evidence.
Before You Field It, Test It: Cognitive Interviewing for Course-Evaluation Questions
Why the wording of a course-evaluation item should be cognitively pretested before it reaches students, what Beatty and Willis (2007) established about think-aloud and verbal probing, and how Koji operationalises probing at scale.