New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods9 min read

Do Interactive Follow-Up Probes Improve Open-Text Course Feedback Quality?

Survey-methodology experiments (Holland & Christian, 2009; Smyth et al., 2009) show that an interactive follow-up probe after a student''s first open-text answer increases response length and the number of distinct themes without raising item non-response. We review the evidence and its implications for conversational, AI-moderated course evaluation.

Koji Education Team

Product

In brief

Adding an interactive follow-up probe after a student''s first open-text answer reliably increases the length and thematic richness of the response, without pushing more students to skip the question. In a web-survey experiment, Holland and Christian (2009) found that respondents who received an automated probe — a short prompt asking them to say more after they submitted an initial answer — wrote longer, more interpretable answers containing more distinct ideas. Smyth, Dillman, Christian and McBride (2009) reached a converging conclusion: motivating instructions and answer-box design changes measurably improve open-ended response quality. The practical implication is that how you ask for open text matters as much as whether you ask, and that a single static comment box leaves a large amount of usable evidence on the table.

What the research says

The anchor study is Holland and Christian (2009), The Influence of Topic Interest and Interactive Probing on Responses to Open-Ended Questions in Web Surveys, in Social Science Computer Review (27(2), 196–212). The authors embedded an experiment in a web survey of university undergraduates. After a respondent submitted an answer to an open-ended question, some respondents were randomly shown an interactive follow-up probe — an automated request to elaborate or clarify — while others were not. They also measured respondents'' interest in the topic.

Two findings stand out. First, interest drives quality: respondents more interested in the topic wrote answers that were longer, more interpretable, and took more time to compose. Second, and more actionable for practitioners, the interactive probe improved response quality on top of interest: probed respondents produced more elaborated answers with more distinct themes. Crucially, the probe did not significantly increase item non-response — students did not abandon the question in protest at being asked for more. In other words, the probe extracted more signal without a measurable cost in coverage.

Smyth, Dillman, Christian and McBride (2009), Open-Ended Questions in Web Surveys: Can Increasing the Size of Answer Boxes and Providing Extra Verbal Instructions Improve Response Quality?, in Public Opinion Quarterly (73(2), 325–337), tested a complementary set of levers across three random-sample web surveys of undergraduates. They found that visual cues (a larger answer box) and, more powerfully, verbal instructions that clarify and motivate — telling respondents why the answer matters and what a useful answer looks like — increased the amount and specificity of information provided. The message is consistent with Holland and Christian: open-ended data quality is a design outcome, not a fixed property of the respondent.

The mechanism generalises beyond static web forms. In the interviewing literature, Conrad and Schober (2000) showed that allowing interviewers to clarify question meaning — conversational interviewing — substantially improved response accuracy for questions where respondents'' situations did not map cleanly onto standardised wording, at the cost of longer interviews. Behr, Kaczmirek, Bandilla and Braun (2012) extended probing into cross-national web surveys ("web probing"), demonstrating that automated probes can elicit interpretable, analysable qualitative material at scale and can reveal how respondents actually understand a question. Together these studies describe a coherent principle: interaction — a well-timed prompt to say more, or to clarify — converts thin, one-shot open answers into richer, more codeable evidence.

Why it matters for course evaluation in practice

Open-text comments are the part of a course evaluation that programme directors actually read and that drive most concrete change, yet they are notoriously thin. A large share of students leave the comment box blank or write a few words ("good", "too much work", "the lecturer was nice"). Those fragments are hard to code, dominated by a vocal minority, and skewed by negativity bias in how they are read. The Holland and Christian result says this thinness is partly an artefact of the instrument: a single static box with a generic prompt suppresses elaboration that a follow-up would have surfaced.

For a quality-assurance process, richer open text has three payoffs. It improves thematic analysis: more distinct themes per respondent means faster thematic saturation and more reliable coding. It improves actionability: "the assessment was unfair" is not actionable, but a probed elaboration — "the rubric wasn''t shared until after the first submission" — is. And it improves triangulation: elaborated text lets you interpret a low scale score rather than guessing at it. In accreditation terms (ESG Standard 1.9 on ongoing monitoring), evidence that you elicit specific, improvement-relevant student feedback — not just numeric averages — is stronger than a wall of one-word comments.

A concrete illustration makes the stakes clear. Imagine two students who both dislike a module''s assessment. On a static form, both write "the assessment was unfair" and stop. Under an interactive design, the first is asked "can you say a bit more about what made it feel unfair?" and elaborates that the marking rubric arrived only after the first submission deadline; the second, similarly probed, explains that the weighting between coursework and exam was not what the syllabus implied. The raw scores are identical and uninformative; the probed responses point to two different, separately fixable problems. Multiplied across a cohort, this is the difference between a word cloud of "unfair" and a ranked list of specific, actionable causes — the difference between data that can be filed and data that can actually drive change. This is why the Holland and Christian result is not a marginal methodological nicety but a lever on the single most-used part of a course evaluation.

Limitations and honest caveats

Generalisability across populations and topics. Holland and Christian and Smyth et al. studied US undergraduates in general-purpose web surveys, not European course evaluations specifically. The direction of the effect is robust and mechanistically plausible, but the magnitude in a given institution, language, and evaluation culture is an empirical question. Effects may be smaller where students are already highly engaged, or where survey fatigue is severe.

Probing has a cost, and it can be overdone. Every probe adds time. Conrad and Schober''s conversational interviewing improved accuracy but lengthened interviews; the analogous risk online is that aggressive or repeated probing increases break-off or annoyance, especially on mobile devices or at the end of a long instrument. The 2009 studies found non-response was not significantly harmed by a single well-designed probe — that finding should not be over-extrapolated to interrogative, multi-probe designs.

Quantity is not automatically quality. More words and more themes are proxies for informativeness, not guarantees of it. A probe can also elicit repetition, venting, or socially desirable elaboration. The construct being improved is "amount and interpretability of information," which correlates with usefulness but is not identical to it.

Probe content matters, and a bad probe can bias. A leading probe ("What did you dislike about the assessment?") imports the very framing effects that good survey design tries to avoid, and can manufacture negativity. The evidence supports neutral, elaboration-seeking probes ("Can you say a bit more about what led you to that?"), not directive ones. This is the central risk to manage when probing is automated by an AI.

How Koji incorporates this

This literature is, in effect, the research warrant for Koji''s core method — and we are careful to frame it as designed to elicit better evidence, not to manufacture consensus.

AI-moderated conversational probing. Where a traditional form offers one static comment box, Koji conducts an AI-moderated conversational interview: after a student''s first open-text answer, the moderator can ask a neutral follow-up to elaborate, clarify, or give an example — precisely the interactive probe that Holland and Christian showed increases length and distinct themes. The probe is adaptive: a rich first answer needs no probe; a thin one ("too much work") invites a single "what specifically felt like too much?" This targets effort where it adds signal, mirroring the finding that interest and elaboration travel together.

Neutral, non-leading probe design. Because the caveat above is real, Koji''s moderation is designed to ask open, non-directive follow-ups rather than leading questions, to avoid importing framing effects and manufacturing negativity. Institutions can review and constrain probe behaviour, and probes are logged for audit.

Guardrails against over-probing. Consistent with the non-response caveat, the interview is bounded — probes are limited and skippable — so the design captures the elaboration benefit documented in 2009 without drifting into the fatigue and break-off risks that heavier probing invites.

Downstream analysis. Richer probed text feeds Koji''s automatic thematic analysis and quality scoring, so the additional themes an interactive probe surfaces are actually coded and reported, not lost. This closes the loop from elicitation to insight.

Koji''s core research platform at koji.so applies the same AI-moderated probing engine to product, UX, and customer research, where thin open-ended answers are an identical and well-known problem.

Related resources

References

  • Holland, J. L., & Christian, L. M. (2009). The influence of topic interest and interactive probing on responses to open-ended questions in web surveys. Social Science Computer Review, 27(2), 196–212. https://doi.org/10.1177/0894439108327894
  • Smyth, J. D., Dillman, D. A., Christian, L. M., & McBride, M. (2009). Open-ended questions in web surveys: Can increasing the size of answer boxes and providing extra verbal instructions improve response quality? Public Opinion Quarterly, 73(2), 325–337. https://doi.org/10.1093/poq/nfp029
  • Behr, D., Kaczmirek, L., Bandilla, W., & Braun, M. (2012). Asking probing questions in web surveys: Which factors have an impact on the quality of responses? Social Science Computer Review, 30(4), 487–498. https://doi.org/10.1177/0894439311435305
  • Conrad, F. G., & Schober, M. F. (2000). Clarifying question meaning in a household telephone survey. Public Opinion Quarterly, 64(1), 1–28. https://doi.org/10.1086/316757