New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods9 min read

Before You Field It, Test It: Cognitive Interviewing for Course-Evaluation Questions

Why the wording of a course-evaluation item should be cognitively pretested before it reaches students, what Beatty and Willis (2007) established about think-aloud and verbal probing, and how Koji operationalises probing at scale.

Koji Education Team

Product

In brief: A course-evaluation question that looks clear to the committee that wrote it can still be misread, differently interpreted, or answered from the wrong information by students. Cognitive interviewing — administering draft items to a small sample while collecting verbal evidence of how respondents comprehend, retrieve, judge, and respond — is the established method for catching those defects before fielding. Beatty and Willis (2007) synthesised the practice and showed it reliably surfaces comprehension and response problems that expert review alone misses. For course evaluation, this means pretesting your instrument is not optional polish; it is the difference between measuring what you intend and measuring noise.

The question you never tested is the question you cannot trust

Most course-evaluation instruments are written by a quality-assurance committee, approved by a teaching-and-learning board, and then fielded to thousands of students — without anyone ever checking whether a student reads "the course was intellectually stimulating" the way the committee meant it. This is a measurement risk hiding in plain sight. If a sizeable fraction of respondents interpret an item differently from its authors, the resulting mean is an average over incompatible interpretations, and no amount of downstream statistical sophistication can repair it.

Cognitive interviewing is the survey-methodology answer to this problem. It is a pretesting technique in which draft questions are administered to a small, purposive sample of respondents who are asked to verbalise their thinking, and who are probed about how they arrived at their answers. The goal is not to collect data on the topic; it is to collect data about the question.

What the research says

The authoritative synthesis is Beatty and Willis (2007), "Research Synthesis: The Practice of Cognitive Interviewing" (Public Opinion Quarterly, 71(2), 287-311). They define cognitive interviewing as the administration of draft survey questions while collecting additional verbal information about the survey responses, and review the field around three questions: what the dominant paradigms are, what design decisions a study requires, and how the resulting qualitative data should be evaluated. Their core conclusion is that cognitive interviewing is a powerful but under-standardised tool: it reliably surfaces problems, but the quality of what it surfaces depends heavily on probe design and on disciplined analysis rather than anecdote.

The method rests on Tourangeau's four-stage response model — comprehension, retrieval, judgement, and response — which gives interviewers a map of where a question can fail. A student may fail to comprehend an ambiguous term ("contact hours"), may be unable to retrieve the relevant memory (a lecture from twelve weeks ago), may judge using the wrong frame (rating the module against an easy elective rather than in absolute terms), or may struggle at the response stage (no scale point fits their view).

Two techniques dominate. In think-aloud, respondents narrate their thought process as they answer. In verbal probing, the interviewer asks targeted follow-ups — concurrently (right after the item) or retrospectively (after the whole questionnaire). Drawing on Willis (2005), Cognitive Interviewing: A Tool for Improving Questionnaire Design, the evidence is that the two techniques catch different defects: probing tends to expose comprehension problems, while think-aloud reveals more about retrieval and judgement. The practical implication is to combine them rather than rely on either alone.

A more recent experimental test corroborates the value of the method while sharpening expectations of it. Cyba and colleagues (2022), "An experimental test of the effectiveness of cognitive interviewing in pretesting questionnaires" (Quality & Quantity, 56), compared revised and unrevised items in a field setting and found that cognitive-interview-driven revisions changed response distributions and reduced item problems — evidence that the fixes are real, not cosmetic, though effects vary by item type.

Why it matters for course evaluation in practice

Course evaluation is unusually exposed to question-comprehension failure for three reasons. First, the vocabulary is institutional jargon ("learning outcomes", "constructive alignment", "formative assessment") that students may not parse as intended. Second, the recall window is long and the events are many — a single end-of-term survey asks students to summarise twelve weeks of varied experience in one judgement. Third, the stakes and framing are ambiguous: students often do not know whether they are rating the teacher, the module design, or their own engagement, and an untested item lets each respondent decide for themselves.

Cognitive interviewing addresses each. Pretesting a handful of students with think-aloud will quickly reveal whether "the assessment was fair" is read as "the marking was lenient", "the rubric was clear", or "the workload was reasonable" — three different constructs that a committee may have conflated. It will reveal whether a double-barrelled item ("the lecturer was well-prepared and approachable") forces students who agree with one half but not the other into an uninterpretable middle response. And it will reveal terms that simply do not land with a first-year or an international cohort.

Crucially, the sample is small and cheap. Willis (2005) notes that cognitive-interview studies typically use five to thirty respondents, recruited purposively to cover the cohorts you care about (first-year vs final-year, domestic vs international, STEM vs humanities). Five well-run interviews per major subgroup will catch the large, common comprehension defects that would otherwise corrupt thousands of responses.

Limitations and honest caveats

A PhD reader will rightly raise several objections. Small, purposive samples do not generalise statistically — cognitive interviewing identifies whether a problem can occur, not how prevalent it is in the population. A defect seen in two of eight interviews flags a hypothesis to fix or field-test, not a population estimate.

The method is interviewer-dependent. Beatty and Willis (2007) are explicit that under-standardisation is the central weakness of the field: leading probes, over-interpretation of a single comment, and inconsistent coding can manufacture findings that are not really there. Reliability improves with structured probe protocols, multiple coders, and written analysis rather than recollection.

Think-aloud is reactive. Asking people to narrate their reasoning can change that reasoning, so a problem visible in the lab may behave differently in a silent online survey. This is why the strongest evidence (e.g., Cyba et al., 2022) pairs cognitive pretesting with a quantitative field comparison.

Finally, cognitive interviewing tests the question, not the construct. It can confirm that students read "intellectually stimulating" consistently; it cannot tell you whether "intellectually stimulating" is a valid indicator of teaching quality. Pretesting complements, but does not replace, validity evidence such as that discussed in our coverage of what student evaluations actually measure.

How Koji incorporates this

Koji is built around the insight at the heart of cognitive interviewing: a single number is far more trustworthy when you can see the reasoning behind it. Several mechanisms map directly onto the findings above.

  • AI-moderated conversational interviews apply probing at scale. Where a traditional survey fields a fixed Likert item and hopes every student read it the same way, the Koji interview engine can follow a scale answer with a targeted, neutral probe — "What made you choose that?" — surfacing the interpretation behind the rating for every respondent, not just a pretest sample. This is verbal probing operationalised in production, and it is designed to mitigate (not eliminate) the comprehension and judgement failures Tourangeau's model predicts.
  • Structured question types reduce known defects by construction. Koji supports open_ended, scale, single_choice, multiple_choice, ranking, and yes_no items, which lets evaluation designers avoid double-barrelled and forced-frame wording by splitting a conflated item into clean components.
  • Automatic thematic analysis of open text turns probe responses into evidence. The qualitative material that cognitive interviewing generates by hand, Koji codes systematically across the full cohort, with quality scoring to down-weight low-effort answers — a partial, scalable analogue to the disciplined coding Beatty and Willis demand.
  • Pretesting support before a full launch. Because a Koji study can be run on a small purposive pilot first, an institution can use the conversational engine itself as a cognitive-pretest instrument: field the draft items to a handful of students, read how they reason, and revise before the census wave.

Koji is positioned as designed to mitigate comprehension and interpretation error, not to remove it — the human discipline of writing good items and reading the evidence honestly still matters. For teams who run customer and product research alongside course evaluation, the core Koji research platform at koji.so applies the same AI-moderated interview engine to discovery and concept-testing work.

Related Resources

References

  • Beatty, P. C., & Willis, G. B. (2007). Research Synthesis: The Practice of Cognitive Interviewing. Public Opinion Quarterly, 71(2), 287-311. https://doi.org/10.1093/poq/nfm006
  • Willis, G. B. (2005). Cognitive Interviewing: A Tool for Improving Questionnaire Design. Thousand Oaks, CA: Sage.
  • Cyba, K., et al. (2022). An experimental test of the effectiveness of cognitive interviewing in pretesting questionnaires. Quality & Quantity, 56. https://doi.org/10.1007/s11135-022-01489-4
  • Tourangeau, R., Rips, L. J., & Rasinski, K. (2000). The Psychology of Survey Response. Cambridge University Press.

Related articles

research-methods

What Can Open-Text Student Comments Tell You That Likert Scores Cannot?

A research-grounded guide to open-ended student comments in course evaluation: what Alhija and Fresko found about who writes them and what they contain, how thematic analysis surfaces issues numbers miss (Stupans et al.), how rare abusive comments actually are (Tucker), and how to turn free text into reliable evidence.

research-methods

What Do Student Evaluations Actually Measure? Marsh, the SEEQ, and the Case for Multidimensional Feedback

Herbert Marsh spent decades showing that well-built student evaluations are multidimensional, reliable, and stable — and that a single global score throws away most of what they can tell you. What his work establishes, where critics push back, and how to design feedback that is actually usable.

research-methods

Question Order and Context Effects: How the Sequence of Items Shapes Course-Evaluation Answers

The order in which you ask evaluation questions changes the answers you get. Drawing on Schwarz (1999), Strack, Martin & Schwarz (1988) and Tourangeau, Rips & Rasinski (2000), we explain part-whole and assimilation/contrast effects, what they do to your data, and how conversational evaluation reduces the damage.

research-methods

Should Every Point on a Course-Evaluation Scale Be Labelled? The Evidence on Verbal Anchors

Labelling every response option, not just the endpoints, tends to raise reliability and tame response styles — but it also interacts with the number of categories and the language of your students. What the rating-scale-format research says for course evaluation.