The Double-Barreled Question: Why "Knowledgeable and Approachable?" Is Two Questions, Not One
Course-evaluation items that ask about two things at once force students into a single, uninterpretable answer. Here is the measurement evidence on double-barreled questions and how to split them.
Koji Education Team
Product
In brief
A double-barreled question asks about two distinct things in one item ("The instructor was knowledgeable and approachable") but allows only one answer. A student who finds the lecturer expert but aloof has no valid response, so the data become ambiguous and the item's validity drops. Menold (2020), in two randomised experiments, showed that single-stimulus versions of double-barreled items produced different means and factor parameters, broke measurement invariance, and in at least one case achieved higher validity than the double-barreled original. The fix is simple and high-impact: split every "and"/"or" item into separate questions. Koji's structured question library and conversational probing are designed to surface and separate these conflated constructs.
What the research says
Double-barreled questions (DBQs) are flagged in every serious questionnaire-design text — Fowler's Improving Survey Questions (1995), Saris & Gallhofer's Design, Evaluation, and Analysis of Questionnaires for Survey Research (2014), and the Encyclopedia of Survey Research Methods — as one of the most common and most damaging wording errors. The theoretical problem is that the item violates a basic requirement of measurement: a single stimulus mapped to a single response. When two stimuli are bundled, respondents may attend to one and ignore the other, attend to both and average, or be unable to answer — and the analyst cannot tell which.
The strongest recent empirical test is Natalja Menold (2020), "Double Barreled Questions: An Analysis of the Similarity of Elements and Effects on Measurement Quality" (Journal of Official Statistics, 36(4), 855–886). Menold ran two randomised split-ballot experiments comparing DBQs against revised versions in which one stimulus was retained and the other dropped. Key findings:
- Observed means and the parameters of latent-variable models differed across versions — i.e. the double-barreled wording changed the numbers, not just the interpretation.
- Metric and scalar measurement invariance did not hold between the DBQ and single-stimulus versions, meaning scores were not on a comparable scale.
- At least one single-stimulus version showed higher validity than the double-barreled form.
- The damage depended on how similar the two bundled elements were: when the two ideas were closely related, respondents could cope better; when they were distinct, the item degraded sharply.
That last point is important for course evaluation, where bundled attributes are frequently not the same construct — "knowledgeable" (competence) and "approachable" (warmth) are two well-established and partly independent dimensions of how students perceive teaching.
Why it matters for course evaluation in practice
Double-barreled items are endemic in teaching surveys, often hiding behind an innocent "and":
- "The instructor was well-prepared and enthusiastic."
- "Assessments were fair and clearly explained."
- "The course was interesting and useful for my career."
- "Feedback was timely and helpful."
Each conflates two things students can rate very differently. The consequences:
- Uninterpretable scores. A 3/5 on "fair and clearly explained" could mean "fair but unclear", "clear but unfair", or "moderately both". Quality committees cannot act on it, because they cannot tell what to fix.
- Hidden in the lowest-scoring items. Bundled assessment-and-feedback items are often where evaluations score worst (see our note on why assessment and feedback score lowest); double-barreling makes the diagnosis impossible precisely where you most need it.
- Broken comparability. Menold's invariance result means you cannot safely compare a double-barreled item across cohorts or against a single-stimulus version used elsewhere — undermining benchmarking and trend analysis.
- Contaminated factor structure. If you run reliability or factor analysis (Cronbach's alpha, dimensionality checks), DBQs distort the loadings, as Menold's latent-variable parameters showed.
Limitations and honest caveats
A rigorous reader should note the boundaries of the evidence:
- Not every "and" is double-barreled. "Clear and well-organised" may, for some respondents, name a single coherent impression rather than two separable judgments. Menold's own finding — that closely-related elements degrade quality less — means the severity is item-specific, not automatic.
- Splitting has costs. Breaking one item into two lengthens the questionnaire, and questionnaire length itself reduces data quality through fatigue and breakoff (see how long should a course evaluation be). There is a genuine trade-off between item purity and survey burden.
- Generalizability. Menold's experiments were general-population survey studies, not course evaluations specifically; the mechanism transfers cleanly, but the exact effect sizes for a given teaching item should be checked locally.
- Detection is harder than it looks. Some double-barreling is subtle ("the lecturer made the subject accessible and engaging"). Cognitive interviewing is the most reliable way to catch it before fielding.
How Koji incorporates this
Koji is designed to mitigate the double-barreled problem rather than claim to abolish it:
- A structured, single-construct question library. Koji's question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) are built around one stimulus per item, nudging designers to ask about "preparedness" and "enthusiasm" as separate scale questions instead of bundling them. Where a template ships with a bundled item, it can be split into two cleanly comparable questions.
- Conversational unbundling. This is the distinctive mechanism. If a student gives a middling answer to a question that touches two ideas, Koji's AI-moderated interview can follow up to separate them ("You rated feedback a 3 — was that about how quickly it came, or how useful it was?"). The conversational layer effectively de-conflates in real time what a static double-barreled item would permanently fuse.
- Thematic analysis that separates dimensions. Koji's automatic thematic analysis of open text can distinguish "competence" comments from "warmth/approachability" comments even when a single Likert item blurred them, recovering the two signals Menold shows a DBQ collapses.
- Bias- and quality-aware reporting. Koji's quality scoring is designed to flag items and responses that look internally inconsistent, helping committees notice when an item is doing two jobs at once.
As always we frame this honestly: good upstream item design — one idea per question — remains the primary defence, and Koji's tooling supports that discipline rather than substituting for it. The same AI-moderated interview engine powers Koji's core research platform at koji.so for product and customer research, where double-barreled questions are an equally common source of muddy data.
Practical checklist
- Search every item for "and", "or", commas, and slashes — they are the usual tells.
- For each, ask: could a student feel one way about the first part and another way about the second? If yes, split it.
- Prioritise splitting the items you act on most (assessment, feedback, organisation).
- Where you must keep length down, drop a low-value item rather than bundling two into one.
- Pre-test suspected double-barreled items with cognitive interviews.
Related Resources
- Designing Open-Ended Questions That Get Useful Answers
- Cognitive Interviewing for Course-Evaluation Questions
- Agree/Disagree or Item-Specific Scales
- Single-Item vs Multi-Item Measures
- What Do Student Evaluations Actually Measure?
- Response-Order Effects in Course Evaluations
A worked audit: spotting and splitting
Here is how a short audit looks in practice. Each original item is scanned for conjunctions and for two separable ideas, then split.
- Original: "The instructor was well-prepared and enthusiastic." → Split: (1) "The instructor was well-prepared for each class." (2) "The instructor taught with enthusiasm." These are competence and affect — students routinely rate them differently, exactly the case Menold shows degrades most.
- Original: "Assessments were fair and clearly explained." → Split: (1) "The marking criteria were explained clearly in advance." (2) "Marking was fair relative to those criteria." A student can find criteria crystal clear yet feel marking was harsh; the bundled item hides which.
- Original: "The course was interesting and useful for my career." → Split: (1) "I found the course intellectually interesting." (2) "I expect the course to be useful for my career." Interest and instrumental value are distinct constructs and often diverge for required courses.
Note the borderline case: "clear and well-organised" may, for many students, name one coherent impression of structure rather than two judgments. Menold's finding — that closely-related elements degrade quality less — means this one is lower-risk and a candidate to keep if length is tight. The audit is therefore not mechanical: it asks, item by item, whether a student could plausibly feel one way about each half.
Managing the length trade-off
Splitting doubles item count, and longer questionnaires lose data to fatigue and breakoff. Three tactics keep the instrument lean: drop genuinely low-value items rather than bundling them; reserve splitting for the items you actually act on (assessment, feedback, organisation); and use a short open-text or conversational layer to capture the secondary dimension without adding another scale item. The goal is one idea per question where it matters most, not maximal granularity everywhere.
References
- Menold, N. (2020). Double barreled questions: An analysis of the similarity of elements and effects on measurement quality. Journal of Official Statistics, 36(4), 855–886. https://doi.org/10.2478/jos-2020-0041
- Fowler, F. J. (1995). Improving Survey Questions: Design and Evaluation. Sage.
- Saris, W. E., & Gallhofer, I. N. (2014). Design, Evaluation, and Analysis of Questionnaires for Survey Research (2nd ed.). Wiley.
- Lavrakas, P. J. (Ed.). (2008). Double-barreled question. In Encyclopedia of Survey Research Methods. Sage. https://doi.org/10.4135/9781412963947
Related articles
How to Design Open-Ended Course Evaluation Questions That Get Useful Answers
Open-ended evaluation questions usually fail not because students have nothing to say, but because the prompt asks for too little. The survey-methodology evidence on answer-box size, verbal instructions, and specific framing — and what it means for the comment box.
What Do Student Evaluations Actually Measure? Marsh, the SEEQ, and the Case for Multidimensional Feedback
Herbert Marsh spent decades showing that well-built student evaluations are multidimensional, reliable, and stable — and that a single global score throws away most of what they can tell you. What his work establishes, where critics push back, and how to design feedback that is actually usable.
Is One "Overall" Question Enough? Single-Item vs Multi-Item Measures in Course Evaluation
Can a single global question replace a multi-item battery in course evaluation? Gogol et al. (2014) and the single-item-measure literature show when one item is defensible and when it is not.
Response-Order Effects: Does Where an Answer Sits Change How Students Rate Your Course?
The order in which answer options appear can shift course-evaluation responses independent of what students think. We unpack Krosnick & Alwin (1987) on primacy and recency, why it matters for instrument design, and how to limit it.