Stop Arguing About Wording — Test It: Split-Ballot Experiments for Course-Evaluation Questions
Committees spend hours debating whether to word an evaluation item one way or another. Schuman and Presser showed that small wording changes can move survey answers by double-digit margins — and that the way to settle the debate is a randomised split-ballot experiment, not opinion.
Koji Education Team
Product
In short
When a committee cannot agree whether to word a course-evaluation item one way or another, the answer is not to argue harder — it is to run a split-ballot experiment: randomly assign half your students to each version and measure the difference. Schuman and Presser's landmark programme of survey experiments demonstrated that seemingly trivial changes to question form, wording and order can shift responses by margins large enough to change conclusions. The split-ballot design is the gold-standard, low-cost way to find out whether your wording choice matters — before you field it on a whole cohort and lock a decade of trend data to a flawed item.
What the research says
A split-ballot experiment (also called a randomised questionnaire experiment or embedded survey experiment) is elegantly simple: you create two or more versions of a question, randomly assign respondents to receive one version, and compare the distributions of answers. Because assignment is random, any systematic difference in the answers is caused by the wording difference and nothing else. It is the direct experimental analogue of the more familiar A/B test.
The foundational evidence comes from Howard Schuman and Stanley Presser's Questions and Answers in Attitude Surveys: Experiments on Question Form, Wording, and Context (1981; reissued by Sage in 1996). Across a decade of embedded split-ballot experiments, they demonstrated repeatedly that survey answers are not simply "read off" from stable internal attitudes; they are partly constructed in response to the exact question asked. Their classic findings include: offering an explicit "no opinion" or middle option substantially changes the proportion choosing it; whether a question is asked in an "allow" versus "forbid" frame produces different results for what is logically the same attitude; the presence and order of response options shift answers; and open versus closed formats yield different top answers because closed lists constrain what respondents consider. The magnitudes were often large — differences of 10–20 percentage points from a single wording change were not unusual.
The method and its lessons were consolidated and extended by later methodologists. Krosnick and Presser (2010), in their chapter "Question and Questionnaire Design" in the Handbook of Survey Research, treat the split-ballot experiment as the standard empirical tool for choosing between question versions, and lay out the design principles that minimise these effects. Fowler (1995), in Improving Survey Questions, similarly argues that ambiguity and framing effects are best diagnosed empirically rather than by armchair reasoning. The through-line of this literature is a humbling one for questionnaire designers: intuition about how a question "will be read" is unreliable, and the only dependable arbiter is a randomised test on real respondents.
This matters because the effects Schuman and Presser catalogued are exactly the ones that recur in course evaluation: whether to include a neutral midpoint or a don't-know option, whether to use agree/disagree or item-specific formats, how response order affects ratings, and whether an item is double-barrelled. Each of these is a design fork that a split-ballot experiment can resolve empirically.
Why it matters for course evaluation in practice
Course-evaluation instruments are unusually high-stakes to get right, for a reason that survey researchers in other fields rarely face: the trend line. A market researcher can revise a question next quarter. A university that changes an evaluation item breaks comparability with years of historical data and often with sector benchmarks. That asymmetry makes it far more important to test wording before it becomes the institutional standard — and split-ballot experiments are the tool for doing so.
Practically, a QA office can embed split-ballot tests in several ways:
- Piloting a new item. Before adopting a new question institution-wide, field two candidate wordings to a random half of one term's cohort each. If the distributions differ meaningfully, you have learned that the wording is doing work — and you can choose the version that best matches the construct you intend to measure.
- Deciding a contested design choice. Instead of a committee debating whether to drop the neutral midpoint, run it: half the cohort sees a five-point scale, half sees a four-point forced-choice scale, and you observe how many students would have sat in the middle and where they go instead.
- Quantifying a known risk before a redesign. If you suspect an item is double-barrelled or leading, a split-ballot against a cleaned-up version tells you how much the flaw is actually distorting your data — which informs whether a disruptive change to the trend line is worth it.
- Validating a translation. For multilingual European institutions, a split-ballot can test whether two language versions of an item behave equivalently, complementing formal translation procedures.
Split-ballot experiments also sit naturally alongside — and after — cognitive interviewing. Cognitive interviews tell you why an item might be misread by probing a handful of students in depth; a split-ballot tells you whether and how much the wording actually changes answers at scale. Used together, one diagnoses and the other quantifies.
Limitations and honest caveats
Split-ballot experiments are powerful but not a cure-all, and a rigorous evaluator should keep several limits in view:
- They tell you that answers differ, not which is "correct." A split-ballot shows that version A and version B produce different distributions. Deciding which better measures the intended construct still requires a validity argument — the experiment informs that judgement but does not settle it. This is where a validity framework remains essential.
- Adequate power is needed. Detecting a real but modest wording effect requires a sufficiently large sample split across arms. In small modules, a single-term split-ballot may be underpowered; you may need to pool across modules or terms.
- Randomisation must be genuine and concealed. If students can see that peers received a different question (e.g. in a shared paper setting), or if assignment is not truly random, the comparison is compromised. Online administration makes clean randomisation easy; paper does not.
- Ethical and governance clearance. Even low-risk methodological experiments on students may require institutional approval and transparency about why some students see a different form. This is usually straightforward but should not be skipped.
- You can only test what you think to test. A split-ballot answers a specific pre-specified question; it will not reveal wording problems you did not anticipate. That is precisely why it pairs with exploratory methods like cognitive interviewing rather than replacing them.
How Koji incorporates this
Koji for Education is built on infrastructure that makes split-ballot experimentation practical rather than a once-a-decade research project — because it administers evaluations digitally and adaptively:
- Randomised assignment across question variants. Because Koji delivers evaluations online and controls question presentation, it can randomly assign students to alternative wordings, scales, or item orders and hold the assignment stable within a respondent — the essential machinery of a split-ballot.
- Structured question types make variants easy to define. Koji's open_ended, scale, single_choice, multiple_choice, ranking and yes_no types let a QA team define two clean, comparable versions of an item (for example, a five-point scale versus a four-point forced choice, or an agree/disagree versus item-specific frame) without bespoke engineering.
- Built-in analytics quantify the difference. Koji's reporting compares response distributions across cohorts and segments, so the output of a split-ballot — the size and direction of the wording effect — is surfaced directly rather than requiring a separate statistical pipeline.
- AI-moderated probing complements the experiment. Where a split-ballot quantifies an effect, Koji's AI-moderated conversational follow-ups can probe why students interpreted an item as they did, blending the diagnostic power of cognitive interviewing with the quantification of an experiment in a single administration. This is designed to help teams choose wording that measures the intended construct, not merely the wording that scores highest.
- Protecting the trend line. By enabling controlled pilots on a random subset before an institution-wide change, Koji is designed to help QA offices avoid the costly mistake of locking a flawed item into a multi-year trend series.
Koji frames these as tools to test and improve evaluation instruments, never as a guarantee of a "bias-free" question — the research is clear that no wording is neutral, only better or worse understood. Teams running the same experimental discipline on customer or product surveys can use Koji's core research platform at koji.so, which applies the identical randomisation and analytics engine outside the education context.
Related resources
- Before You Field It, Test It: Cognitive Interviewing for Course-Evaluation Questions
- The Double-Barreled Question: Why "Knowledgeable and Approachable?" Is Two Questions
- Question Order and Context Effects in Course-Evaluation Surveys
- Agree/Disagree or Item-Specific? The Scale Choice That Quietly Degrades Data
- Should a Course-Evaluation Scale Have a Neutral Midpoint?
References
- Schuman, H., & Presser, S. (1981). Questions and Answers in Attitude Surveys: Experiments on Question Form, Wording, and Context. Academic Press. (Reissued 1996, Sage.)
- Krosnick, J. A., & Presser, S. (2010). Question and Questionnaire Design. In J. D. Wright & P. V. Marsden (Eds.), Handbook of Survey Research (2nd ed., pp. 263–314). Emerald.
- Fowler, F. J. (1995). Improving Survey Questions: Design and Evaluation. Sage.
- Tourangeau, R., Rips, L. J., & Rasinski, K. (2000). The Psychology of Survey Response. Cambridge University Press. https://doi.org/10.1017/CBO9780511819322
Related articles
The Double-Barreled Question: Why "Knowledgeable and Approachable?" Is Two Questions, Not One
Course-evaluation items that ask about two things at once force students into a single, uninterpretable answer. Here is the measurement evidence on double-barreled questions and how to split them.
Before You Field It, Test It: Cognitive Interviewing for Course-Evaluation Questions
Why the wording of a course-evaluation item should be cognitively pretested before it reaches students, what Beatty and Willis (2007) established about think-aloud and verbal probing, and how Koji operationalises probing at scale.
Agree/Disagree or Item-Specific? The Scale Choice That Quietly Degrades Course-Evaluation Data
"The lecturer was well organised: Strongly disagree to Strongly agree" feels natural, but the agree/disagree format invites acquiescence and lower data quality. Saris et al. (2010) on why item-specific scales measure better.
Response-Order Effects: Does Where an Answer Sits Change How Students Rate Your Course?
The order in which answer options appear can shift course-evaluation responses independent of what students think. We unpack Krosnick & Alwin (1987) on primacy and recency, why it matters for instrument design, and how to limit it.