New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods10 min read

Question Order and Context Effects: How the Sequence of Items Shapes Course-Evaluation Answers

The order in which you ask evaluation questions changes the answers you get. Drawing on Schwarz (1999), Strack, Martin & Schwarz (1988) and Tourangeau, Rips & Rasinski (2000), we explain part-whole and assimilation/contrast effects, what they do to your data, and how conversational evaluation reduces the damage.

Koji Education Team

Product

Answer box

Respondents do not answer each survey item in isolation — they read every question in the context of the ones before it. Decades of survey-methodology research (Schwarz 1999; Strack, Martin & Schwarz 1988; Tourangeau, Rips & Rasinski 2000) show that question order systematically shifts answers: a specific item placed before a general one (a "part-whole" sequence) can either pull the general judgement toward it or push it away, depending on how respondents interpret the conversational intent. For course evaluation this means a fixed questionnaire can manufacture or mask an apparent "overall satisfaction" score simply through item sequencing. The robust mitigations are to put global judgements before specific probes, randomise or rotate item order where possible, and — most powerfully — let respondents construct their own narrative rather than march through a fixed list.

What the research says

Schwarz (1999), in American Psychologist, synthesised the cognitive and communicative processes behind self-reports. His core thesis: "self-reports of behaviours and attitudes are strongly influenced by features of the research instrument, including question wording, format, and context." Answering a survey is an act of communication, and respondents apply ordinary conversational rules — they assume each question is relevant, non-redundant, and informed by what was just asked. Those assumptions make earlier questions part of the context for later ones.

The canonical demonstration is Strack, Martin & Schwarz (1988) in the European Journal of Social Psychology. Students were asked two questions — how often they were dating, and how satisfied they were with life overall — in varying order. When the general life-satisfaction question came first, the dating question had little influence. When the specific dating question came first, the correlation between dating frequency and overall life satisfaction jumped dramatically: the just-activated information about dating was incorporated into the global judgement. Critically, the effect depended on conversational framing — when both questions were explicitly bracketed as one related block, respondents excluded the already-given dating information from the "overall" answer (to avoid being redundant), reversing the effect. Order did not just add noise; it changed the meaning respondents assigned to "overall."

Tourangeau, Rips & Rasinski (2000), in The Psychology of Survey Response, provide the integrative framework: answering any question involves comprehension, retrieval, judgement and response. Context effects arise at each stage. Earlier items can make certain information accessible (priming), set a standard of comparison (anchoring), or change what a later question is understood to mean. They distinguish assimilation effects (later answers pulled toward the context) from contrast effects (pushed away), and show the direction is governed by whether respondents treat prior information as included in or excluded from the target judgement — exactly the part-whole logic Strack et al. demonstrated.

For evaluation specifically, this literature predicts concrete artefacts: ask several pointed questions about a stressful exam, then ask "overall, how satisfied were you with this course?", and the global rating will be dragged down by the freshly-primed exam (assimilation) — or, if students read the overall question as "apart from the exam," dragged up (contrast). Either way, the "overall satisfaction" number is partly an artefact of sequence, not a stable summary of the course.

Why it matters for course evaluation in practice

Almost every institutional SET instrument is a fixed list of items presented in the same order to everyone. That design choice has measurement consequences:

  • The "overall" score is order-dependent. Where the global satisfaction item sits relative to specific probes (workload, assessment, the lecturer) can shift it in either direction. Two institutions with identical teaching but differently ordered forms can report different headline numbers.
  • Cross-year and cross-instrument comparisons are fragile. If a form is redesigned and items reordered, an apparent change in scores may be a context artefact, not a real change in student experience — a serious problem for trend reporting and accreditation evidence.
  • Negative-then-global sequencing depresses headline scores. Front-loading detailed problem-focused items primes dissatisfaction before the summary judgement, a common but rarely-noticed design flaw.
  • It interacts with other biases. Context effects compound satisficing and straightlining: a respondent who has anchored on an early impression is more likely to carry it mechanically down the rest of the form.

Practical design guidance that follows from the research: place broad/global judgements before narrow ones when you want an uncontaminated summary; group related items and signal the grouping explicitly so respondents know what to include or exclude; randomise the order of non-dependent blocks across respondents so order effects average out at the aggregate level; and treat any single fixed-order "overall satisfaction" mean with appropriate suspicion.

Limitations and honest caveats

  • Effect sizes vary with topic and salience. Order effects are largest for ambiguous, attitude-type questions where respondents lack a pre-formed answer; for concrete, well-rehearsed judgements they shrink. Not every course-evaluation item is equally vulnerable.
  • Much foundational evidence is from general social psychology, not course evaluation per se. The mechanisms are well established and domain-general, but the precise magnitude in a given SET instrument is an empirical question worth testing locally (e.g., via split-ballot experiments).
  • Randomisation has costs. Rotating item order complicates item-level longitudinal comparison and can confuse respondents if dependencies exist; it is a mitigation, not a free lunch.
  • Context effects cannot be fully eliminated in any sequential instrument — they are intrinsic to how people answer questions. The goal is to manage and disclose them, not to claim a "neutral" questionnaire exists.

How Koji incorporates this

Koji's conversational, AI-moderated format changes the structural conditions that produce the worst order effects:

  • Respondent-led narrative reduces fixed-sequence priming. Instead of forcing every student through the same item order, Koji's AI moderator can open with a broad, open_ended prompt and let the student raise what matters to them first — capturing the global impression before specific probes contaminate it, the sequence Strack et al. show is safest.
  • Adaptive, conversation-aware probing. Because follow-ups are generated in response to what the student actually said, Koji avoids the rigid "specific-then-global" templates that manufacture assimilation/contrast artefacts, and can explicitly bracket a topic ("setting the exam aside, how was the teaching?") to control inclusion/exclusion the way the research describes.
  • Order rotation for structured items. For the structured question types it does use (scale, single_choice, multiple_choice, ranking, yes_no), Koji can vary block order across respondents so residual order effects wash out in aggregate rather than biasing every record the same way.
  • Reasoning capture over single summary numbers. By recording why a student reached an overall view, Koji lets evaluators see whether a low "overall" reflects the whole course or a single primed grievance — distinguishing genuine summary judgement from a context artefact.
  • Comparable-design reporting. Koji's analytics flag when instrument design has changed between cycles, so QA teams do not misread an order-driven shift as a real change in student experience.

Honestly stated, no instrument escapes context effects entirely; a conversation has its own order. What Koji changes is who controls the sequence and how much context is captured, moving from a one-size-fixed-order form toward an adaptive exchange that is structurally less prone to manufacturing artefacts. The same conversational engine underpins customer and product research at koji.so, where order and framing effects distort feedback just as much.

Designing a split-ballot test

The cleanest way to find out how much order is distorting your own data is a split-ballot experiment, and it costs almost nothing to run inside an existing evaluation cycle. Randomly assign students to two (or more) versions of the same instrument that differ only in the placement of the global "overall satisfaction" item — version A asks it first, version B asks it last, after the specific probes. If the two versions produce materially different overall means for the same courses, you have direct, local evidence of an order effect rather than a borrowed estimate from the literature.

Three design points make such a test trustworthy. First, randomise at the respondent level so the two groups are comparable; do not give one cohort version A and another cohort version B, or course differences will confound the comparison. Second, hold everything else constant — wording, scale points, channel and timing — so sequence is the only thing that varies. Third, pre-register the comparison you will make (which means, which items) so the analysis is not a fishing expedition.

If you find a meaningful gap, the remedy follows directly: standardise on global-before-specific ordering for headline items, and rotate the order of independent specific blocks so residual effects average out across respondents. If you find no gap, you have earned the right to trust your overall score — and documented that diligence for accreditation review. Either outcome is more defensible than assuming a fixed-order form is neutral. The same split-ballot logic is how rigorous customer-research teams validate their own survey instruments before trusting the numbers.

Related Resources

References

Related articles

research-methods

Why Students Click Straight Down the Middle: Satisficing in Course Evaluations

A large share of students do not answer evaluation questions carefully — they straightline, speed through, and pick the easy option. Barge and Gehlbach (2012) showed this "satisficing" doesn't just add noise; it inflates your reliability and validity statistics, making bad data look good.

research-methods

What Do Student Evaluations Actually Measure? Marsh, the SEEQ, and the Case for Multidimensional Feedback

Herbert Marsh spent decades showing that well-built student evaluations are multidimensional, reliable, and stable — and that a single global score throws away most of what they can tell you. What his work establishes, where critics push back, and how to design feedback that is actually usable.

research-methods

Response Styles and Likert Scales: Why Cross-Cultural Evaluation Needs More Than Numbers

Acquiescence and extreme response styles vary systematically by culture (Harzing, 2006; Baumgartner & Steenkamp, 2001), which means raw Likert averages are not directly comparable across nationalities in Europe's multinational classrooms. This article explains the evidence and how to evaluate fairly across diverse cohorts.

research-methods

Online vs. Paper Course Evaluations: Do Lower Response Rates Mean Worse Data?

Online course evaluations consistently draw lower response rates than in-class paper forms, but the research shows the resulting scores are largely equivalent. Here is what Dommeyer, Nulty, and Stowell actually found, and what an adequate response rate really requires.