New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Evaluation bias9 min read

Question Order Effects: How the Sequence of Your Evaluation Questions Quietly Changes the Answers

The order you ask course-evaluation questions in is not neutral. Decades of survey research show that an earlier item can prime, anchor, or contrast with a later one — shifting scores without anything about the teaching changing. Here is what the evidence says, and how to design around it.

Koji for Education

Editorial Team · June 30, 2026

Answer up front: The sequence in which you ask course-evaluation questions changes the answers you get. A general question placed before specific ones is answered differently than the same question placed after them; an item near the top of a list is endorsed at a different rate than the identical item near the bottom; and a strongly worded question early in the form can colour everything a student reports afterwards. None of this involves the teaching changing — only the questionnaire. For institutions that compare scores across courses, terms, and instructors, an un-standardised question order is a quiet source of non-comparability hiding inside data that looks clean.

Why order is not a neutral design choice

Survey responses are not retrieved from memory like files from a cabinet. Respondents construct an answer on the spot, drawing on whatever considerations are most accessible at that moment — and the preceding questions are the single biggest determinant of what is accessible. This is the core finding of decades of cognitive survey methodology, summarised in Norbert Schwarz and colleagues' work on context effects and in Krosnick and Presser's canonical chapter on question and questionnaire design. When you ask a student "How clear was the assessment?" immediately before "How would you rate the course overall?", you have made assessment-clarity artificially accessible, and the overall rating moves toward it.

Three mechanisms matter for course evaluation specifically.

1. Assimilation and contrast (the part–whole problem)

Tourangeau and Rasinski (1988) and Schwarz, Strack and Mai (1991) showed that the relationship between a specific question and a general one depends entirely on which comes first. Ask the general question first ("Overall, how satisfied are you with this course?") and students answer it broadly. Ask several specific questions first — workload, feedback, the lecturer's availability — and the general question is now interpreted in light of them. Depending on framing, students either assimilate (pull the general rating toward the specifics they just considered) or contrast (treat the general question as "apart from everything I just mentioned, how was the rest?"). The same overall-satisfaction item can therefore yield systematically different averages purely as a function of what preceded it.

2. Primacy and recency in item and option lists

Krosnick and Alwin (1987) demonstrated that the position of a response option changes how often it is chosen. In visually presented, self-administered surveys — which is what almost every modern course evaluation is — respondents tend to a primacy effect: options and items near the top of a list are endorsed more often, because students engage most deeply with what they read first and satisfice thereafter. In orally administered surveys, the pattern flips to a recency effect. For a long evaluation form, the items you place first get more careful, more differentiated answers; items buried at the bottom get flatter, more agreeable ones. If your "overall" question sits at the end of a fatiguing 30-item form, you are measuring something different than if it sits at the top.

3. Priming and mood carry-over

A single emotionally loaded early question can set a frame. If the first thing you ask is "Did the course workload feel excessive?", you have primed a grievance lens; subsequent neutral items inherit some of that negativity. Conversely, opening with "What did you enjoy most?" primes a charitable lens. This is not hypothetical: a 2025 randomised controlled trial on survey order and cognitive-load measures found that simply reordering instruments shifted the subjective ratings respondents gave.

What this means for the numbers on your dashboard

The practical consequence is uncomfortable for anyone running institutional comparisons. Two departments that use the same questions in a different order are not running the same instrument. A year-on-year trend can move because someone reordered the form during a "minor revision." An instructor flagged for a low "overall" score may simply have had that item positioned after the workload question while a colleague had it first. When you aggregate to a programme mean, these order artefacts do not cancel out — they bias in a consistent direction set by the form's design.

This compounds the more familiar problems we have written about elsewhere: that averaging Likert scores discards most of the information in the distribution, that the halo effect makes ten questions behave like one impression, and that acquiescence and straightlining flatten genuine signal. Question order is the upstream cousin of all three: it shapes the considerations students bring to every later item.

"But doesn't randomising the order just fix it?"

This is the strongest counterargument, and it is half right. Randomising item order across respondents does prevent a systematic bias from contaminating the aggregate — the assimilation pull on one student is offset by a contrast push on another, and the group mean is cleaner. Many survey platforms now offer this. So why is order still a problem?

Three reasons. First, most institutions do not randomise — they ship a fixed PDF-era form, identical for everyone, which locks the bias in rather than averaging it out. Second, randomisation increases noise even as it removes bias: every student now answers a slightly different instrument, so individual responses are less comparable and you need larger samples to detect real effects — a serious problem given the small class sizes typical in higher education. Third, randomisation does nothing for the part–whole problem when the logic of the form requires a fixed sequence (you cannot sensibly ask "overall satisfaction" before the student has thought about anything). Randomisation is a patch on a static-form paradigm, not a solution to the underlying fact that a one-size questionnaire forces every student down the same cognitive path.

It is also worth being precise about scale. Order effects are real but usually modest — they move averages by fractions of a scale point, not whole points. The honest framing is not "your data is worthless" but "your data carries a measurement artefact of roughly the same size as the differences you routinely treat as meaningful when you rank instructors." When the signal you are chasing is a 0.2-point gap, a 0.1-point order artefact is not negligible.

How a conversational approach sidesteps the problem

Order effects are a symptom of a deeper design constraint: the fixed-form survey asks every student the same questions in a pre-set sequence, so the only way to manage context is to randomise or to accept the bias. An AI-moderated conversational interview works differently. Instead of marching every respondent through an identical item list, Koji for Education conducts a structured but adaptive interview: it opens with a neutral, open-ended prompt, lets the student raise what is salient to them, and then probes specifics in response to what they actually said — rather than priming them with a pre-ordered checklist of grievances and virtues.

Concretely, this matters in a few ways. Because the opening is open-ended rather than a loaded closed item, there is no fixed early question priming a frame for everything after it. Because the moderation is standardised by the same AI across every interview, there is no human-moderator inconsistency layered on top of order effects. And because Koji supports six structured question types — open_ended, scale, single_choice, multiple_choice, ranking, and yes_no — alongside automatic thematic analysis of the open text, the "overall" judgement can be reconstructed from what students actually emphasised, not anchored to whichever specific item happened to come first. Koji is careful here: this mitigates order and priming artefacts; it does not eliminate context effects entirely, because any act of asking shapes the answer. The claim is that an adaptive interview gives you more levers to manage that than a frozen form does.

Teams who run general customer and user research face the identical questionnaire-design trap, which is why the same adaptive interview engine powers the main Koji platform for product and market research.

A short, practical checklist

Whatever tool you use, you can reduce order artefacts today:

  • Standardise the sequence across every course you intend to compare. An un-versioned, inconsistently ordered form makes benchmarking meaningless before bias even enters.
  • Put the global/overall question first, before specific items, if you want a holistic judgement — or last and clearly framed ("aside from the points above…") if you want a residual one. Decide which you mean.
  • Front-load the items you care most about in self-administered forms, where primacy and fatigue degrade later answers.
  • Avoid opening with an emotionally loaded item; lead with a neutral or open prompt.
  • If you randomise, log it and account for the added noise in your sample-size planning.
  • Treat open-text first. Letting students say what matters before you show them your categories captures their salience instead of yours.

The uncomfortable truth is that there is no order-free survey. Every sequence encodes an assumption about what students should think about first. The question is whether you have chosen that sequence deliberately and consistently — or inherited it from a form nobody has audited in a decade.


Koji for Education runs AI-moderated, GDPR/AVG-compliant course evaluations that probe beyond a fixed question list. See how the conversational approach works.