New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
analysis-reporting10 min read

The Course Reviewed Straight After a Brilliant One Looks Worse: Contrast Effects and Narrow Bracketing in Evaluation Review

Course-evaluation bias research focuses on the student filling in the form. The evidence on sequential judgement says the committee reading twenty reports in an afternoon is biased too — by what it read immediately before.

Koji Education Team

Product

In short: When professionals evaluate a stream of cases in sequence, their judgement of each case is distorted by the cases just before it. Simonsohn and Gino (2013), analysing more than 9,000 MBA admission interviews, found that interviewers who had already rated several applicants highly that day became reluctant to do so again — narrow bracketing against an expected distribution. Bhargava and Fisman (2014) and Chen, Moskowitz and Shue (2016) document the same contrast and negative-autocorrelation patterns in speed dating, asylum courts, loan approval and umpiring. If your teaching committee reads course-evaluation reports back-to-back, the order of the pile is a variable in the outcome.

The bias nobody screens for

The course-evaluation literature is unusually well developed on the biases of the respondent: gender, accent, attractiveness, grading leniency, class size, discipline. Institutions have absorbed this. Reports carry caveats; committees are trained to discount small samples.

Almost nothing is said about the biases of the reader. Yet the reading is where the consequence happens. A programme director triaging thirty module reports on a Tuesday afternoon, or a promotion panel working through a stack of teaching dossiers, is doing exactly the task that the judgement literature has shown to be systematically distorted — evaluating a sequence of cases under time pressure, with an implicit sense of how many should come out favourably.

What the research says

Narrow bracketing: the quota you did not know you were applying

Simonsohn and Gino (2013) analysed ten years of MBA admissions interviews — over 9,000 of them — at a top school where interviewers see only a handful of candidates per day. They predicted and found that interviewers bracket narrowly: they assess each day's small subset as if it should resemble the overall distribution of applicant quality. An interviewer who has already given three strong recommendations that day becomes measurably less likely to give a fourth.

The mechanism is not fatigue and not learning. It is a mistaken application of the expected long-run distribution to a short local run. The consequence is that an applicant's outcome depends partly on who else happened to be scheduled that morning.

Contrast effects: the case before you moves the standard

Bhargava and Fisman (2014) used the quasi-experimental structure of speed dating — participants meet partners in effectively random sequence — to isolate contrast effects. A more attractive prior partner reduced the likelihood of an affirmative decision about the next one. Crucially, the effect was confined to recent interactions, which is the signature of a perceptual error rather than learning or a deliberate quota. (In their data the effect was driven almost entirely by male evaluators, a heterogeneity worth noting rather than over-reading.)

Negative autocorrelation in high-stakes professional judgement

Chen, Moskowitz and Shue (2016) extended this to three consequential field settings: US asylum court decisions, loan officer reviews, and Major League Baseball umpire pitch calls. Across all three they found negatively autocorrelated decisions — a favourable decision made one case less likely to be followed by another favourable decision, beyond what case merits explained. They attribute this to the gambler's fallacy: decision-makers underestimate how often streaks occur by chance and correct against them.

Their moderators map uncomfortably well onto a teaching committee. The effect was stronger among less experienced decision-makers, after longer streaks, when consecutive cases were similar in characteristics or close in time, and when the decision-maker faced weaker incentives for accuracy. A newly appointed director of studies reviewing a batch of similar modules in one sitting, with no external check on the judgements, sits at the intersection of every one of those conditions.

Why it matters for course evaluation in practice

Course-evaluation review has the exact structural features these studies identify as risky: a queue of superficially similar cases, judged in one sitting, by someone with a rough prior about how many modules "should" be flagged as needing action.

Five defences follow directly from the evidence.

1. Randomise and record the review order. If order is a variable, treat it as one. Randomise the sequence in which modules are reviewed each cycle, and log the order. Reviewing consistently in alphabetical or timetable order builds a systematic advantage for whoever sits early in the list.

2. Break the batch. The effects strengthen with streak length and with proximity in time between similar cases. Reviewing eight modules in a sitting rather than thirty is a cheap structural mitigation.

3. Judge against criteria, not against the previous case. Contrast effects operate when the comparison standard is implicit and local. An explicit, absolute rubric — this is what "action required" means, stated in advance — anchors the judgement outside the sequence. This is the practical value of criterion-referenced rather than norm-referenced reporting.

4. Refuse implicit quotas. If a process expects "the bottom 10% of modules" to be identified, it has built narrow bracketing into policy. Some cycles genuinely contain no failing modules and some contain five. See why ranking instructors by evaluation scores misclassifies them.

5. Audit for it. You can test for this in your own data. Regress the review outcome on the outcome of the immediately preceding case in the review sequence, controlling for the module's own scores. A significant negative coefficient is evidence of the effect operating in your process. This is a genuinely tractable internal analysis and few institutions have run it.

Limitations and honest caveats

This section matters more than usual here, because the sequential-judgement literature contains one of the best-known cautionary tales in behavioural science.

  • The headline example of this genre is contested. Danziger, Levav and Avnaim-Pesso (2011) reported that Israeli parole grant rates fell from roughly 65% to near zero across a decision session and reset after food breaks. Weinshall-Margel and Shapard (2011) showed the case ordering was not random — unrepresented prisoners, who prevail far less often, were systematically heard at the end of sessions. The original authors replied that the trend persists with representation controlled, and Glöckner (2016) argued by simulation that the effect magnitude is implausibly large for a psychological mechanism. The dispute is unresolved. Anyone citing sequential-judgement research in a policy paper should cite this controversy alongside it.
  • No study here is about course evaluation. Every finding is transferred from admissions, dating, courts, lending and sport. The structural analogy is strong but it is an analogy. Direct replication in teaching-quality review panels does not, to our knowledge, exist — which is a research gap worth naming rather than papering over.
  • Effect sizes are modest. These are real but small distortions. They matter most where decisions are close to a threshold and consequences are high, such as promotion or probation, and matter least where an evaluation is purely formative.
  • Some sequencing effects are legitimate. A reviewer who genuinely calibrates over a cycle is learning, not erring. The distinguishing evidence in Bhargava and Fisman is that the effect decayed with recency — a calibration effect would not.
  • Mitigations have costs. Randomised order complicates scheduling; smaller batches take longer. These are real trade-offs, not free wins.

How Koji incorporates this

The bias described here lives in the review workflow, not in the instrument — so the honest claim is narrow. Koji is designed to reduce the conditions under which sequential distortion operates, not to correct the reviewer's cognition.

  • Criterion-anchored reporting. Koji's reporting is built to present a module against stated thresholds and its own history rather than as a rank position within the cycle's batch, removing the implicit local comparison that contrast effects feed on.
  • Bias-aware reporting surfaces sample size, response rate and uncertainty prominently, so a reviewer comparing two consecutive modules sees whether the apparent difference is within noise — the check that a fast sequential read otherwise skips. See why small mean differences are usually noise.
  • Automatic thematic analysis of open text means the substance of a module's feedback is summarised into themes with supporting quotations. A reviewer engaging with named themes is judging content against criteria rather than eyeballing a score relative to the last report in the pile.
  • AI-moderated conversational interviews produce evidence that is harder to reduce to a single comparable number, which structurally resists the ranking behaviour that narrow bracketing exploits.
  • Closing-the-loop action tracking records why a module was flagged and what was done, creating the audit trail that makes the order-effect analysis described above possible after the fact.

None of this eliminates the bias. Sequential judgement effects are properties of human reviewers, and the mitigations that work are procedural — randomised order, smaller batches, explicit rubrics. Koji is designed to make those procedures easier to run, and to make the resulting decisions auditable.

Institutions that run the same kind of batched evidence review outside education — customer research portfolios, product feedback triage — face an identical exposure, and the same interview engine is available at koji.so.

Related resources

References

Related articles