Does the Number Anchor the Committee Before It Reads a Word? Anchoring Bias in Evaluation Review
When a review panel sees an instructor's 3.8 mean before reading the comments, that number quietly pulls every later judgement toward it. What the anchoring literature says, and how to sequence evaluation review to resist it.
Koji Education Team
Product
In brief: Anchoring is one of the most robust findings in judgement research: an initial number pulls subsequent estimates toward it, even when the number is irrelevant, random, or known to be arbitrary — and expertise does not immunise you (Tversky & Kahneman, 1974; Englich, Mussweiler & Strack, 2006). When an evaluation committee reads an instructor's numeric mean before the open-text comments and contextual evidence, that mean becomes an anchor that biases the whole review. The mitigation is procedural: sequence the evidence deliberately, and treat the number as one input read last, not the frame set first.
The question this answers
Most course-evaluation review happens in a predictable order. A committee — a head of department, a teaching-quality panel, a promotion board — opens the report and the first thing on the page is a number: an overall mean, a percentile, a red-amber-green flag. Only afterwards do they read the comments, the response rate, the class size, the difficulty of the cohort. This article asks a narrow but consequential question: does seeing the number first bias how the rest of the evidence is read? The judgement-and-decision-making literature says almost certainly yes — and names the mechanism.
What the research says
Anchoring is foundational and general. Amos Tversky and Daniel Kahneman's landmark paper (1974, Science, 185(4157), 1124–1131) demonstrated that people estimate quantities by starting from an initial value and adjusting — but the adjustment is typically insufficient, so the final estimate stays biased toward the anchor. In their famous demonstration, spinning a wheel of fortune to generate a random number before asking participants to estimate the percentage of African nations in the UN measurably shifted the estimates: a higher random number produced a higher estimate. The anchor was transparently irrelevant, and it worked anyway.
Expertise does not protect you — and neither does knowing the anchor is random. The most alarming evidence for evaluation panels comes from Englich, Mussweiler and Strack (2006, Personality and Social Psychology Bulletin, 32(2), 188–200). Experienced legal professionals set criminal sentences that assimilated toward a sentencing demand even when that demand came from an irrelevant source, was described as randomly determined, or was generated by the participants themselves throwing dice. Expertise and experience did not reduce the effect. The proposed mechanism is selective accessibility: an anchor makes anchor-consistent information more cognitively available — a high anchor brings incriminating arguments to mind, a low anchor brings mitigating ones. Translated to evaluation: a low score primes a reviewer to notice the critical comments; a high score primes them to notice the praise.
The mechanism is about what comes to mind, not just arithmetic. Mussweiler and Strack's selective-accessibility work reframes anchoring as a confirmatory search: once the anchor is present, the mind tests whether the target is consistent with it, and that hypothesis-testing surfaces anchor-consistent evidence. This is why anchoring interacts so dangerously with open-text comments — the number does not just bias the final rating, it biases which comments the reader weights.
The convergent finding across four decades: an initial number sets an unearned frame, adjustment from it is insufficient, the effect survives irrelevance and randomness, and professional expertise offers little defence.
Why it matters for course evaluation in practice
Reading order is a design choice with consequences. If the mean is the first thing a panel sees, it becomes the lens through which every subsequent piece of evidence is interpreted — the comments, the response rate, the cohort difficulty. A 3.6 anchors the reader to look for what went wrong; a 4.6 anchors them to explain away the one scathing comment. The evidence is the same; the number changes how it lands. This compounds the related contrast effects and negativity bias already documented in this knowledge base: anchoring sets the baseline, contrast distorts comparison between adjacent cases, and negativity bias over-weights the worst comment.
Anchors are often statistically fragile to begin with. The number doing the anchoring is frequently a small-sample mean with a wide confidence interval — exactly the kind of estimate our guidance on small mean differences and empirical-Bayes shrinkage warns against treating as precise. A 3.8-versus-4.1 gap that is statistically noise can still anchor a promotion decision if it is the first thing on the page.
High-stakes review is where it bites hardest. For formative feedback to a teacher, an anchoring nudge is low-cost. For tenure, promotion, or contract-renewal decisions, an unearned anchor can tilt a career. This is why responsible-reporting guidance (see interpreting and reporting student ratings responsibly) increasingly stresses how evidence is presented, not just what is collected.
Limitations and honest caveats
The direct evidence is analogical, not SET-specific. There is no large randomised trial showing that course-evaluation panels who see the mean first reach different decisions than those who read comments first. The claim rests on transferring a robust, replicated general phenomenon (anchoring in estimation and in judicial sentencing) to the evaluation-committee context. That transfer is reasonable — the cognitive machinery is the same — but it is an inference, and a rigorous reader should treat the recommended mitigations as plausible de-biasing, not proven remedy.
Some anchoring de-biasing attempts fail. The literature on countering anchoring is mixed: simply warning people about the bias, or telling them to "consider the opposite," has inconsistent effects, and the Englich et al. finding that randomness and expertise did not help is sobering. Procedural fixes (changing what information is presented, and in what order) are more promising than exhortation, but no fix fully eliminates the effect.
Order manipulation has its own costs. Withholding the number until last can frustrate reviewers, slow high-volume review, and — if done clumsily — simply move the anchor to whatever appears first instead. The goal is not to hide the number but to ensure it is contextualised before it frames everything, and that its statistical uncertainty is visible when it is read.
Not every anchor is a bias. A prior-year score or a departmental benchmark can be legitimately relevant context. The problem is not that numbers inform judgement; it is that an early, uncertain, decontextualised number exerts influence out of proportion to its evidential weight.
How Koji incorporates this
Koji for Education is designed so that reporting sequences evidence deliberately rather than leading with a decontextualised number.
- Evidence-first, number-in-context reporting. Koji's reporting is designed to present the qualitative picture — themed open-text findings, representative quotes, response-rate and cohort context — alongside the quantitative summary, rather than opening with a bare mean flag. The aim is to prevent a single early number from anchoring interpretation of everything that follows.
- Uncertainty made visible. Because many anchoring numbers are fragile small-sample means, Koji is designed to report scores with their uncertainty (confidence intervals, sample sizes, response rates) so a reviewer sees a range, not a false point, at the moment they read it — consistent with our shrinkage and small-difference guidance.
- Structured, criterion-referenced summaries. Rather than a single global anchor, Koji supports structured multi-dimensional reporting (clarity, workload, feedback, belonging) so no one number dominates the frame, reducing the pull any single anchor can exert.
- Automatic thematic analysis reduces selective reading. By surfacing all major themes and their prevalence — not just the comments that happen to fit the score — Koji's thematic analysis is designed to counteract the selective-accessibility mechanism, giving reviewers a systematic account of the open text instead of an anchor-consistent sample of it.
- Bias-aware guidance for committees. Koji can accompany reports with reporting conventions that encourage reading context and comments before the summary number in high-stakes review.
We frame this precisely: sequencing and uncertainty-visible reporting are designed to mitigate anchoring, not to eliminate a cognitive bias that survives expertise and warnings. Koji's core research platform at koji.so applies the same evidence-first reporting discipline to customer and product research, where a headline metric read first can anchor a product decision just as easily.
A practical protocol for de-anchoring high-stakes review
Because exhortation is a weak defence against anchoring, the fix has to live in the process a committee follows, not in a reminder at the top of the page. A workable protocol has four moves. First, lead with context, not the number: present the response rate, class size, cohort difficulty, and the themed open-text summary before the headline mean, so the reader forms a picture before the anchor arrives. Second, show the number as a range: report the mean with its confidence interval and sample size so a 3.8 built on eleven responses cannot masquerade as a precise verdict — a fragile point estimate is the most dangerous kind of anchor. Third, have reviewers record an initial written impression from the qualitative evidence before the numeric summary is revealed; committing to a reading first makes the later number a check rather than a frame. Fourth, in genuinely high-stakes cases — promotion, tenure, non-renewal — have two reviewers read independently, one context-first and one number-first, and compare: a large divergence is itself a signal that the number was doing too much work.
None of these steps abolishes anchoring; Englich et al. are clear that the effect survives expertise and awareness. But they change what is available to mind first, and since the selective-accessibility mechanism is precisely about what the anchor makes salient, controlling the order of exposure is the most defensible lever a committee actually has. The alternative — opening every file with a red-amber-green flag — hands the anchor maximum power over decisions that deserve deliberation.
Related Resources
- Contrast Effects and Narrow Bracketing in Evaluation Review
- Negativity Bias in Reading Open-Text Comments
- The 4.2 vs 4.4 Trap: Small Differences Are Usually Noise
- Empirical-Bayes Shrinkage for Course-Evaluation Scores
- Interpreting and Reporting Student Ratings Responsibly
- Thin-Slice Judgments and First Impressions
References
- Tversky, A., & Kahneman, D. (1974). Judgment under uncertainty: Heuristics and biases. Science, 185(4157), 1124–1131. https://doi.org/10.1126/science.185.4157.1124
- Englich, B., Mussweiler, T., & Strack, F. (2006). Playing dice with criminal sentences: The influence of irrelevant anchors on experts' judicial decision making. Personality and Social Psychology Bulletin, 32(2), 188–200. https://doi.org/10.1177/0146167205282152
- Mussweiler, T., & Strack, F. (1999). Hypothesis-consistent testing and semantic priming in the anchoring paradigm: A selective accessibility model. Journal of Experimental Social Psychology, 35(2), 136–164. https://doi.org/10.1006/jesp.1998.1364
- Furnham, A., & Boo, H. C. (2011). A literature review of the anchoring effect. The Journal of Socio-Economics, 40(1), 35–42. https://doi.org/10.1016/j.socec.2010.10.008
Related articles
Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees
The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.
The 4.2 vs 4.4 Trap: Why Small Differences in Course-Evaluation Means Are Usually Noise
Faculty and administrators routinely read meaning into tiny gaps between course-evaluation means. Boysen (2015) and Boysen et al. (2014) show this happens even when confidence intervals say the difference is nothing — and that warnings barely help. How to report uncertainty honestly.
Why One Cruel Comment Outweighs Twenty Kind Ones: Negativity Bias in Reading Course Evaluations
Instructors and committees fixate on the harshest open-text comment and discount the praise. Baumeister''s "bad is stronger than good" and the negativity-bias literature explain why, and how to read qualitative course feedback fairly.
Stop Comparing Raw Averages: Empirical-Bayes Shrinkage for Course-Evaluation Scores
Why raw course-evaluation means mislead — especially for small classes — and how empirical-Bayes shrinkage (partial pooling), the method Kane and Staiger applied to teacher value-added, produces fairer, more accurate instructor estimates.