One Vivid Comment Is Not a Pattern: Base-Rate Neglect in Reading Course Evaluations
A single scathing open-text comment can outweigh forty neutral ones in a reviewer's mind. Bar-Hillel and Kahneman & Tversky showed why: people ignore base rates in favour of vivid, individuating detail. Here is how base-rate neglect distorts course-evaluation review — and how to design against it.
Koji Education Team
Product
In brief
When a committee reads course evaluations, one dramatic comment — a furious paragraph, a shocking anecdote — routinely outweighs the far larger number of neutral or positive responses it sits among. This is base-rate neglect (the base-rate fallacy): the well-documented human tendency to under-weight background frequency information in favour of vivid, specific, individuating detail. The research says the fix is not "try harder to be objective" — it is to make the base rate impossible to ignore: report how many responses said each thing, present the denominator alongside every quote, and force the reader to place the vivid case against the distribution before drawing a conclusion.
What the research says
The foundational demonstration is Kahneman & Tversky (1973), On the Psychology of Prediction (Psychological Review, 80, 237–251). Given a personality sketch and told it was drawn from a pool of 70 engineers and 30 lawyers (or the reverse), participants judged the person's profession almost entirely from how well the description matched their stereotype — and essentially ignored the 70/30 base rate. The individuating story dominated the frequency.
Maya Bar-Hillel (1980), The Base-Rate Fallacy in Probability Judgments (Acta Psychologica, 44, 211–233), sharpened the account. Her key contribution is the relevance explanation: people do not ignore base rates because they cannot do arithmetic. They order information by perceived relevance and let high-relevance information dominate low-relevance information. Specific, individuating detail feels more relevant to the case at hand than an abstract background frequency — so it wins, even when it is statistically less diagnostic. Bar-Hillel showed the neglect weakens when the base rate is made to feel causally or specifically relevant to the case.
A vivid clinical corroboration is Casscells, Schoenberger & Graboys (1978) (New England Journal of Medicine): given a disease with a 1-in-1,000 prevalence and a test with a 5% false-positive rate, most Harvard Medical School staff and students estimated the chance that a positive-testing person actually had the disease at about 95%. The correct answer, dominated by the low base rate, is roughly 2%. Highly trained professionals neglected the base rate wholesale.
The mechanism travels directly into evaluation review, where it compounds with the availability heuristic (Tversky & Kahneman, 1973): an emotionally charged comment is easier to recall and mentally simulate, so it feels more frequent and more representative of the cohort than it is.
Why it matters for course evaluation in practice
Course-evaluation data is exactly the environment base-rate neglect exploits — a small number of vivid free-text comments embedded in a large field of ratings.
- The loud minority rewrites the story. In a cohort of 60 where 55 responses are neutral-to-positive and 3 are furious, the committee's memory of the course is disproportionately built from the 3. The base rate (55/60 not complaining) is abstract; the furious paragraph is concrete and emotionally sticky.
- Rare events look like trends. A single report of a specific problem can be read as "students are saying X" when the base rate is one mention in dozens. Without the denominator, a reviewer cannot tell a genuine pattern from a one-off.
- It penalises honest, detailed feedback. The richest, most specific comments are the most vivid — and therefore the most over-weighted. Instructors quickly learn that a course can be dragged down by a handful of articulate complaints regardless of the distribution, which corrodes trust in the whole exercise.
- It interacts with who responds. Because dissatisfied students are often more motivated to write long comments (a non-response and self-selection issue covered in our companion pieces), the vivid tail is already over-represented before base-rate neglect amplifies it a second time.
- Committees are not immune. As Casscells et al. showed, expertise does not protect against it. Seniority and good intentions are not a control; structure is.
Limitations and honest caveats
A methodologically careful reader should note the boundaries of this claim:
- Base-rate neglect is not universal. Decades of follow-up work (Gigerenzer, Koehler, and others) show the effect shrinks substantially when information is presented as natural frequencies ("3 out of 60 students") rather than probabilities or percentages, and when the base rate is framed as causally relevant. The fallacy is a design-dependent tendency, not an immutable law — which is precisely why reporting format is the lever.
- A vivid comment can be a true signal. Rarity is not falsity. A single credible report of misconduct, safety risk, or discrimination must be escalated on its content, not discounted because it is one in sixty. Correcting for base-rate neglect means weighting frequency claims correctly — not ignoring low-frequency events that matter on their own terms.
- The original studies are lab tasks. Kahneman & Tversky and Bar-Hillel used stylised judgment problems, not committee meetings. The extrapolation to evaluation review is a well-motivated analogy supported by the availability literature, but it has not been tested as a controlled experiment on evaluation committees specifically.
- Debiasing effects can be modest and fade. Simply warning reviewers about the bias has weak, short-lived effects; changing the information they see is far more reliable than exhorting them to think harder.
How Koji incorporates this
Koji is an AI-native course-evaluation platform, and countering base-rate neglect is fundamentally a reporting-design problem — one Koji is built to address:
- Every theme carries its denominator. Koji's automatic thematic analysis of open-text responses reports each theme as a natural frequency — "mentioned by 4 of 58 respondents" — rather than surfacing quotes in isolation. Putting the base rate next to the quote is the single most evidence-backed debiasing move, and it is the default, not an option a reader has to request.
- Quotes are shown as instances of a distribution. When Koji surfaces an illustrative comment, it presents it explicitly as one example within a sized theme, so the reviewer sees the vivid case and the base rate in the same glance — the framing Bar-Hillel showed reduces the fallacy.
- Salience-flattening layout. Bias-aware reporting is designed so that a long, emotive comment does not occupy more visual weight than its frequency warrants; the distribution comes first, the exemplar second.
- Structured questions anchor the base rate. Pairing open text with
scaleandsingle_choiceitems gives a quantitative backbone, so a dramatic comment can be checked against the numeric distribution rather than standing in for it. - Genuine low-frequency risks are routed, not buried. Because base-rate correction must never suppress a real one-off safety or misconduct signal, Koji can flag sensitive individual responses for escalation on content, keeping the frequency-weighting fix separate from the duty to act on rare but serious reports.
Koji's core research platform at koji.so applies the same frequency-anchored thematic reporting to customer and product research, where a single loud user is just as capable of hijacking a roadmap as a single loud student is of hijacking a course review.
A debiasing checklist for evaluation committees
Because warning reviewers about the bias has weak and short-lived effects, the practical response is procedural — build the base rate into the room rather than into people's willpower:
- Read the numbers before the comments. Present the numeric distribution (means, spread, response rate) first, then the open text, so the quantitative base rate is established before any vivid quote can anchor the discussion.
- Never show a quote without its denominator. Every illustrative comment should be labelled with how many respondents raised the same theme out of how many responded — "raised by 4 of 58" — turning an anecdote into a rate.
- Count before you conclude. Ask explicitly: "How many students actually said this?" Requiring a count converts a felt impression into a checkable frequency and is the single most effective in-meeting habit.
- Use natural frequencies, not percentages. "3 out of 60" is processed more accurately than "5%" (Gigerenzer & Hoffrage, 1995), and far better than a bare quote.
- Separate the escalation lane. Decide in advance that credible reports of harm, safety, or misconduct are escalated on content regardless of frequency, so that frequency-weighting the ordinary feedback never becomes an excuse to dismiss a serious rare signal.
- Read across cohorts. A single term is itself a small sample; a theme's frequency this term should be read against its base rate across previous cohorts before it is treated as new.
A committee that follows these steps is not relying on individual reviewers to out-think a robust cognitive bias — it is changing the information environment, which the evidence says is what actually works.
Related resources
- Anchoring bias: does the number anchor the committee?
- The dilution effect: how extra data weakens a strong signal
- Confirmation bias in interpreting course feedback
- Contrast effects and narrow bracketing in review
- Response latency: spotting low-quality responses
- Simpson's paradox in aggregating scores
References
- Kahneman, D., & Tversky, A. (1973). On the psychology of prediction. Psychological Review, 80(4), 237–251. https://doi.org/10.1037/h0034747
- Bar-Hillel, M. (1980). The base-rate fallacy in probability judgments. Acta Psychologica, 44(3), 211–233. https://doi.org/10.1016/0001-6918(80)90046-3
- Casscells, W., Schoenberger, A., & Graboys, T. B. (1978). Interpretation by physicians of clinical laboratory results. New England Journal of Medicine, 299(18), 999–1001. https://doi.org/10.1056/NEJM197811022991808
- Tversky, A., & Kahneman, A. (1973). Availability: A heuristic for judging frequency and probability. Cognitive Psychology, 5(2), 207–232. https://doi.org/10.1016/0010-0285(73)90033-9
- Gigerenzer, G., & Hoffrage, U. (1995). How to improve Bayesian reasoning without instruction: Frequency formats. Psychological Review, 102(4), 684–704. https://doi.org/10.1037/0033-295X.102.4.684
Related articles
Some of Your Course Evaluations Were Answered in Ten Seconds: Response Latency as a Data-Quality Signal
Response times ('paradata') are a free, objective quality signal in online course evaluations. Evidence from Zhang & Conrad and others shows speeders straightline more — and how to use latency to screen data without deleting honest fast responses.
When Combining Sections Reverses the Result: Simpson's Paradox in Course-Evaluation Data
Aggregating course-evaluation scores across sections, cohorts, or years can reverse the very conclusion you are trying to draw. What Simpson's paradox is, how it appears in evaluation data, and how to report so the reversal cannot bite you.
The Course Reviewed Straight After a Brilliant One Looks Worse: Contrast Effects and Narrow Bracketing in Evaluation Review
Course-evaluation bias research focuses on the student filling in the form. The evidence on sequential judgement says the committee reading twenty reports in an afternoon is biased too — by what it read immediately before.
Does the Number Anchor the Committee Before It Reads a Word? Anchoring Bias in Evaluation Review
When a review panel sees an instructor's 3.8 mean before reading the comments, that number quietly pulls every later judgement toward it. What the anchoring literature says, and how to sequence evaluation review to resist it.