New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
analysis-reporting9 min read

Reading Evaluations to Confirm What You Already Believe: Confirmation Bias in Interpreting Course Feedback

Whoever reads a course evaluation already has a hypothesis about the instructor. Confirmation bias shapes which comments they weight, how ambiguity is resolved, and what the "data" is taken to show. Here is the evidence and the guardrails.

Koji Education Team

Product

In brief

Confirmation bias — the tendency to seek, weight and interpret evidence in ways that favour a belief already held — operates on the readers of course evaluations, not only the students who write them, and it can turn the same set of comments into "confirmation" of whatever a committee expected to find. Nickerson's (1998) authoritative review documents the bias across domains as a robust, largely unintentional feature of human reasoning; Klayman and Ha (1987) show much of it stems from a default "positive test strategy" — checking cases expected to fit the hypothesis. Applied to evaluation reading, this means a chair who expects a colleague to be weak notices the critical comments and discounts the praise, while a champion does the reverse. The fix is procedural, not attitudinal: structure the reading so the hypothesis cannot quietly steer it.

What the research says

Confirmation bias is, in Raymond Nickerson's (1998) Review of General Psychology synthesis, "the seeking or interpreting of evidence in ways that are partial to existing beliefs, expectations, or a hypothesis in hand." His review — one of the most cited treatments of the phenomenon — catalogues its guises: preferential search for confirming information, over-weighting of evidence that fits, under-weighting or explaining-away of evidence that does not, and the interpretation of genuinely ambiguous information as supportive of a prior view. Crucially, Nickerson argues the bias is typically not deliberate; people are not consciously cherry-picking, which is precisely why good intentions do not neutralise it.

Joshua Klayman and Young-Won Ha (1987), in Psychological Review, sharpened the mechanism. Much of what is labelled "confirmation bias," they argued, reflects a general positive test strategy — a tendency to test a hypothesis by examining cases you expect to have the property of interest. In many natural environments this heuristic is efficient, but when a hypothesis is wrong it systematically fails to expose the disconfirming cases that would correct it. The lesson for evaluation is that the bias is not moral weakness; it is the default way people gather evidence, and it needs a countervailing structure.

Two further strands corroborate the applied risk. First, biased assimilation and attitude polarisation: Lord, Ross and Lepper (1979) showed that people presented with mixed evidence on a topic they hold views about rate the confirming studies as more convincing and end up more entrenched — not less — after seeing balanced data. A committee reading a folder of mixed comments can leave more certain of a prior impression than it started. Second, belief perseverance and primacy: initial impressions colour the interpretation of everything that follows, so the first few comments (or the headline mean) function as a hypothesis that later evidence is read to fit — a mechanism related to anchoring and to the fundamental attribution error already familiar in evaluation reading.

The empirical core, then, is well established: interpretation of ambiguous evidence is hypothesis-driven, the pull is largely unconscious, and exposure to mixed evidence can entrench rather than correct.

Why it matters for course evaluation in practice

Course evaluations are an almost ideal substrate for confirmation bias. The data are mixed and ambiguous (most instructors get a spread of glowing, lukewarm and hostile comments), the reader usually arrives with a prior (reputation, a previous score, a personal impression, a rumour), and the stakes — promotion, tenure, module review — make the reader motivated. Every precondition the literature specifies is present.

The consequences are concrete. A promotion panel that expects a candidate to be strong will read "challenging but fair" as evidence of rigour, while for a candidate it doubts the same phrase reads as "students found it hard." A chair investigating a low mean scrolls the open text looking for the problems that would explain it, and — per the positive test strategy — finds them, without ever sampling the counter-evidence that might show the mean was driven by two outliers. Selective quotation compounds this: a report that pulls three verbatim comments to "illustrate" a score almost always illustrates the reader's hypothesis, because those are the comments that came to mind. And because mixed evidence can polarise, a "balanced" folder of feedback does not automatically produce a balanced judgement — it can harden whatever the committee walked in believing.

This is a fairness and validity problem for quality assurance. If interpretation is hypothesis-driven, then two committees can reach opposite conclusions from identical evidence, and the evaluation stops functioning as a measurement and starts functioning as a mirror. It also interacts with documented biases in the underlying ratings (gender, accent, discipline): a reader who already holds a stereotype will find the ambiguous comments that seem to confirm it.

Limitations and honest caveats

Not every disagreement is bias. Priors are often legitimate — a chair who has observed a colleague teach has relevant information, and using it is not a fallacy. The bias is the selective weighting of new evidence to protect the prior, not the possession of one. Overcorrecting — treating every confirming reading as suspect — can itself become a bias and can wrongly discount valid signals.

Evidence transfer. Much of the strongest experimental work (Lord, Ross & Lepper; classic hypothesis-testing paradigms) was not conducted on course-evaluation committees specifically. The mechanism is general and well replicated, but the exact effect sizes in a real promotion panel reading real evaluations are not precisely quantified, and field conditions (accountability, deliberation, written justification) can attenuate the bias. Claims here are therefore about a well-supported risk, not a measured coefficient in this setting.

Debiasing is only partly effective. Simply warning readers to "be objective" does little; the more effective interventions (consider-the-opposite, structured criteria, blinding) reduce but do not eliminate the effect, and some add cost or reduce the contextual judgement that experienced readers legitimately bring. There is no procedure that makes interpretation belief-free; the realistic goal is to make the prior visible and contestable.

How Koji incorporates this

Koji cannot make a reader's mind neutral, but it is designed to change the structure of the evidence so a prior cannot silently steer the reading.

  • Whole-distribution reporting, not curated quotes. Koji's reporting is built to present the full distribution of responses and systematically surfaced themes rather than a human-chosen handful of verbatims, removing the selective-quotation channel through which a hypothesis usually enters.
  • Systematic thematic analysis instead of impressionistic scanning. Koji's automatic thematic analysis clusters all open-text responses and reports theme prevalence, so a reader sees "18% mention pacing, 4% mention the instructor's manner" rather than scrolling until they find what they expected. This directly counters the positive-test-strategy failure of only inspecting confirming cases.
  • Consider-the-opposite by construction. Because Koji reports both supporting and dissenting themes with their frequencies, the disconfirming evidence is placed in front of the reader rather than left to be sought — the single most effective debiasing move the literature identifies, built into the artefact.
  • AI-moderated probing reduces ambiguity at the source. Much confirmation bias exploits ambiguity. Koji's AI moderator follows a vague comment with an open_ended probe that clarifies what the student actually meant, so the reader interprets a specific statement rather than projecting onto a blank. Structured scale, single_choice and yes_no items further reduce the interpretive latitude a prior can exploit.
  • Bias-aware framing. Koji's reports are designed to foreground base rates and distribution shape (see also small-mean-difference and confidence-interval reporting), discouraging the leap from an anecdote to a verdict. Koji is designed to mitigate confirmation bias in interpretation, not to eliminate it.

Koji's core research platform at koji.so applies the same theme-first, whole-distribution reporting to product and customer research, where stakeholders reading user feedback face the identical temptation to find what they came to confirm.

Related Resources

A reading protocol that constrains the prior

Because the bias is structural rather than a matter of goodwill, the countermeasures that work are procedural. Committees that must reach defensible judgements from evaluation evidence can borrow three moves the literature supports. First, fix the criteria before opening the folder: agreeing in advance what would count as strong or weak teaching prevents a reader from quietly redefining the standard to fit the impression the first few comments create. Second, read the distribution before the anecdotes: theme frequencies and the full spread of ratings anchor the reader in what the cohort actually said, so individual quotes are interpreted against a base rate rather than standing in for it. Third, require a written counter-case: asking each reader to articulate the strongest reading against their initial impression operationalises the consider-the-opposite instruction, one of the few debiasing techniques with consistent empirical support. None of these makes a reader neutral, but each removes a degree of freedom through which an unexamined prior turns mixed evidence into false confirmation.

References

  • Nickerson, R. S. (1998). Confirmation bias: A ubiquitous phenomenon in many guises. Review of General Psychology, 2(2), 175–220. https://doi.org/10.1037/1089-2680.2.2.175
  • Klayman, J., & Ha, Y.-W. (1987). Confirmation, disconfirmation, and information in hypothesis testing. Psychological Review, 94(2), 211–228. https://doi.org/10.1037/0033-295X.94.2.211
  • Lord, C. G., Ross, L., & Lepper, M. R. (1979). Biased assimilation and attitude polarization: The effects of prior theories on subsequently considered evidence. Journal of Personality and Social Psychology, 37(11), 2098–2109. https://doi.org/10.1037/0022-3514.37.11.2098
  • Ross, L., Lepper, M. R., & Hubbard, M. (1975). Perseverance in self-perception and social perception: Biased attributional processes in the debriefing paradigm. Journal of Personality and Social Psychology, 32(5), 880–892. https://doi.org/10.1037/0022-3514.32.5.880