New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
analysis-reporting9 min read

The Availability Heuristic: Why a Few Vivid Comments Distort How You Read Course Evaluations

The availability heuristic means the most memorable, extreme open-text comment feels more frequent than it is. What the research says, why it corrupts qualitative course-evaluation review, and how to read comments by prevalence instead of vividness.

Koji Education Team

Product

The short answer

When a committee reads the open-text section of a course evaluation, the comment that comes to mind most easily — the vivid horror story, the furious rant, the single unforgettable phrase — feels far more representative than it actually is. This is the availability heuristic: people judge how frequent or likely something is by how easily examples come to mind (Tversky & Kahneman, 1973). In practice, three angry comments out of sixty can dominate a programme review while fifty-five constructive ones are forgotten. The fix is not to read harder; it is to read by counted prevalence rather than by recalled vividness — tabulate how often each theme actually appears before you form a judgement.

BLUF: The availability heuristic makes memorable, emotionally charged, or extreme open-text comments feel more common than they are, biasing qualitative review toward the loudest voices. Because vividness and ease-of-recall — not frequency — drive the estimate, the antidote is to quantify theme prevalence (what share of respondents actually raised an issue) before interpreting it. Koji is designed to mitigate this by auto-coding every open-text response into themes and reporting each theme''s prevalence and representative-but-weighted quotes, so a review reads the distribution, not just the outliers.

What the research says

The availability heuristic was introduced by Amos Tversky and Daniel Kahneman in Cognitive Psychology (1973). Their thesis: when asked how frequent or probable a class of events is, people substitute an easier question — how readily do instances come to mind? — and answer that instead. The substitution is often reasonable, because frequent things are usually easier to recall. But it fails systematically whenever ease of recall is driven by something other than frequency: recency, emotional salience, vividness, or distinctiveness. Their classic demonstration asked whether English words are more likely to start with the letter K or to have K as the third letter; most people say the former (words starting with K are easier to generate), when in fact K appears more often in the third position.

Norbert Schwarz and colleagues sharpened the mechanism in a landmark Journal of Personality and Social Psychology paper (Schwarz et al., 1991). They showed the driver is the metacognitive experience of ease, not the content recalled. Participants asked to recall six examples of their own assertive behaviour — an easy task — rated themselves as more assertive than participants asked to recall twelve examples, even though the second group listed more evidence of assertiveness. Struggling to produce the last few examples felt like disconfirmation. The implication for reading feedback is direct: a reviewer who can effortlessly summon two scathing comments will over-estimate how negative the cohort was, precisely because those two came so easily.

The heuristic also distorts frequency estimates for whole categories. Lichtenstein, Slovic, Fischhoff, Layman and Combs (1978), in the Journal of Experimental Psychology: Human Learning and Memory, found that people over-estimate the frequency of dramatic, well-publicised causes of death (accidents, homicide, tornadoes) and under-estimate quiet, common ones (diabetes, stroke) — a bias that tracks media coverage and memorability rather than actual statistics. Translated to a course-evaluation dataset, a single graphic complaint about a lab safety scare will loom larger in the reader''s frequency estimate than a diffuse, widely-shared grumble about assessment timing, even if ten times more students mentioned the latter.

Taken together, three findings matter for evaluation practice: (1) frequency judgements are anchored to recall ease, not counts; (2) ease is manufactured by vividness, extremity and recency; and (3) the bias survives even when the full evidence is in front of you, because you weight the feeling of retrieval over the tally.

Why it matters for course evaluation in practice

Open-text comments are the richest part of most course evaluations, and also the most vulnerable to this bias, because they are almost never counted — they are read. A programme director skimming 200 comments does not compute a distribution; they form an impression, and that impression is disproportionately built from the handful of comments that were easiest to remember: the funniest, the cruellest, the most specific, the last one they read (a recency effect that compounds availability).

The practical damage is threefold. First, misallocated action. If the vivid-but-rare complaint captures the committee''s attention, the closing-the-loop action plan targets an issue three students raised while ignoring the workload problem forty students raised in calmer language. Second, instructor harm. A single abusive or personal comment is maximally available and can colour an entire personnel discussion, which is one reason abusive comments are both a well-being issue and a validity issue. Third, false consensus. Reviewers who each remember the same one or two striking comments will feel they independently converged on a shared conclusion, mistaking correlated recall for corroborated evidence.

The bias is worse under exactly the conditions QA panels operate in: time pressure, large comment volumes, and no requirement to quantify. It is also asymmetric — negative and extreme comments are more vivid than lukewarm-positive ones, so availability and negativity bias reinforce each other.

Limitations and honest caveats

A PhD reader should hold several caveats. First, the availability heuristic is a heuristic, not a guarantee of error: when recall ease genuinely tracks frequency, it produces fast and accurate judgements. The problem is specifically the decoupling of ease from frequency, which open-text feedback invites but does not force. Second, much of the foundational evidence comes from laboratory tasks (letter frequencies, autobiographical recall) rather than from committees reading real evaluations; the generalisation to evaluation review is theoretically strong but relies on translation, and field replications specific to QA panels are scarce. Third, the ease-of-retrieval effect itself has boundary conditions — later work shows it can reverse or disappear when people attribute the difficulty of recall to an external cause, or when they are highly motivated and knowledgeable about the topic. Fourth, quantifying prevalence is not a neutral cure-all: a rare comment can be diagnostically critical (a credible safety or misconduct allegation must not be down-weighted just because it is infrequent). Prevalence should inform interpretation, not mechanically override judgement. Counting themes also imposes a coding frame, which introduces its own inter-rater-reliability questions. The honest position is that availability is one well-documented distortion among several, and that quantification manages it rather than eliminating it.

How Koji incorporates this

Koji is built to let reviewers read the distribution of feedback, not just the most available fragments of it, while preserving the qualitative richness that makes open text worth collecting.

  • Automatic thematic coding with prevalence counts. Every open-ended response is coded into recurring themes, and each theme is reported with the share of respondents who raised it. A reviewer sees "assessment timing — raised by 38% of respondents" next to "lab equipment — raised by 5%," which directly substitutes a count for a recall-ease impression. This is the core mitigation for availability: it re-couples the frequency estimate to an actual tally.
  • Representative, not cherry-picked, quotes. For each theme Koji surfaces quotations chosen to illustrate the range of a theme, not the single most extreme example, so vividness stops standing in for typicality. The striking comment is still visible — but it is visible as one instance of a 5% theme, which is the context availability strips away.
  • Aspect-based prevalence and sentiment. Because themes are tagged by aspect (content, workload, assessment, teaching, environment) and by sentiment, a panel can see that a vivid negative comment sits inside an otherwise positive aspect, rather than letting one memorable line define the whole picture.
  • AI-moderated conversational interviews that probe beyond a first vivid reaction. Where a static box captures whoever felt most strongly, Koji''s conversational format follows up ("how often did that happen?", "did it affect your learning?"), which yields calibrated, comparable statements instead of a few unforgettable outbursts — reducing the raw material of the availability bias at the source.
  • Bias-aware reporting. Reports foreground prevalence and flag when a conclusion rests on a small number of comments, nudging reviewers to ask whether a striking finding is frequent or merely available.

Framed honestly, none of this removes human judgement or the diagnostic weight of a serious rare comment — it is designed to mitigate the availability distortion by making counts as easy to see as the memorable lines. Koji''s core research platform at koji.so applies the same theme-prevalence engine to product and customer research, where a single loud user interview can hijack a roadmap in exactly the same way.

Frequently asked questions

What is the availability heuristic in the context of course evaluations? It is the tendency to judge how common a piece of feedback is by how easily an example comes to mind. Because vivid, extreme, or recent comments are easier to recall, reviewers over-estimate how frequent those views are and under-weight quieter, more common feedback — even when all the comments are in front of them.

How is the availability heuristic different from negativity bias or base-rate neglect? Negativity bias is about weighting negative information more heavily than positive; base-rate neglect is about ignoring prior probabilities. Availability is about frequency estimation driven by recall ease — you think something happened a lot because a striking instance is easy to bring to mind. They often co-occur (negative comments are more vivid, hence more available), but the mechanisms are distinct.

Does counting comment themes solve the problem completely? No. Quantifying prevalence re-anchors frequency judgements to actual counts, which addresses the core distortion, but a rare comment can still be critically important (for example a credible safety allegation). Prevalence should inform interpretation, not mechanically override it, and thematic coding introduces its own reliability considerations.

Why are open-text comments especially exposed to this bias? Because they are read, not tabulated. Numeric ratings force an implicit average; comments do not. Without a count, the reader defaults to impression-formation, which weights the most available fragments — the funniest, cruellest, or last-read lines.

What can a QA committee do about it without special software? Tally themes before discussing them, agree a simple coding frame in advance, read comments in a randomised order to blunt recency, and explicitly ask "what share of respondents said this?" for any claim that shapes an action plan. Separating a striking comment''s vividness from its frequency is the whole game.

Related resources

References

  • Tversky, A., & Kahneman, D. (1973). Availability: A heuristic for judging frequency and probability. Cognitive Psychology, 5(2), 207–232. https://doi.org/10.1016/0010-0285(73)90033-9
  • Schwarz, N., Bless, H., Strack, F., Klumpp, G., Rittenauer-Schatka, H., & Simons, A. (1991). Ease of retrieval as information: Another look at the availability heuristic. Journal of Personality and Social Psychology, 61(2), 195–202. https://doi.org/10.1037/0022-3514.61.2.195
  • Lichtenstein, S., Slovic, P., Fischhoff, B., Layman, M., & Combs, B. (1978). Judged frequency of lethal events. Journal of Experimental Psychology: Human Learning and Memory, 4(6), 551–578. https://doi.org/10.1037/0278-7393.4.6.551