Denominator Neglect: Why Raw Comment Counts and Percentages Mislead Course-Evaluation Readers
Denominator neglect and ratio bias make "five complaints" and "20% negative" feel worse than they are. Here is what the research says about this reasoning error and how to report course evaluations so it does not distort decisions.
Koji Education Team
Product
In brief
When a dean reads "five students complained about the workload," the natural reaction is alarm — but five out of how many? Reading a numerator without weighting it by its denominator is a well-documented reasoning error called denominator neglect, closely tied to ratio bias (Reyna & Brainerd, 2008). It systematically distorts how course evaluations are interpreted: a handful of vivid negative comments from a low-response survey can outweigh a large silent majority, and a "20% negative" figure feels worse than "2 in 10" even though they are identical. The fix is not more data but better presentation — always pairing counts with their base, showing response rates prominently, and using visual formats that make the denominator impossible to ignore.
Answer box. Denominator neglect is the tendency to focus on the numerator of a ratio (how many times something happened) while under-weighting the denominator (out of how many opportunities). In course evaluation it makes raw complaint counts and bare percentages mislead readers: "five negative comments" or "18% dissatisfied" trigger stronger reactions than the underlying proportion warrants, especially when response rates are low. Reporting that always shows counts over their base, foregrounds response rates, and uses icon arrays or part-of-whole visuals measurably reduces the error.
What the research says
The anchor study is Reyna and Brainerd's (2008) "Numeracy, ratio bias, and denominator neglect in judgments of risk and probability," in Learning and Individual Differences. Working within fuzzy-trace theory, they define ratio bias as the tendency to judge an event more likely when it is expressed as a large-numbered ratio (e.g., 10 in 100) than as an equivalent small-numbered ratio (1 in 10), and denominator neglect as the default tendency to attend to the numerator — the count of target events — while paying too little attention to the denominator that gives that count its meaning. Their review synthesises evidence that these errors are widespread, are moderated by numeracy (people lower in numeracy are more susceptible), and arise because reasoners extract a "gist" dominated by the salient numerator rather than the full ratio.
Two corroborating studies show the error is correctable through presentation, which is the actionable finding for reporting. Garcia-Retamero, Galesic and Gigerenzer (2010), in "Do icon arrays help reduce denominator neglect?" (Medical Decision Making), found that adding icon arrays — grids of symbols where the filled portion represents the numerator against the whole denominator — to numerical risk information significantly reduced denominator neglect. The effect was largest for lower-numeracy participants: in one condition the proportion of low-numeracy people who misjudged a comparison fell from 74% to 42% when an icon array accompanied the numbers. Galesic and Garcia-Retamero (2011), in "Graph literacy" (Medical Decision Making), further showed that people vary in their ability to read such visuals, so format choices must match the audience's graph literacy rather than assuming a chart automatically helps.
The convergent picture across these sources: the number itself is not the problem; the format that hides or reveals the denominator determines whether readers reason about a proportion or react to a count.
Why it matters for course evaluation in practice
Course-evaluation reports are dense with ratios whose denominators are easy to drop, and the people reading them — programme directors, promotion committees, quality officers — are subject to exactly the error Reyna and Brainerd describe. Four concrete failure modes recur.
-
Complaint counts without the base. "Seven students said the pace was too fast" reads as a serious signal until you learn the class had 210 students and 160 respondents. The absolute count is salient; the denominator is invisible unless the report forces it into view. Comment counts should never be presented without the number of respondents who addressed that theme and the number who could have.
-
Percentages that hide small denominators. In a seminar of 12 students with 6 respondents, "33% rated the course poorly" means two people. The percentage inherits false authority precisely because it strips the denominator. This is the same problem behind ratio bias: a proportion detached from its base invites over-reaction. Low-N percentages should be shown as fractions (2 of 6) or suppressed entirely below a threshold.
-
Cross-course comparison with unequal response. Comparing a "15% negative" course at 40% response against a "10% negative" course at 90% response treats two very differently-grounded numbers as commensurable. Denominator neglect makes the raw percentages feel comparable when the sampling behind them is not, compounding the non-response and small-sample cautions discussed elsewhere in this knowledge base.
-
The vivid-comment override. A single articulate, angry open-text comment can dominate a committee's impression of a course rated positively by the silent majority — a fusion of denominator neglect with negativity bias. Without a visible reminder that this is one voice among many, the numerator wins.
A worked example makes the distortion concrete. A first-year lecture of 240 students receives 96 responses (a 40% response rate); of those, 12 mention that assessment deadlines clustered awkwardly. Reported as "12 complaints about deadlines," it reads as a crisis worth an urgent redesign. Reported as "12 of 96 respondents, or 5% of those who answered — and 12 of 240 enrolled, or 2%," it reads as a minor, possibly worth-noting theme. Nothing in the data changed; only the visibility of the denominator did, and with it the decision it invites. Multiply that across every theme in a multi-section programme report and the aggregate risk of denominator-driven over-reaction becomes substantial.
Because these are reader errors, not data errors, they are fixed at the reporting layer: show counts over their base, put response rate next to every percentage, express low-N results as fractions, and use part-of-whole visuals so the denominator is structurally present.
Limitations and honest caveats
-
Denominator neglect is not universal or fixed. Reyna and Brainerd emphasise that susceptibility varies with numeracy and with how information is framed. Highly numerate readers, and reports that already present rates well, may show little of the effect. Claiming that every reader will misjudge every count over-states the evidence.
-
The primary evidence is from risk and health, not evaluation. The anchor and corroborating studies concern medical risk and probability judgments. The transfer to reading course evaluations is well-motivated — the cognitive mechanism is domain-general — but it is an inference, not a finding from evaluation research itself. Treat it as a strong theoretical prior, not proof.
-
Visuals are not a panacea. Galesic and Garcia-Retamero's graph-literacy work shows that icon arrays and part-of-whole charts help only readers who can decode them; poorly designed or unfamiliar graphics can add confusion. Format choices should be tested with the actual audience, and text and visual should reinforce, not replace, each other.
-
Suppressing small-N results has its own cost. Hiding low-response percentages protects against ratio bias but can also conceal genuine signals from small classes. The remedy is transparent thresholds and fraction-based display, not silent deletion — a reporting-governance decision the institution should make explicitly.
How Koji incorporates this
Koji for Education is designed so that the denominator travels with every number, addressing the reader-side error at the point of reporting rather than hoping readers self-correct.
-
Base-anchored reporting by default. Koji's analysis and reporting layer pairs every theme count and every percentage with its denominator — respondents who addressed the theme, total respondents, and the response rate — so a figure like "18% raised pacing concerns" is never shown stripped of the sample it came from. This is a bias-aware reporting choice aimed squarely at denominator neglect.
-
Part-of-whole and fraction display for small cohorts. For low-response or small-class results, Koji can present findings as fractions (2 of 6) and as part-of-whole visuals rather than bare percentages, echoing the icon-array evidence that making the denominator structurally visible reduces the error. Reporting thresholds are configurable so institutions can set transparent rules instead of silently dropping data.
-
Thematic counts contextualised, not amplified. Koji's automatic thematic analysis of open-text reports how many respondents expressed a theme against how many responded, so a vivid single comment is shown as one voice within a quantified whole — mitigating the fusion of denominator neglect and negativity bias that lets one articulate complaint override a satisfied majority.
-
AI-moderated depth over raw counts. Because Koji's AI-moderated conversational interviews probe the reasoning behind a response, a low count of a concern can be examined for severity rather than merely tallied, helping decision-makers weigh a small-but-serious signal appropriately instead of reacting to, or dismissing, a number in isolation.
Koji frames these as designed to mitigate a documented reasoning error, not to eliminate human judgement. Koji's core research platform at koji.so applies the same reporting discipline to product and customer research, where "how many users complained" is meaningless without "out of how many," and the same denominator-neglect trap awaits every stakeholder readout.
Frequently asked questions
What is denominator neglect?
Denominator neglect is the tendency to focus on the numerator of a ratio — how many times something happened — while under-weighting the denominator, the number of opportunities it could have happened. Reyna and Brainerd (2008) show it is widespread, stronger in people lower in numeracy, and arises because reasoners extract a gist dominated by the salient count.
How is denominator neglect different from base-rate neglect?
Base-rate neglect is ignoring the prior probability or general prevalence when judging a specific case. Denominator neglect is narrower: it is dropping the base of a specific ratio you are looking at — reacting to "five complaints" or "20% negative" without weighting the sample size behind them. The two often co-occur in evaluation reading but are distinct errors.
Why do bare percentages mislead in small classes?
A percentage strips its denominator, so "33% rated the course poorly" in a six-respondent seminar means just two people while inheriting the authority of a proportion. For small cohorts, expressing results as fractions (2 of 6) or suppressing percentages below a transparent threshold prevents ratio bias from over-amplifying tiny counts.
Do charts fix denominator neglect?
Partly. Garcia-Retamero and colleagues (2010) found icon arrays that show the numerator against the whole denominator significantly reduced the error, especially for lower-numeracy readers. But Galesic and Garcia-Retamero (2011) showed graph literacy varies, so visuals help only readers who can decode them and should be tested with the actual audience and paired with text.
Does this research come from course-evaluation studies?
No. The anchor and corroborating evidence come from risk and probability judgments in health and decision-making contexts. The cognitive mechanism is domain-general, so applying it to reading course evaluations is a well-motivated inference and a strong prior, but not a finding from evaluation research itself — a caveat worth stating in any report that relies on it.
How does Koji stop denominator neglect in its reports?
Koji pairs every theme count and percentage with its denominator and the response rate by default, displays small-cohort results as fractions and part-of-whole visuals with configurable thresholds, contextualises thematic counts so a single vivid comment is shown as one voice within a quantified whole, and uses AI-moderated probing to weigh a small-but-serious concern by severity rather than by raw tally.
Related resources
- Base-rate neglect in interpreting course evaluations
- Negativity bias when reading open-text comments
- Top-box vs mean: how you report changes the story
- Small mean differences and confidence intervals
- Interpreting and reporting student ratings responsibly
- Is ranking instructors by evaluation scores fair?
References
- Reyna, V. F., & Brainerd, C. J. (2008). Numeracy, ratio bias, and denominator neglect in judgments of risk and probability. Learning and Individual Differences, 18(1), 89–107. https://doi.org/10.1016/j.lindif.2007.03.011
- Garcia-Retamero, R., Galesic, M., & Gigerenzer, G. (2010). Do icon arrays help reduce denominator neglect? Medical Decision Making, 30(6), 672–684. https://doi.org/10.1177/0272989X10369000
- Galesic, M., & Garcia-Retamero, R. (2011). Graph literacy: A cross-cultural comparison. Medical Decision Making, 31(3), 444–457. https://doi.org/10.1177/0272989X10373805
Related articles
Can You Fairly Rank Instructors by Their Course-Evaluation Scores? What Esarey & Valdes (2020) Proved
Even if course evaluations were unbiased, reliable, and valid, ranking instructors by their scores would still misclassify many good teachers. A walk through the Esarey & Valdes (2020) simulation and what it means for how you report and use evaluation data.
Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees
The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.
The 4.2 vs 4.4 Trap: Why Small Differences in Course-Evaluation Means Are Usually Noise
Faculty and administrators routinely read meaning into tiny gaps between course-evaluation means. Boysen (2015) and Boysen et al. (2014) show this happens even when confidence intervals say the difference is nothing — and that warnings barely help. How to report uncertainty honestly.
Why One Cruel Comment Outweighs Twenty Kind Ones: Negativity Bias in Reading Course Evaluations
Instructors and committees fixate on the harshest open-text comment and discount the praise. Baumeister''s "bad is stronger than good" and the negativity-bias literature explain why, and how to read qualitative course feedback fairly.