How Padding a Report With Extra Data Weakens a Strong Signal: The Dilution Effect in Evaluation Review
A classic finding (Nisbett, Zukier & Lemley, 1981) shows that adding irrelevant, non-diagnostic information makes judgements less extreme. For evaluation committees, a padded report can dilute a genuinely strong signal.
Koji Education Team
Product
In brief
The dilution effect (Nisbett, Zukier, & Lemley, 1981) is a robust judgement bias: adding non-diagnostic (irrelevant) information to genuinely diagnostic (predictive) information makes people''s judgements less extreme — the weak signal dilutes the strong one. For anyone who reads course-evaluation reports — heads of department, promotion panels, accreditation reviewers — this is a direct warning. A report padded with irrelevant charts, demographic breakdowns, and boilerplate can water down a clear finding, and, worse, accountability makes it worse, not better. The remedy is disciplined report design that foregrounds diagnostic evidence and strips non-diagnostic clutter.
What the research says
In their classic Cognitive Psychology paper, Nisbett, Zukier, and Lemley (1981) asked participants to predict outcomes about target individuals. One group received only diagnostic information — details genuinely predictive of the outcome. Another group received the same diagnostic information plus additional non-diagnostic details — facts with no predictive value. The result was counter-intuitive: participants given the extra, useless information made markedly less extreme predictions than those given the diagnostic facts alone. Irrelevant information did not just get ignored; it actively pulled judgements toward the middle.
Nisbett et al. explained this through similarity-based reasoning: people judge by comparing a target to a prototype, and irrelevant details make the target seem less prototypical, diluting the perceived strength of the diagnostic cue. The effect has since replicated across consumer judgement, legal decision-making, fraud risk, and negotiation.
Two follow-ups sharpen the lesson for evaluation committees:
- Tetlock and Boettger (1989), in the Journal of Personality and Social Psychology, found that accountability magnifies the dilution effect. Participants who expected to justify their judgements to others diluted more in response to non-diagnostic information, not less. This is the alarming part: the very people whose judgements carry formal consequences — panels who must document their reasoning — are more susceptible, not less.
- Peters and Rothbart (2000) showed the effect is moderated by typicality: non-diagnostic information dilutes when it makes the target seem less typical of the category, and can even reverse under some conditions. In other words, how irrelevant detail is framed changes how much it dilutes.
The convergent message: mixing low-value information with high-value information systematically degrades judgement, and formal accountability — the hallmark of QA and personnel review — amplifies the damage.
Why it matters for course evaluation in practice
Course-evaluation reports are frequently the opposite of lean. In an effort to look thorough, institutions produce dashboards crammed with every item mean, demographic cut, benchmark, and rating distribution — much of it non-diagnostic for the decision at hand. The dilution effect predicts a concrete harm: a genuinely strong signal gets watered down by the surrounding noise.
Consider a lecturer with an unambiguous, well-evidenced problem — consistent, specific student comments about inaccessible materials, corroborated across cohorts. Presented alone, that is a clear, actionable diagnostic signal. Buried inside a 20-page report alongside irrelevant breakdowns (ratings by time-of-day, decorative charts, items unrelated to the concern), the committee''s judgement drifts toward "mixed picture, hard to say." The signal did not weaken; the presentation diluted it.
The effect cuts both ways, and both are damaging:
- A strong positive signal (clear evidence of excellent teaching) can be diluted into "solid but unremarkable," disadvantaging staff in promotion cases.
- A strong negative signal (a real quality problem) can be diluted into "nothing decisive," so nothing gets fixed — undermining the entire point of quality assurance.
And because promotion panels, accreditation reviewers, and QA committees are formally accountable, Tetlock and Boettger''s finding says they are especially exposed. This compounds related reading-stage biases already in the literature — anchoring on the numeric score before reading comments, and contrast effects between consecutively reviewed cases. Dilution is the third member of that family: it is about what surrounds the signal in a single report.
The practical corrective is report design as a bias-control instrument, not just a data dump:
- Lead with the diagnostic evidence. Put the decision-relevant signal first and prominently; relegate context to appendices.
- Cut non-diagnostic clutter. If a chart does not bear on the question the reader must answer, it is not neutral — it is dilutive. This reinforces the case for reporting responsibly and for choosing summary statistics that carry signal, such as top-box versus mean reporting.
- Separate diagnostic from descriptive. Label what is decision-relevant versus merely descriptive, so readers weight accordingly.
Limitations and honest caveats
- External validity. The classic dilution studies used artificial prediction tasks with student participants, not committees reading real evaluation reports. The extension to QA panels is a well-motivated inference, not a direct demonstration; treat it as a strong hypothesis.
- "Non-diagnostic" is a judgement call. Deciding what is truly irrelevant to a decision is itself contestable. Demographic breakdowns can be diagnostic for equity questions and non-diagnostic for a specific pedagogical one. The label depends on the decision being made, so stripping information requires care, not a blanket rule.
- Under-inclusion has its own risks. Aggressively removing context can hide legitimate confounds (class size, difficulty) that a fair reading needs. The goal is relevance, not minimalism for its own sake — omitting genuinely diagnostic context is its own error.
- Moderators matter. Peters and Rothbart (2000) show framing and typicality change the effect''s size and direction; it is not a fixed constant, and its magnitude in a given report is uncertain.
- Not a substitute for good judgement. Report design reduces one bias; it does not replace trained, deliberate interpretation of evaluation evidence.
How Koji incorporates this
Koji for Education treats the report as a decision instrument, and several capabilities work directly against dilution.
- Signal-first synthesis, not a data dump. Instead of emitting every possible cross-tab, Koji''s AI-assisted reporting is designed to synthesise the diagnostic findings — the themes and patterns that bear on the question — and present them first, with descriptive detail available but subordinate. This inverts the usual clutter-forward dashboard.
- Automatic thematic analysis with evidence. Koji clusters open-text responses into themes and surfaces representative quotations, so a committee sees a concentrated, corroborated signal ("materials inaccessible, reported across three cohorts") rather than raw comment dumps that invite dilution.
- Quality scoring separates signal from noise. Careless-response and low-effort screening mean the material that reaches the report is more likely to be diagnostic, reducing the volume of non-diagnostic content in the first place.
- Structured, scoped reporting. Because studies are built from explicit questions (
open_ended,scale,single_choice, and others), reports can be scoped to the decision at hand rather than dumping the entire item bank — keeping non-diagnostic items out of the reviewer''s field of view. - Consistency across cases. Standardised report structure helps panels weigh comparable evidence across staff and courses, mitigating the accountability-magnified dilution that Tetlock and Boettger document. Framed honestly, Koji is designed to reduce dilution risk through disciplined presentation; it cannot replace a reviewer''s judgement about what is diagnostic.
The same discipline benefits any research reporting: Koji''s core platform at koji.so applies the same AI-moderated interview and synthesis engine to product and customer research, where burying the key insight under secondary data is an everyday failure mode.
Related resources
- Anchoring on the Numeric Score Before Reading the Comments
- Contrast Effects and Narrow Bracketing in Reviewing Evaluation Reports
- Negativity Bias When Reading Open-Text Comments
- Interpreting and Reporting Student Ratings Responsibly
- Top-Box vs Mean: Choosing a Course-Evaluation Reporting Statistic
- Small Mean Differences and Confidence Intervals in Course Evaluation
References
- Nisbett, R. E., Zukier, H., & Lemley, R. E. (1981). The dilution effect: Nondiagnostic information weakens the implications of diagnostic information. Cognitive Psychology, 13(2), 248–277. https://doi.org/10.1016/0010-0285(81)90010-4
- Tetlock, P. E., & Boettger, R. (1989). Accountability: A social magnifier of the dilution effect. Journal of Personality and Social Psychology, 57(3), 388–398. https://doi.org/10.1037/0022-3514.57.3.388
- Peters, E., & Rothbart, M. (2000). Typicality can create, eliminate, and reverse the dilution effect. Personality and Social Psychology Bulletin, 26(2), 177–187. https://doi.org/10.1177/0146167200264005
Related articles
Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees
The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.
The 4.2 vs 4.4 Trap: Why Small Differences in Course-Evaluation Means Are Usually Noise
Faculty and administrators routinely read meaning into tiny gaps between course-evaluation means. Boysen (2015) and Boysen et al. (2014) show this happens even when confidence intervals say the difference is nothing — and that warnings barely help. How to report uncertainty honestly.
Why One Cruel Comment Outweighs Twenty Kind Ones: Negativity Bias in Reading Course Evaluations
Instructors and committees fixate on the harshest open-text comment and discount the praise. Baumeister''s "bad is stronger than good" and the negativity-bias literature explain why, and how to read qualitative course feedback fairly.
Top-Box vs Mean: How to Report Course Evaluation Scores Without Throwing Away Information
Reporting the percentage of students who chose the top box feels intuitive, but collapsing a scale to favorable/unfavorable discards information. Here is what the measurement evidence says and how to report responsibly.