Why One Cruel Comment Outweighs Twenty Kind Ones: Negativity Bias in Reading Course Evaluations
Instructors and committees fixate on the harshest open-text comment and discount the praise. Baumeister''s "bad is stronger than good" and the negativity-bias literature explain why, and how to read qualitative course feedback fairly.
Koji Education Team
Product
In short: When an instructor reads thirty positive open-text comments and one cutting one, it is the cutting one they remember, reread, and lose sleep over. This is not a character flaw — it is negativity bias, one of the most robust findings in psychology. Baumeister, Bratslavsky, Finkenauer and Vohs (2001) summarised it as "bad is stronger than good": negative information is weighted more heavily, processed more thoroughly, and resists disconfirmation more stubbornly than positive information of equal magnitude. Applied to course evaluation, this means both instructors and the committees who read their feedback systematically over-weight a handful of harsh comments. The remedy is to make qualitative analysis prevalence-based: report how common a theme is, not how vividly it was phrased.
What the research says
In 2001, Roy Baumeister and colleagues published "Bad Is Stronger Than Good" in Review of General Psychology — a sweeping review that gathered evidence across an unusually wide range of domains and reached a single, blunt generalisation. Bad events, bad emotions, bad feedback, and bad information exert a greater psychological impact than good events, emotions, feedback, and information of comparable size. Their evidence spanned everyday interactions, major life events, close relationships, social networks, and — directly relevant here — learning and feedback processes. Among the specific patterns they document: bad feedback has more impact than good; bad information is processed more thoroughly and remembered better; and bad impressions form faster and are more resistant to disconfirmation than good ones. The authors frame the asymmetry as adaptive — organisms that attend more to threats than to rewards survive longer — but note that adaptiveness does not make it accurate as a basis for judgement.
Independently and at almost the same moment, Paul Rozin and Edward Royzman published "Negativity Bias, Negativity Dominance, and Contagion" (2001, Personality and Social Psychology Review), arriving at the same hypothesis from a different route. They decompose the phenomenon into components — greater negative potency, steeper negative gradients, and negativity dominance (the tendency for the negative element of a mixed set to dominate the overall impression). The two papers cite each other as convergent discoveries of the same principle, which is part of why negativity bias is now treated as a foundational regularity rather than a single-study curiosity.
The mechanism matters for how feedback is read, not just how it is given. Negativity dominance predicts that a set of comments that is 95% positive will not produce a "95% positive" impression; the negative 5% will colour the whole, because the bad simply weighs more in the integration. This is the cognitive engine behind the familiar staffroom phenomenon of the lecturer who received glowing feedback all term and can quote only the one anonymous insult verbatim.
Why it matters for course evaluation in practice
Open-text comments are the most valued and most volatile part of a course evaluation. They carry the richest diagnostic information — and they are read by exactly the human cognitive system the negativity-bias literature describes. Two distinct harms follow.
The first is to instructor wellbeing and behaviour. A well-documented consequence of negativity bias in evaluations is that faculty disproportionately remember and ruminate on harsh or abusive comments, which can drive defensive, risk-averse teaching and real distress (a concern the duty-of-care literature on abusive comments takes up directly). When the one venomous line outweighs the term's worth of constructive praise, instructors optimise to avoid the venom rather than to teach well.
The second harm is to decision quality. Committees and heads of department reading qualitative feedback for appraisal are no more immune to negativity dominance than anyone else. A single memorably negative comment can anchor an entire judgement of a course, crowding out a representative reading of the whole. Because negative impressions also "resist disconfirmation," subsequent positive evidence is discounted rather than allowed to rebalance the picture. The result is qualitative analysis that is vivid-driven rather than prevalence-driven — the loudest comment wins, regardless of how typical it is.
The corrective principle is straightforward to state and surprisingly hard to practise: weight themes by how common they are, not by how strongly they are phrased. A criticism raised once, however eloquently, is a data point of n=1; a mild concern raised by a third of the cohort is a signal. Good qualitative practice — systematic thematic coding with prevalence counts, and explicit separation of frequency from intensity — exists precisely to discipline the reader's negativity bias. Reporting the proportion of comments expressing each theme, and presenting representative rather than extreme exemplars, structurally counteracts negativity dominance.
How large is the asymmetry? The literature does not fix a single exchange rate, but the practical lesson of negativity dominance is that positive and negative comments do not cancel one-for-one — it can take many affirming remarks to offset the felt weight of one harsh line. That is precisely why counting matters. If a thematic report shows that "pacing too fast" appears in thirty-one per cent of comments while a single respondent calls the lecturer "the worst I have ever had," the responsible reading foregrounds the pacing theme and treats the insult as the unrepresentative outlier it is. Numbers discipline the instinct; without them, the most vivid comment wins by sheer salience rather than by how many students actually shared it.
Limitations and honest caveats
Several caveats keep this from being a licence to dismiss criticism. First, negativity bias describes a tendency in how feedback is weighted, not a claim that negative feedback is wrong. Some harsh comments are accurate and important; the bias to be corrected is over-weighting by vividness, not the existence of criticism. A prevalence-based reading can itself fail if it buries a rare but serious signal — a single credible report of discrimination or misconduct is not noise to be averaged away, and frequency-weighting must always be overridden for safeguarding-relevant content.
Second, the Baumeister and Rozin–Royzman papers are broad reviews synthesising heterogeneous studies, not a single controlled experiment on course evaluations specifically. Their generalisation is well-supported but, like any review, aggregates effects of varying size and quality, and the magnitude of negativity bias is moderated by context, individual differences, and how feedback is framed. Third, more recent work has urged refinements — negativity bias is not perfectly universal and can be attenuated by expectations and by positive framing — so "bad is stronger than good" is a strong default, not an exceptionless law. Finally, the practical inference (read by prevalence) is itself a judgement call about how to aggregate; it does not tell you which themes matter most, which still requires substantive expertise.
How Koji incorporates this
Koji for Education is built to give a reader the prevalence of a theme before its vividness can hijack the interpretation — the structural counter to negativity dominance:
- Prevalence-based thematic analysis. Koji's automatic thematic analysis of open-text responses is designed to report how many respondents raised each theme and what proportion of comments are positive, negative, or mixed — so a single scorching comment is visibly one data point, not the headline. This reframes qualitative reading from "what was the worst thing said?" to "what did most students actually experience?"
- Representative exemplars, not just extremes. Where a naive read gravitates to the most quotable insult, Koji surfaces representative quotes for each theme alongside frequency, giving committees a balanced evidence base rather than the outlier that negativity bias would otherwise elevate.
- Sentiment proportions over anecdote. Koji reports the distribution of sentiment across the cohort, so a 95%-positive set reads as 95% positive instead of being recoloured by the negative 5%.
- Conversational depth that contextualises criticism. Because Koji's AI-moderated interviews probe why a student felt as they did, a harsh reaction is captured with its reasoning and severity, helping a reader distinguish a fixable irritation from a serious problem rather than reacting to tone. Koji's core research platform at koji.so applies the same prevalence-aware analysis to product and customer feedback, where one furious review can otherwise dominate a roadmap.
- Duty-of-care filtering. Koji is designed to flag abusive or personal-attack content so it can be handled appropriately, protecting instructor wellbeing while preserving the substantive signal — with safeguarding-relevant content escalated rather than averaged away.
Koji is designed to mitigate negativity dominance in how feedback is read; it does not, and should not, suppress genuine criticism — the goal is a representative reading, not a flattering one.
Related Resources
- Abusive Open-Text Comments in Course Evaluations: The Evidence and the Duty of Care
- What Can Open-Text Student Comments Tell You That Likert Scores Cannot?
- Text Analytics for Open-Ended Student Comments: What NLP Can and Cannot Tell You
- How Reliable Is Your Coding of Open-Text Feedback? Inter-Rater Reliability and Thematic Analysis
- Interpreting and Reporting Student Ratings Responsibly
References
- Baumeister, R. F., Bratslavsky, E., Finkenauer, C., & Vohs, K. D. (2001). Bad is stronger than good. Review of General Psychology, 5(4), 323–370. https://doi.org/10.1037/1089-2680.5.4.323
- Rozin, P., & Royzman, E. B. (2001). Negativity bias, negativity dominance, and contagion. Personality and Social Psychology Review, 5(4), 296–320. https://doi.org/10.1207/S15327957PSPR0504_2
Related articles
Text Analytics for Open-Ended Student Comments: What NLP Can and Cannot Tell You
Natural-language processing can turn thousands of free-text course-evaluation comments into themes and sentiment at scale — but the research (Cunningham-Nelson 2019; Sunar & Khalid 2023) shows where automated analysis is reliable and where human judgement is still required.
Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees
The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.
How Reliable Is Your Coding of Open-Text Feedback? Inter-Rater Reliability and Thematic Analysis
When you turn thousands of free-text comments into themes and counts, how do you know the coding is trustworthy? O Connor and Joffe (2020) on intercoder reliability, Braun and Clarke on thematic analysis, and what rigorous qualitative QA looks like.
Do Students'' Written Comments Match Their Ratings? What Concordance Tells You
Open-text comments and Likert scores usually agree — but the gaps are where the insight lives. What Alhija & Fresko (2009) and Brockx et al. (2012) found about the consistency between qualitative and quantitative course-evaluation data, and how to read it.