Chocolate, Sunshine, and a Free Half-Point: How Transient Mood Contaminates Course Evaluations
A meaningful slice of end-of-term ratings reflects a student's mood at the instant they hit submit — mood that has nothing to do with your teaching. A famous experiment moved ratings with a handful of chocolate. Here is what the irrelevant-affect research shows, where it is contested, and why standardising how you collect feedback matters as much as fixing the questions.
Koji Education Team
Product ·
Bottom line up front: Some of the variance in your course-evaluation scores is not about the course. It reflects the student''s transient emotional state at the moment of judgement — a state shaped by the weather, the room, whether the previous class went badly, and, in one now-famous experiment, whether someone handed them chocolate first. Because this "irrelevant affect" is often correlated with how and when you collect evaluations rather than being purely random, it does not reliably average away. Standardising the collection context is therefore not administrative housekeeping; it is a measurement decision that protects the validity of the number.
The experiment that should unsettle you
In 2007, Robert Youmans and Benjamin Jee published a small, elegant study in Teaching of Psychology with the title "Fudging the Numbers: Distributing Chocolate Influences Student Evaluations of an Undergraduate Course." Six discussion sections of the same course, taught by the same instructor, completed evaluations administered by an independent experimenter. Three sections were offered chocolate immediately before filling in the forms; three were not. The sections offered chocolate rated the instructor and course more positively — despite evaluating a class they had experienced identically to the sections that got nothing.
Sit with what that means. A variable with zero relationship to teaching quality — a piece of confectionery handed out at the door — moved the aggregate rating. The students were not corrupt or careless; they were human, and human judgements of complex, hard-to-summarise experiences lean on whatever affect is salient at the moment of asking. The chocolate did not change the teaching. It changed the mood, and the mood changed the number.
Why this happens: mood as information
The mechanism has a well-developed theoretical home. Norbert Schwarz and Gerald Clore''s 1983 studies, "Mood, misattribution, and judgments of well-being," proposed the "mood-as-information" account: when people evaluate something diffuse and effortful to assess — how satisfied am I with my life? how good was this course? — they often consult their current feelings as a shortcut, implicitly asking "how do I feel about this?" and reading the answer off their present mood. In their best-known demonstration, respondents interviewed on sunny days reported higher life satisfaction than those interviewed on rainy days — unless the interviewer first drew attention to the weather, at which point people discounted their mood and the effect disappeared.
A course evaluation is precisely the kind of judgement that invites this shortcut. "Overall, how would you rate this course?" asks a student to compress thirteen weeks of varied experience into a single number, quickly, at the end of a busy term. That is cognitively expensive, so the current mood does some of the work — and the current mood was set by things the course did not cause.
An honest caveat, because it strengthens the point
Intellectual honesty requires a flag here: the specific weather finding has a contested replication record. Later large-sample studies (and a well-known 2011 re-examination, "Good Weather for Schwarz and Clore," plus subsequent reanalyses) failed to reproduce a robust sunshine-to-life-satisfaction effect, and the size of any weather effect appears small and condition-dependent. If we cited the rainy-day study as hard proof that weather swings your evaluations, we would be committing exactly the over-interpretation this blog keeps warning against.
But notice that the caveat does not rescue the evaluation form — it relocates the problem. The principle the chocolate study demonstrates does not depend on weather at all: it depends on transient affect being read into a global judgement. Whether the mood comes from sunshine, a handful of chocolate, a warm room, a friendly administrator, or the fact that the class before was a disaster, the pathway is the same. The chocolate experiment is a clean, direct, on-target demonstration in the exact setting we care about; the weather literature is a reminder that any single source of mood is noisy and easily overstated. Both lessons point the same way: control the collection context, and stop treating the overall rating as if it were sealed off from the moment it was produced.
Noise or bias? The distinction that decides whether it matters
The natural objection is: "Fine, mood adds error — but random error averages out. With a whole class responding, the moods cancel and the mean is clean." This is the single most important thing to get right, and it is usually got wrong.
Random mood would average out. But irrelevant affect in course evaluation is frequently not random with respect to the collection procedure — which makes it bias, not noise. Consider:
- A shared shock. If evaluations open the morning after a punishing exam in a different module, an entire cohort arrives in a sour mood simultaneously. That does not cancel; it shifts the whole distribution.
- The room and the moment. Sections evaluated in a cramped, hot afternoon slot versus a bright morning slot are not exchangeable. In-person forms handed out at the end of a gruelling three-hour lab carry the lab''s residue.
- The administrator effect. The chocolate study is really a study of how the form is administered. A warm, apologetic "sorry to keep you, thanks so much" administration primes different affect than a curt one. Where humans administer inconsistently, that inconsistency becomes systematic between-instructor variance.
- Small classes. Averaging only rescues you at scale. In a seminar of twelve, three students in a foul mood is a quarter of your data. As we have argued in What Is a Fair Score for a Class of 12?, small-cohort means are already fragile; a shared mood tips them.
This is why irrelevant affect belongs in the bias conversation alongside its better-known cousins. It is a close relative of the peak-end rule, where the emotional tone of the final moments dominates recall, and of contrast effects, where the class before yours colours your ratings. All three share a root: the evaluation captures a state, and the state is contaminated by things outside the construct you meant to measure — textbook construct-irrelevant variance.
What to actually do about it
You cannot legislate students'' moods, and you should be sceptical of anyone who claims to "eliminate" this bias. But you can shrink and de-systematise it:
- Standardise administration. The lesson Youmans and Jee themselves drew is that evaluations should be delivered in consistent ways so that extraneous factors like how the form is handed out stop varying between instructors. Consistency does not remove mood; it stops mood from becoming a between-instructor bias.
- Decouple from the end-of-term crunch. Collecting some feedback mid-cycle, away from the exam-season emotional spike, samples a different — and often more considered — affective state. This is the formative case made in Formative vs Summative Course Evaluation.
- Ask for reasons, not just ratings. A number absorbs mood silently. A student who has to explain a low rating often reveals that the complaint is about the timetable, or the previous class, or exam stress — information that lets you discount the irrelevant affect instead of banking it as a verdict on teaching.
- Read spread, not just the mean. A shared mood shift and a genuine split of opinion can produce the same average. Watching dispersion and bimodality helps you tell "everyone was a bit grumpy" from "half the class was failed by the course."
Where Koji fits
Two features of Koji for Education bear directly on irrelevant affect. First, its AI moderation is standardised by construction: every student is greeted and prompted in the same calibrated way, removing the human-administrator inconsistency that the chocolate study exposed — no warm-versus-curt hand-out, no confectionery on the desk. Second, because Koji runs AI-moderated conversational interviews rather than a bare Likert form, it can probe beyond the number: when a rating is low, it asks why, and its automatic thematic analysis separates "the course had real problems" from "I was in a terrible mood because of the exam this morning." Add mid-cycle formative collection and you sample mood at more than one moment rather than freezing the whole judgement at the single most stressful point in the calendar.
The same conversational engine powers the main Koji platform for customer and user research, where momentary mood contaminates satisfaction scores in exactly the same way — a five-star rating left in a bad ten minutes is no more about the product than a low course score is about the teaching.
None of this makes evaluations mood-proof. It makes them mood-aware — collected consistently, sampled across time, and read with the humility to ask whether a dip is the teaching or just a rainy Tuesday.