New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods8 min read

Do Cookies, Treats, and Mood Bias Course Evaluations?

Two controlled studies show that giving students chocolate or cookies before an evaluation measurably raises teaching scores. What the affect heuristic means for the validity of course evaluations — and how to design around it.

Koji Education Team

Product

In brief: Yes. In two controlled studies, students who were given chocolate or had cookies available before completing a course evaluation rated the teaching, the material, and the course overall significantly higher than students who were not — despite identical instruction. The cleanest experiment found medium effect sizes (Cohen's d ≈ 0.5–0.7). The lesson is not "ban snacks": it is that a single-number satisfaction rating absorbs transient mood, so evaluation design and interpretation must separate how students felt in the room from how well they were taught.

When a department compares two lecturers and one scores 4.6 while the other scores 4.2, the difference is routinely treated as a signal about teaching quality. A small but methodologically clean body of evidence suggests that a difference of that size can be manufactured by something as trivial as a plate of cookies. This matters because course evaluations are increasingly used in promotion, probation, and quality-assurance decisions, where a few tenths of a point can change an outcome. If mood can move the needle as much as pedagogy, the number is measuring something other than what we think.

What the research says

The most-cited demonstration is Youmans and Jee (2007), published in Teaching of Psychology. Six discussion sections of the same undergraduate course completed standard evaluations administered by an independent experimenter (not the instructor). Before three of the sections filled in the forms, the experimenter offered them chocolate; the other three received nothing. Students who were offered chocolate rated the instructor more positively than those who were not. Because the sections were taught identically and the manipulation was the chocolate alone, the difference is attributable to the treat — or, more precisely, to the mood and reciprocity it induced — rather than to teaching.

The stronger design is Hessler et al. (2018), published in Medical Education. The team randomly allocated 118 third-year medical students into 20 groups across an emergency-medicine course; ten groups had free access to 500 g of chocolate cookies during the session and ten did not. The same teachers delivered the same content with the same materials. The cookie groups rated teachers significantly higher (113.4 vs 109.2 on the teaching scale, p = 0.001, effect size 0.68), rated the course material higher (effect size 0.66), and gave higher overall summation scores (224.5 vs 217.2, p = 0.008, effect size 0.51). Randomisation, identical content, and a pre-registered-style comparison make this close to a true experiment, and the effect sizes are not trivial — they are in the medium range that, in any other context, would be reported as a meaningful intervention effect.

Both findings sit on a well-established psychological foundation: the affect-as-information or mood-as-information hypothesis. Schwarz and Clore (1983) showed in Journal of Personality and Social Psychology that people lean on their current mood as a shortcut when forming evaluative judgements — when asked "how is your life going?" on sunny versus rainy days, respondents reported higher life satisfaction on sunny days, unless their attention was drawn to the weather as the source of their mood. Applied to course evaluations: a question like "overall, how would you rate this course?" invites students to consult their momentary feeling, and a pleasant, well-fed, mildly grateful feeling is read as evidence that the course was good. The cookie is not bribery in any cynical sense; it is a small positive affect that gets misattributed to teaching quality.

This connects to a broader literature the field already documents: instructor expressiveness inflating ratings independent of content (the Dr Fox effect), rapid first-impression judgements (thin-slice effects), and the general halo effect by which one salient positive feature colours every rating on the form. Cookies are simply a particularly tidy, manipulable instance of the same underlying mechanism.

Why it matters for course evaluation in practice

The practical danger is not that lecturers will start bribing students en masse. It is what these studies reveal about the instrument. If a transient mood manipulation produces a half-point shift, then any uncontrolled source of mood — the room temperature, the time of day, a popular guest speaker the week before, whether the evaluation followed a returned assignment with good marks — is silently embedded in the score. For quality assurance, three consequences follow.

First, small numeric differences between courses or instructors are not trustworthy signals. A 0.3–0.5 gap is within the range that mood alone can generate. Ranking staff by mean score, or setting a numeric threshold for "satisfactory teaching," treats noise as if it were measurement. This is the same caution our guidance on small mean differences and confidence intervals and on interpreting and reporting student ratings responsibly develops in detail.

Second, standardised administration matters more than most institutions enforce. Youmans and Jee explicitly framed their study as an argument for controlling the conditions under which evaluations are collected. If one cohort completes the form in the last five minutes of a tense exam-prep session and another completes it after a relaxed wrap-up, the comparison is contaminated before a single response is logged. Consistency of timing, channel, and framing is a precondition for comparability — a theme that also drives our analysis of timing effects.

Third, a single global satisfaction item is the most mood-susceptible question you can ask. The more an item invites a holistic gut feeling ("rate this course overall"), the more room there is for affect to substitute for evidence. Specific, behaviourally anchored questions ("how often did you receive feedback in time to use it?") give the mood heuristic less to work with, because they ask the student to recall a concrete fact rather than consult a feeling.

Limitations and honest caveats

A critical reader should not over-read these studies. Both are single-institution, single-discipline experiments — undergraduate psychology and medical education respectively — and neither is large by epidemiological standards. The Hessler cookie effect of d ≈ 0.5–0.7 is impressive but comes from 20 groups; the confidence intervals around such estimates are wide, and the true population effect could be appreciably smaller. Youmans and Jee report a directional effect but with a modest sample of sections.

Second, ecological validity cuts both ways. A deliberate chocolate manipulation is more salient than the ambient mood fluctuations of ordinary evaluation administration; we should not assume every real-world course gap reflects a cookie-sized confound. It is equally possible that in routine practice, mood effects partly cancel out across a large, varied cohort. The studies establish that the channel exists and can be large; they do not establish how much of any particular institution's variance it explains.

Third, there is a publication-and-novelty bias risk: a "cookies change evaluations" finding is memorable and citable, which can inflate its apparent importance relative to less surprising null results that never get written up. The responsible reading is the conservative one: treat the studies as a proof that the satisfaction rating is mood-permeable, not as a precise estimate of contamination.

Finally, mood effects do not mean evaluations are worthless. Decades of work (see our summary of the validity-versus-reliability question) show that well-constructed, multidimensional ratings carry real information about aspects of teaching. The cookie studies sharpen, rather than refute, that conclusion: they tell you which parts of the instrument to distrust (global, mood-laden items read in isolation) and which to lean on (specific, triangulated, well-aggregated evidence).

How Koji incorporates this

Koji is an AI-native course-evaluation platform built around an AI-moderated conversational interview rather than a static form, and several of its design choices are aimed squarely at the mood-misattribution problem this research identifies.

  • Probing beyond the gut number. When a student gives a high or low global rating, Koji's conversational engine can follow up — "what specifically led you to that?" — eliciting a concrete reason. A rating propped up only by good mood tends to produce thin, non-specific justification, which is visible in the transcript and in the automatic quality scoring of each response. This is designed to surface affect-driven answers, not to eliminate them, so reviewers can weight them accordingly.
  • Structured, behaviourally anchored questions. Koji supports open_ended, scale, single_choice, multiple_choice, ranking, and yes_no question types. Evaluation designers can lean on specific, evidence-seeking items (e.g. a yes_no on whether feedback arrived in time, a ranking of which course elements helped learning most) that give the mood heuristic less leverage than a lone "overall satisfaction" slider.
  • Thematic analysis over single scores. Rather than reducing a course to one mood-sensitive mean, Koji performs automatic thematic analysis across open-text and conversational responses, so the signal comes from recurring, content-specific themes that are harder to fake with a transient good feeling.
  • Bias-aware reporting and uncertainty. Koji's reporting is designed to present distributions and caveats rather than encouraging spurious precision, consistent with the principle that a 0.3-point gap is rarely a real difference.
  • Standardised administration. Because collection is conversational and delivered consistently across cohorts, Koji reduces the uncontrolled administration differences (a rushed end-of-class form here, a relaxed one there) that let mood effects creep in unequally.

None of this eliminates the affect heuristic — no instrument can stop a human from feeling good and reading that feeling as quality. The aim is to make mood-driven responses identifiable and to dilute their weight by triangulating across specific, content-anchored evidence. Beyond the classroom, Koji's core research platform at koji.so applies the same AI-moderated interview engine to product and customer research, where mood-as-information contaminates satisfaction surveys in exactly the same way.

Related Resources

References

  1. Youmans, R. J., & Jee, B. D. (2007). Fudging the numbers: Distributing chocolate influences student evaluations of an undergraduate course. Teaching of Psychology, 34(4), 245–247. https://doi.org/10.1080/00986280701700318
  2. Hessler, M., Pöpping, D. M., Hollstein, H., Ohlenburg, H., Arnemann, P. H., Massoth, C., Seidel, L. M., Zarbock, A., & Wenk, M. (2018). Availability of cookies during an academic course session affects evaluation of teaching. Medical Education, 52(10), 1064–1072. https://doi.org/10.1111/medu.13627
  3. Schwarz, N., & Clore, G. L. (1983). Mood, misattribution, and judgments of well-being: Informative and directive functions of affective states. Journal of Personality and Social Psychology, 45(3), 513–523. https://doi.org/10.1037/0022-3514.45.3.513

Related articles

analysis-reporting

Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees

The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.

analysis-reporting

The 4.2 vs 4.4 Trap: Why Small Differences in Course-Evaluation Means Are Usually Noise

Faculty and administrators routinely read meaning into tiny gaps between course-evaluation means. Boysen (2015) and Boysen et al. (2014) show this happens even when confidence intervals say the difference is nothing — and that warnings barely help. How to report uncertainty honestly.

research-methods

The Halo Effect in Course Evaluations: When One Impression Colours Every Rating

When students like an instructor, that single global impression bleeds into their ratings of unrelated specifics. From Nisbett & Wilson (1977) to Feistauer & Richter (2018) and Cannon & Cipriani (2021), here is how the halo effect distorts itemised course evaluations — and what to do about it.

research-methods

The Dr. Fox Effect: Does Lecturer Charisma Inflate Course Evaluations?

The 1973 Dr. Fox experiment, its meta-analytic re-interpretation, and the 2014 re-revisitation — what the evidence really says about whether expressive delivery seduces students into rating empty teaching highly, and how to evaluate substance over showmanship.