The Chocolate-Chip Evidence: How a Plate of Cookies Inflates Course-Evaluation Scores
A randomised controlled trial found that giving students cookies raised their ratings of the same teacher by an effect size of 0.68 - larger than most documented gender or grading-leniency effects. Here is what the sweetening effect tells us about what a single evaluation score really measures.
Koji Education Team
Product ยท August 18, 2026
If you want a quick, uncomfortable demonstration that a course-evaluation mean is not a clean measure of teaching quality, look at what happens when you hand out cookies.
The short answer (BLUF): In a 2018 randomised controlled trial published in Medical Education, medical students who had free access to chocolate cookies during a course session rated the same teachers, teaching the same content, significantly higher than students in cookie-free sessions - an effect size of 0.68, which is larger than most of the gender or grading-leniency effects that dominate the bias literature. An earlier study replicated the pattern three times with chocolate handed out at evaluation time. The lesson is not "ban snacks." It is that a numeric Student Evaluation of Teaching (SET) score is contaminated by construct-irrelevant factors that have nothing to do with learning, and that a single average cannot tell you which is which. Conversational, probing evaluation can at least ask.
What the trials actually found
The strongest evidence comes from a genuine randomised controlled trial. Hessler and colleagues (2018), working within a curricular emergency-medicine course, randomly allocated 118 third-year medical students into 20 groups. Ten groups had free access to 500 g of chocolate cookies during the session; ten did not. Crucially, the teachers, the educational content and the course materials were identical across both arms - the only thing that varied was the presence of cookies.
The cookie group rated teachers significantly better (113.4 vs 109.2 on the composite; p = 0.001; effect size 0.68). They also rated the course materials as better (10.1 vs 8.4; p = 0.001; effect size 0.66) and gave higher overall summation scores (224.5 vs 217.2; p = 0.008). Read that again: a plate of cookies improved students' rating of the printed handouts they were given. The handouts did not change. The authors concluded plainly that the findings "question the validity of SETs and their use in making widespread decisions within a faculty" (Hessler et al., 2018, Medical Education).
This was not a fluke. Over a decade earlier, Youmans and Jee (2007) had run the low-tech version. An independent experimenter administered evaluations to six discussion sections of the same course; three sections were offered chocolate beforehand, three were not. The chocolate sections gave their instructor higher ratings - and the researchers repeated the manipulation across different classes with the same result each time (Youmans & Jee, 2007, Teaching of Psychology). Their recommendation was telling: because ratings are so easily nudged by extraneous events, evaluations "should be given in more standardized ways to limit the effects of extraneous things like how they are handed out."
Why this is not really about cookies
It is tempting to file this under "amusing study" and move on. That would be a mistake, because cookies are just a clean, randomisable proxy for a whole class of influences that psychometricians call construct-irrelevant variance - systematic score movement caused by something other than the construct you claim to measure (here, teaching quality). If a 30-cent biscuit can move a composite by two-thirds of a standard deviation, then so can:
- Mood and affect. Warm feelings at the moment of rating bleed into the judgement. This is the same mechanism behind the finding that weather, time of day and immediately preceding events shift ratings - see our piece on irrelevant affect, mood and weather bias.
- Reciprocity. A small gift triggers a near-automatic urge to give something back. In a graded course, students cannot reciprocate with money, so they reciprocate with the currency they control: the rating.
- Priming and expressiveness. The cookie result is a cousin of the classic Dr Fox effect, where an expressive, charismatic delivery of empty content earns high ratings. Both show that ratings track the experience of the room, not the transfer of knowledge.
Put together, these are not noise that averages out. They are directional pressures that a savvy instructor can exploit - which is exactly why the effect threatens the fairness of any high-stakes use of SET scores. An instructor who brings snacks, cracks jokes and releases generous grades is not necessarily teaching better; they are running up the score on dimensions the instrument cannot separate from teaching. This is the mechanism underneath grading-leniency bias, too.
Putting the effect size in perspective
Bias in SET is often waved away with "the effects are small." The cookie RCT is a useful corrective. An effect size of 0.68 is large by the conventions of social science - and it sits comfortably above many of the between-group differences that universities happily use to rank instructors, hand out teaching awards, or flag someone for review. When the gap between "excellent" and "needs improvement" on a five-point scale is often a few tenths of a point, and a snack can manufacture a difference of that magnitude, the ranking exercise starts to look indefensible. We make the statistical version of this argument in is a 0.3 difference a real effect? and in the broader review of how strong the evidence for SET bias actually is.
"But critics argue..." - the honest counterarguments
A PhD audience will immediately push back, and rightly so. Let us take the strongest objections head-on.
"These are single-session or single-course studies with small samples." True. The Hessler trial randomised 118 students into 20 clusters; the Youmans work used a handful of sections. Neither is a definitive estimate of a population effect, and cluster-level randomisation with few clusters warrants caution. But the point of an RCT here is not to estimate the "true" cookie coefficient - it is to establish causation under controlled conditions. The internal validity is the whole value: content held constant, allocation randomised, ratings still moved. That is precisely the design SET-defenders demand of bias studies, and it delivered a positive result.
"In the real world, everyone can bring cookies, so it washes out." It does not wash out; it ratchets up. If snacks, leniency and showmanship all raise scores, the equilibrium is that the conscientious instructor who refuses to game the instrument is systematically disadvantaged relative to the one who does. Uniform gaming does not restore fairness; it just moves the whole distribution and rewards the wrong behaviours - a textbook case of Goodhart's law applied to teaching.
"So evaluations are worthless?" No - and this is important. Student voice is essential; students are the only people present for every minute of the course. The problem is not asking students. The problem is compressing a rich experience into a single number and then treating that number as an objective, comparable, decision-grade measure of teaching quality. The cookie studies indict the measurement model, not the act of listening. This is the same distinction we draw in why averaging Likert scores misleads.
What defensible practice looks like
If a snack can move your headline metric, three things follow for any institution that takes its quality assurance seriously.
- Standardise the conditions of collection. Youmans and Jee's own recommendation was procedural: administer evaluations consistently, ideally decoupled from the instructor's physical presence and from reward cues. Online, time-boxed collection with neutral framing removes the most obvious manipulation surfaces.
- Stop treating the mean as the measure. A number cannot distinguish "students learned a great deal" from "students felt good in the room." You need the reasons behind the rating. When a student says the course was excellent, the decision-relevant question is why - and only open, probing follow-up recovers that.
- Read the construct, not the affect. Bias-aware analysis should actively separate comments about learning, challenge and skills gained from comments about likeability, entertainment and, yes, refreshments.
Where Koji fits
This is the gap Koji for Education is built to close. Instead of a static Likert form that captures a mood in a single digit, Koji runs an AI-moderated conversational interview that probes beyond the number: when a student rates a course highly, the AI asks what specifically helped them learn, and follows up until the answer is about pedagogy rather than atmosphere. Because the moderation is standardised and bias-aware, it does not vary with the instructor's charm or catering on the day - removing exactly the human-administration inconsistencies that Youmans and Jee warned about. Automatic thematic analysis then separates substance ("the worked examples finally made recursion click") from affect ("she brought snacks, so nice"), and a quality score flags low-information, purely affective responses. Koji also supports formative, mid-cycle collection, so feedback informs teaching while the course is still running rather than becoming a popularity contest at the end. The same conversational engine powers general user and customer research on the main Koji platform - the difference here is that it is tuned for the evidentiary standards of higher education.
None of this eliminates bias - no instrument can, and any vendor who claims otherwise should be shown the door. But moving from a single manipulable average to a probed, thematically-analysed, consistently-moderated conversation mitigates the exact failure mode the cookie trials expose: the ease with which a construct-irrelevant nudge becomes your official measure of teaching quality.
The plate of cookies is funny. What it reveals about the number on the dashboard is not.