The Peak-End Rule: Why End-of-Term Timing Distorts What Students Remember About Your Course
End-of-term evaluations measure the course students remember, not the course they experienced — and memory is ruled by the peak-end rule and duration neglect. Here is what the evidence says, and what to do about it.
Koji Education Team
Product · July 4, 2026
The end-of-semester evaluation does not measure the course a student experienced. It measures the course a student remembers — and remembering is not a faithful recording. Four decades of research on how people summarise past experiences shows that a retrospective judgement is dominated by two moments: the emotional peak and the ending. The steady middle — most of your teaching — is compressed to almost nothing. If your evaluation window opens in the last week of term, you are sampling memory at its most distorted, and calling the result "the student experience".
This is not a reason to distrust student feedback. It is a reason to be precise about what a single end-of-term number can and cannot tell you, and to change when and how you collect it.
The peak-end rule, briefly and accurately
In 1993, Barbara Fredrickson and Daniel Kahneman proposed the snapshot model of remembered utility: people do not integrate an experience moment-by-moment when they judge it in hindsight. Instead they rely on a few emotionally representative snapshots — above all the most intense moment (the peak) and the final moment. Averaged, these two predict the remembered evaluation remarkably well.
The most cited demonstration is clinical, not educational. Redelmeier and Kahneman (1996) recorded real-time pain from 154 colonoscopy patients and 133 lithotripsy patients, then collected retrospective evaluations. Patients' judgements of total pain correlated strongly with the peak intensity and the intensity during the last three minutes of the procedure — but not with its duration. A longer, objectively worse procedure that ended gently was remembered as less bad than a shorter one that ended at high intensity. Kahneman named the companion effect duration neglect: how long an experience lasts has almost no bearing on how it is later judged.
A 2022 meta-analysis in Organizational Behavior and Human Decision Processes ("All's well that ends (and peaks) well?") confirmed that the peak-end effect on retrospective summary evaluations is large and robust, and consistently stronger than the effect of duration — which was, as the theory predicts, essentially nil. This is not a fragile lab curiosity; it is one of the more durable findings in the psychology of judgement.
Why this is a course-evaluation problem, not a trivia question
A semester is a long experience — twelve or thirteen weeks of lectures, seminars, coursework and, at the end, assessment. When a student completes an evaluation in week 12, the snapshot model predicts their overall rating will be anchored on:
- The peak. The most emotionally intense moment of the term. For many students that is not a brilliant lecture in week 4 — it is the assessment crunch, the hardest problem set, the group project that went sideways, or the exam looming next week. Emotional peaks are disproportionately negative and disproportionately assessment-related.
- The end. The final weeks — revision, deadlines, exam stress — which for structural reasons are among the most stressful of the whole course, and which have nothing to do with the quality of the teaching in weeks 1 through 9.
Meanwhile the eleven competent, well-paced weeks in the middle are flattened by duration neglect. A term that was 80% strong and 20% stressful at the end does not get remembered as "80% good". It gets remembered as "stressful", because the stressful part was the peak and the end. The average-a-Likert-score machinery then converts that distorted memory into a 3.4 and files it as fact.
This interacts with everything else we know about evaluation timing. Collecting feedback after grades are released introduces a different distortion — outcome-driven reappraisal — which is why the timing of course evaluation is a validity decision, not an administrative one. The peak-end problem is the deeper, more general version: even before grades, the retrospective window itself is the problem.
The strongest counterargument — "But doesn't every method have this problem?"
A careful reader will object: All retrospective evaluation is memory-based. You cannot ask people about a past experience without querying memory. So peak-end distortion is unavoidable and therefore not actionable.
This is half right, and the half that is wrong is the important half. Two things follow from the evidence that the objection misses.
First, distortion is not uniform — it is a function of design. The peak-end rule bites hardest when a single summary judgement is demanded about a long, variable experience assessed only at its stressful conclusion. It bites far less when feedback is collected closer to the moments being judged, when questions are specific rather than global, and when the instrument invites narrative reconstruction rather than a one-shot gestalt rating. Timing and question design are levers, and the research tells us which way to pull them.
Second, knowing the mechanism lets you read the data correctly. A department that understands duration neglect will not treat a dip in end-of-term scores during a heavy-assessment semester as evidence that teaching quality fell. It will treat it as partly an artefact of when memory was sampled. The alternative — naïve averaging — launders a predictable cognitive bias into a personnel signal.
What the evidence implies you should do
The peak-end rule does not say "abandon student feedback". It says three concrete things.
1. Move some collection out of the retrospective window. In-the-moment or mid-course feedback sidesteps duration neglect because it does not ask students to compress a whole term into one snapshot. This is the case for experience sampling and momentary feedback over the end-of-term survey, and for formative, mid-cycle evaluation that most universities skip. Feedback gathered in week 6 about week 6 is not governed by the peak-end rule in the way a week-12 verdict on the whole course is.
2. Ask specific questions, not global gestalt ratings. "Overall, how would you rate this course?" is precisely the summary judgement the snapshot model predicts will collapse onto peak and end. Concrete, dimension-specific questions ("How clear were the assessment instructions in the first half of term?") force reconstruction and resist the halo of a single remembered moment — a point that connects directly to why one impression can contaminate ten questions.
3. Separate the affective peak from the quality signal. When feedback is a number, the peak and end are baked in invisibly. When feedback is a conversation, you can hear the difference between "the exam period was brutal" and "the teaching was unclear" — two very different findings that a Likert mean fuses into one lukewarm score.
Where Koji fits
Koji for Education is built for exactly this reconstruction problem. Instead of a single end-of-term Likert form, Koji runs AI-moderated conversational interviews that probe beyond a number — asking students to describe specific moments, not deliver a compressed verdict. That structure works against duration neglect: a student who says "it was stressful" is asked when, and about what, so a peak caused by the exam is not silently attributed to the teaching.
Koji supports formative, mid-cycle collection, so evidence is gathered while the experience is fresh rather than only at the memory-distorting end. Its automatic thematic analysis of open-text responses surfaces whether a negative signal clusters around assessment stress (a peak-end artefact) or around genuine instructional issues that recur across weeks. And because the AI moderator is standardised and bias-aware, it applies the same probing consistently — no human interviewer whose own recency bias shapes which follow-ups get asked. Koji is careful about the claim: this mitigates and surfaces peak-end distortion; it does not eliminate the fact that all retrospective feedback queries memory.
The same conversational interview engine powers the main Koji platform for general user and customer research, where retrospective-recall bias is just as corrosive to product feedback as it is to course feedback.
The bottom line
The end-of-term evaluation is not a neutral instrument reading off "the student experience". It is a memory probe fired at the single worst moment in the semester to fire it. The peak-end rule and duration neglect are not reasons to stop listening to students — they are reasons to collect feedback earlier, ask about specifics, and read a lone end-of-term mean with the humility a well-understood cognitive bias demands. A course that was steadily good for eleven weeks and stressful for two deserves better than to be remembered only for the two.
Ready to collect feedback that captures the whole term, not just its most stressful moment? See how Koji for Education works.