New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods10 min read

The Last Week Counts Double: The Peak-End Rule and What Students Actually Remember About Your Course

Retrospective memory of an experience is dominated by its most intense moment and its ending, not its average or its length. Here is what the peak-end rule means for end-of-semester course evaluations.

Koji Education Team

Product

In brief

When people evaluate an experience after it ends, their summary judgment is driven disproportionately by the most intense moment (the peak) and the final moments (the end), while the experience's total duration is largely ignored — a pattern Kahneman and colleagues named the peak-end rule and duration neglect. A 2022 meta-analysis of 174 effect sizes confirms the peak-end effect is large (r ≈ 0.58) and that duration's effect is essentially nil. Because end-of-semester course evaluations are retrospective judgments of a months-long experience, they are vulnerable to exactly this distortion: a stressful final assessment or a strong closing lecture can colour the whole rating. The defence is not to abandon end-of-term surveys but to triangulate them with in-the-moment (formative) collection — which is where Koji's mid-cycle conversational evaluation is designed to help.

What the research says

The foundational finding comes from Redelmeier and Kahneman (1996), "Patients' memories of painful medical treatments" (Pain, 66(1), 3–8). Patients undergoing colonoscopy reported pain in real time; their retrospective evaluation afterward was predicted almost entirely by the average of the worst moment and the final moment — not by the total amount or duration of pain. In a striking follow-up experiment, deliberately extending a procedure with a few extra minutes of mild (rather than intense) discomfort at the end produced a better remembered experience, even though it added objectively more total discomfort.

The mechanism was established earlier by Kahneman, Fredrickson, Schreiber and Redelmeier (1993), "When More Pain Is Preferred to Less: Adding a Better End" (Psychological Science, 4(6), 401–405), and the broader phenomenon of duration neglect was articulated by Fredrickson and Kahneman (1993). The pattern holds for pleasant as well as unpleasant episodes (Do, Rupert & Wolford, 2008).

Crucially for anyone relying on this, the effect is well-replicated. Alaybek and colleagues (2022), "All's well that ends (and peaks) well? A meta-analysis of the peak-end rule and duration neglect" (Organizational Behavior and Human Decision Processes, 170, 104149) synthesised 174 effect sizes and found:

  • The peak-end effect on retrospective summary evaluations was large (r = 0.581, 95% CI 0.487–0.661) and robust across boundary conditions.
  • The duration effect was essentially nil, confirming duration neglect.
  • Notably, the average of the whole experience predicted summary evaluations about as well as the peak-end score — a nuance we return to in the caveats.

Why it matters for course evaluation in practice

An end-of-semester evaluation asks students to summarise twelve or more weeks in a few ratings — the textbook conditions for peak-end and duration-neglect effects. Practical implications:

  1. The final weeks are over-weighted. A punishing final assessment, a rushed last lecture, or conversely an inspiring closing session can disproportionately shape "overall" ratings. This compounds the well-documented concern about evaluation timing relative to the final exam: run the survey in the shadow of a stressful exam and the "end" you capture is an anxiety peak.
  2. Duration and breadth are under-counted. A course that delivered steady value across the whole term may be rated similarly to one with a single memorable high point, because students neglect duration. Sustained-but-unspectacular teaching can be undervalued.
  3. Comparisons across courses are confounded. Courses differ in where their peaks and endings fall (a capstone project vs a final written exam). Comparing their "overall" scores partly compares the emotional shape of their endings, not their average quality.
  4. It interacts with other memory biases. Peak-end sits alongside the halo effect and thin-slice first impressions: the beginning makes the first impression, the peak and end make the lasting one, and the long middle is compressed.

Limitations and honest caveats

Intellectual honesty requires several qualifications a PhD reader will demand:

  • The average predicts about as well. The Alaybek meta-analysis found that the mean of the whole experience predicted retrospective judgments roughly as strongly as the peak-end composite. This is a genuinely important caveat: peak-end does not prove the middle is ignored entirely — it shows ends and peaks are over-weighted relative to duration, not that everything but the end is discarded.
  • Most evidence is from short, affective episodes. Colonoscopies and lab tasks last minutes; a course lasts months and is cognitively, not just affectively, evaluated. Direct peak-end studies on course evaluations specifically are sparse, so applying the rule to a semester is a reasoned extrapolation, not a settled finding. We flag this honestly rather than overclaiming.
  • Boundary conditions exist. Some studies find the effect weakens when episodes are clearly segmented or when respondents are prompted to consider duration — both of which a well-designed evaluation can partly engineer.
  • Not necessarily 'bias.' If what an institution cares about is the remembered student experience (e.g. for alumni loyalty), peak-end weighting may reflect something real. It is a threat mainly when the evaluation is meant to measure average instructional quality.

How Koji incorporates this

Koji is designed to mitigate peak-end and duration-neglect distortion, without pretending memory can be bypassed:

  • Mid-cycle, formative collection. The most direct defence against retrospective distortion is to collect during the course, not only at the end. Koji supports mid-semester and pulse evaluations so that the steady middle of the term is captured while it is still experienced, not reconstructed from memory months later. This complements the formative-feedback evidence.
  • Probing the basis of a summary judgment. Koji's AI-moderated interview can ask why a student rates the course as they do, surfacing whether an "overall" score is anchored to the final assessment or reflects the whole term ("What stood out across the semester, not just recently?"). This makes peak-end anchoring visible rather than silent.
  • Triangulation across the cycle. Because Koji can hold formative and summative data together, its reporting can compare an end-of-term "overall" rating against in-the-moment signals, flagging when the summary diverges from the lived experience — a triangulation approach no single end-point survey can offer.
  • Thematic analysis of what is remembered. Koji's automatic thematic analysis of open text can reveal recency-weighted content (lots of comments about the final weeks, little about early material), giving committees a cue that duration neglect may be in play.

As always, we frame these as mitigations. Koji cannot remove how human memory works; it can reduce reliance on a single retrospective snapshot and make its limitations legible. Koji's core research platform at koji.so applies the same AI-moderated engine to product and customer research, where peak-end effects shape post-experience satisfaction surveys just as strongly.

Practical checklist

  • Do not rely solely on a single end-of-term "overall" rating; pair it with mid-term/formative collection.
  • Avoid fielding evaluations immediately after a high-stakes final exam if you want a whole-course judgment.
  • When comparing courses, remember their endings differ in emotional shape; weight overall scores accordingly.
  • Use open-text and conversational probes to check whether a summary rating reflects the whole term or just the last weeks.
  • Treat "remembered experience" and "average instructional quality" as different constructs and be explicit about which one a given item measures.

Related Resources

A worked example: two courses, same average, different scores

Imagine two courses that, week by week, delivered identical average value. Course A built steadily and ended with a calm, well-signposted revision session that left students feeling capable. Course B delivered the same average but ended with a chaotic final fortnight and a punishing, poorly-briefed exam. Peak-end logic predicts Course B will score materially lower on "overall" items even though the average experience was equivalent — the painful ending dominates the remembered whole, and the months of equivalent value are compressed by duration neglect.

Now add timing. If both surveys are fielded in the 48 hours after each exam, Course B's "end" is captured at an anxiety peak, amplifying the gap. Move Course B's survey to a week later, after the result is known and the stress has receded, and the same students may rate it noticeably higher. The score moved; the teaching did not. This is why the timing decision is not administrative housekeeping but a measurement choice with real consequences.

Designing around the effect

Three design moves reduce reliance on a single distorted snapshot. First, segment the experience: ask students to reflect on distinct phases of the course rather than only "the course overall", which counteracts duration neglect by forcing attention back onto the middle. Second, capture in the moment: brief pulse check-ins during the term record value while it is being experienced, not reconstructed. Third, separate the constructs explicitly: if you care about remembered experience (relevant to alumni sentiment and recommendation) report it as such, and if you care about average instructional quality, build items and timing that resist end-weighting. Conflating the two is where peak-end quietly does its damage.

References

  • Redelmeier, D. A., & Kahneman, D. (1996). Patients' memories of painful medical treatments: Real-time and retrospective evaluations of two minimally invasive procedures. Pain, 66(1), 3–8. https://doi.org/10.1016/0304-3959(96)02994-6
  • Kahneman, D., Fredrickson, B. L., Schreiber, C. A., & Redelmeier, D. A. (1993). When more pain is preferred to less: Adding a better end. Psychological Science, 4(6), 401–405. https://doi.org/10.1111/j.1467-9280.1993.tb00589.x
  • Alaybek, B., Dalal, R. S., Fyffe, S., Aitken, J. A., Zhou, Y., Qu, X., Roman, A., & Baines, J. I. (2022). All's well that ends (and peaks) well? A meta-analysis of the peak-end rule and duration neglect. Organizational Behavior and Human Decision Processes, 170, 104149. https://doi.org/10.1016/j.obhdp.2022.104149
  • Do, A. M., Rupert, A. V., & Wolford, G. (2008). Evaluations of pleasurable experiences: The peak-end rule. Psychonomic Bulletin & Review, 15(1), 96–98. https://doi.org/10.3758/PBR.15.1.96