New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Methodology9 min read

Experience Sampling: Why In-the-Moment Course Feedback Beats the End-of-Term Survey

The end-of-term evaluation asks students to reconstruct a 12-week experience from memory in five minutes. Decades of research on retrospective self-report say that memory is a poor witness. Experience sampling is the methodological alternative — and it changes what feedback can tell you.

Koji for Education

Research & Editorial Team · June 24, 2026

Bottom line up front: The standard course evaluation asks a student, in the last week of term, to summarise twelve weeks of teaching into a handful of numbers. That is a retrospective self-report — and the measurement literature is emphatic that retrospective reports are distorted by predictable memory biases: recency, peak-end effects, and mood-congruent recall. Experience sampling (also called ecological momentary assessment) is the alternative: brief, repeated measurements taken during the experience, in context. It does not just improve accuracy; it changes what feedback can do — turning evaluation from an autopsy into a live signal you can still act on. This post explains the method, the evidence, the honest limitations, and how an AI-native platform makes it practical at course scale.

The hidden assumption in the end-of-term survey

Every end-of-term evaluation rests on an assumption that is rarely stated: that a student can accurately reconstruct a whole semester from memory at a single moment. Cognitive psychology has been dismantling that assumption for forty years.

Daniel Kahneman's distinction between the experiencing self and the remembering self is the crux. What we remember about an extended experience is not its average; it is dominated by its emotional peaks and by how it ended — the peak-end rule. A semester that was clear and well-paced for ten weeks but chaotic in the final fortnight will be remembered, and rated, as chaotic. The reverse is also true. (We cover the rating-specific version of this in our piece on recency and peak-end effects in course evaluation.)

This is not a minor wrinkle. It means the end-of-term mean is a biased estimator of the term — systematically over-weighting the final weeks and the most emotionally charged moments, and under-weighting the long stretches of ordinary, effective teaching.

What the measurement literature says about retrospective reports

The case against retrospection is not speculative; it comes from a large methodological literature, much of it in clinical and health psychology where the stakes for accurate self-report are high.

  • Real-time and retrospective reports diverge. Reviews of ecological momentary assessment (EMA) document systematic discrepancies between moment-by-moment reports and later recall of the same mood, symptoms and behaviours. Shiffman, Stone & Hufford's (2008) Annual Review of Clinical Psychology overview is the standard reference: people preferentially recall experiences that were recent, salient, unusual, or congruent with their current mood.
  • EMA is designed to fix exactly this. EMA and the experience sampling method (ESM), developed by Csikszentmihalyi and Larson and refined by Stone and Shiffman, involve repeated sampling of current experience in the natural environment, explicitly to minimise recall bias and maximise ecological validity.
  • It travels to education. A systematic review of ESM in schools (Frontiers in Psychology, 2022) documents its growing use to capture students' in-the-moment engagement and social experience — evidence that the method is viable in classrooms, not only clinics.

The logic transfers cleanly to course evaluation. If you want to know whether week-six problem sets were too hard, asking in week six beats asking in week twelve.

What changes when you sample the experience

Moving from one retrospective survey to repeated momentary measurement changes three things.

  1. Accuracy. You measure the experience closer to when it happened, before memory has compressed and distorted it. The peak-end and recency distortions shrink because you are no longer asking the remembering self to stand in for the experiencing self.

  2. Resolution. A single end-of-term number tells you a course was rated 3.6. A term-long signal tells you it was 4.2 until the assessment brief landed in week seven, then fell — which is an actionable finding, not a verdict. This is the difference between knowing a patient's average temperature and seeing the chart.

  3. Timing of action. The most damning fact about the end-of-term survey is that it arrives too late to help the students who completed it. Experience sampling is formative by construction: it surfaces problems while the cohort is still in the room, which is the whole point of the quality-assurance loop (see our piece on the course-evaluation action gap).

"But doesn't this just create survey fatigue?"

This is the strongest and most honest objection, and it has real teeth. Asking students to respond repeatedly risks the very survey fatigue and over-evaluation that already depresses response rates. Three points in response.

First, frequency is a design variable, not a fixed cost. EMA research is explicit that burden must be managed — short prompts, sensible sampling schedules, and a clear payoff for the respondent. A well-designed momentary check is two questions, not twenty; the goal is more occasions, not more items per occasion.

Second, relevance reduces fatigue. Fatigue is driven less by frequency than by the feeling that responses vanish into a void. Momentary feedback that visibly changes the course — "you said the pace was too fast, so here is an extra worked example" — is experienced as participation, not burden. Closing the loop is the antidote.

Second objection worth naming: momentary reports can be noisy and reactive. A single bad day distorts a single measurement, and being asked to reflect can itself change behaviour (reactivity). The answer is that momentary sampling is meant to be aggregated into a trajectory, where idiosyncratic noise averages out, while still preserving the time-structure that a one-shot survey destroys. The trajectory is the signal; any single ping is not.

None of this makes experience sampling a free lunch. It is a trade: you accept design complexity and respondent-burden management in exchange for less recall bias and earlier, more actionable signal. For most high-value courses, that trade is worth making.

Where Koji fits

Experience sampling has always been methodologically attractive and operationally painful — coordinating repeated, short, well-timed touchpoints across hundreds of students, then making sense of the resulting open-text stream, was simply too much work for a paper-and-Likert evaluation office. AI changes that calculus.

Koji for Education supports formative, mid-cycle collection rather than a single terminal survey, so feedback can be gathered at the points in the term that matter — after a major assessment, midway through a difficult module — instead of only at the end. Its AI-moderated conversational interviews keep each touchpoint short but deep: a two-minute exchange that probes the why in the student's own words, captured close to the experience rather than reconstructed from memory months later. Automatic thematic analysis turns a term's worth of momentary, open-text responses into a readable trajectory — showing when and why sentiment moved — and quality scoring filters the noise so a single off-day does not masquerade as a trend. Because the moderation is standardized and bias-aware, the repeated measurements are comparable over time, which is what makes a trajectory trustworthy.

The same conversational engine runs continuous and pulse research on the main Koji platform, where product and UX teams have long known that in-the-moment feedback beats the quarterly retrospective. Course evaluation is catching up to a method other fields already trust.

To be precise about the claim: experience sampling reduces recall bias and surfaces problems earlier. It does not eliminate bias, and it introduces its own design challenges. Used well — short prompts, sensible cadence, visible action — it is a strict improvement on asking students to remember a semester in the final week.

The takeaway

The end-of-term survey is built on a memory it cannot trust. Experience sampling replaces one distorted retrospective summary with many small, well-timed, in-context measurements, trading some design complexity for less recall bias, higher resolution, and feedback that arrives in time to matter. For institutions serious about acting on the student voice rather than merely archiving it, sampling the experience — not interrogating the memory of it — is the methodologically sounder path.

Want to collect course feedback while the term is still running, not after it's too late to help? See how Koji for Education does formative, in-the-moment evaluation.