New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods11 min read

Stop Waiting for the End of Term: Experience Sampling for In-the-Moment Course Feedback

Experience-sampling methods capture what students feel and think during a course, not their reconstructed memory of it months later. Here is why in-the-moment data can be more valid than the end-of-term survey — and how to use it responsibly.

Koji Education Team

Product

The short answer

The end-of-term survey asks students to remember a whole course after it is over — and memory is not a neutral recording of experience. Experience-sampling methods (ESM), also called ecological momentary assessment (EMA), instead capture short reports of students' feelings, thoughts, and engagement as the course is happening, in the settings where learning actually occurs. Because momentary reports sidestep the memory distortions and peak-end weighting that contaminate retrospective ratings, ESM can yield a more valid picture of the lived course experience — and, uniquely, it can catch problems while there is still time to fix them. It is not a replacement for summative evaluation, but it fills a blind spot that the single end-of-term instrument cannot.

What the research says

Zirkel, Garcia and Murphy's 2015 methodological review in Educational Researcher ("Experience-sampling research methods and their potential for education research," 44(1), 7–16) is the clearest case for bringing ESM into education. ESM, they explain, lets researchers "learn about individuals' lives in context by measuring participants' feelings, thoughts, actions, context, and/or activities as they go about their daily lives." By capturing experience "in the moment and with repeated measures," ESM expands the questions education research can ask — how engagement, interest, and effort actually rise and fall across a term — rather than collapsing a semester into one remembered summary judgment.

The method is not new or fringe. Csikszentmihalyi and Larson's foundational 1987 paper in the Journal of Nervous and Mental Disease ("Validity and reliability of the experience-sampling method") established ESM as a psychometrically defensible instrument, presenting evidence for its short- and long-term reliability and for its validity through correlations with physiological measures, one-time psychological tests, and behavioural indices. Decades of subsequent work — much of it summarised in Hektner, Schmidt and Csikszentmihalyi's Experience Sampling Method: Measuring the Quality of Everyday Life (2007) — has refined how to sample moments (signal-contingent, interval-contingent, or event-contingent prompts) and how to model the resulting nested, longitudinal data.

Why does the timing matter so much? Because momentary and retrospective reports can diverge sharply, and the gap is systematic rather than random. Goetz, Bieg, Lüdtke, Pekrun and Hall's 2013 study in Psychological Science ("Do girls really experience more anxiety in mathematics?") is a striking demonstration: on trait (retrospective, habitual) measures, female students reported markedly higher maths anxiety than male students — but when the same emotion was captured in the moment with experience sampling during actual maths classes and tests, the gender difference disappeared. The retrospective self-report reflected students' beliefs about their competence, not their real-time emotional experience. The lesson for course evaluation is direct: what students remember feeling about a course, filtered through their self-concept and the peak-end rule, can differ substantially from what they actually felt while enrolled. If you only ever ask at the end, you measure the reconstruction, not the experience.

Why it matters for course evaluation in practice

The dominant course-evaluation design — one questionnaire in the final week — is retrospective by construction, and it inherits every known weakness of retrospective self-report: the peak-end rule (the most intense moment and the final moment dominate the summary), recency effects (the last topic or a single bad exam colours the whole rating), and reconstruction biased by the student's self-concept and current mood. ESM attacks the root cause by moving measurement inside the term.

Three practical payoffs follow. First, validity: repeated momentary reports are less distorted by memory, so trends in engagement and difficulty reflect the course as lived. Second, granularity: instead of one number for "the course," you see which weeks, topics, or activities drove engagement up or down — a map, not a verdict. Third, and most valuable for quality assurance, timeliness: an ESM signal that engagement collapsed in week five is actionable this term, for these students, not a post-mortem finding delivered after they have moved on. This is the decisive advantage over even a well-run mid-semester survey: ESM samples repeatedly and lightly, tracking the trajectory rather than taking a single interior snapshot.

Limitations and honest caveats

Participant burden and reactivity are real. Repeatedly prompting students risks fatigue, dropout, and — because you are asking them to attend to their experience — reactivity, where the act of measuring changes the thing measured. Poorly designed ESM can annoy students into non-response or nudge them to think about a course differently than they otherwise would. Prompt frequency must be deliberately light.

Compliance and sampling bias. ESM's validity depends on students actually responding to prompts, and non-response is rarely random — the disengaged students whose data you most need are the likeliest to ignore a prompt. Missed prompts can bias the picture toward the conscientious, exactly as low response rates do in conventional surveys.

Momentary is not automatically 'truer' for every purpose. Retrospective summary judgments are not merely errors; for some decisions (an overall verdict students would stand behind, a considered reflection) the reconstruction is the relevant construct. The Goetz finding shows momentary and retrospective measures answer different questions; ESM complements the end-of-term survey rather than dethroning it.

Analytic and ethical complexity. ESM produces nested, autocorrelated, longitudinal data that require multilevel modelling to analyse properly — a raw average across prompts is misleading. And repeated in-context prompting raises heightened privacy and consent obligations, especially where responses could be linked to a small class or an identifiable student, engaging GDPR data-minimisation duties.

How Koji incorporates this

Koji's conversational, mid-cycle architecture is well suited to lightweight in-semester sampling that keeps burden low and turns momentary signals into action.

  • Light, event-contingent conversational check-ins. Rather than a single long end-of-term form, Koji can run short AI-moderated check-ins tied to points in the course — after a key assessment, a difficult module, or a specific week — capturing feeling and engagement close to the experience while keeping each touch brief to limit fatigue and reactivity.
  • Momentary depth through conversation. A short check-in that flags falling engagement can trigger an in-the-moment probe — "What happened this week that made it harder to stay engaged?" — capturing the cause while the memory is fresh, not reconstructed months later through the peak-end filter.
  • Trajectory reporting over verdicts. Koji's analysis is built to show how engagement, difficulty, and sentiment move across the term rather than collapsing them into one end-of-term number, giving a programme team the week-by-week map ESM is prized for.
  • Formative, act-now timing. Because the signal arrives during the course, closing-the-loop action tracking can record an intervention made this term and test whether the next check-in shows recovery — the timeliness payoff that a retrospective survey structurally cannot offer.
  • Burden- and bias-aware design. Koji supports deliberately low prompt frequency to protect response quality, and its reporting can flag when in-term responses are skewing toward the most engaged students, mitigating (never eliminating) the compliance-bias problem.
  • Privacy safeguards. In-context sampling in small cohorts raises re-identification risk, so Koji's anonymity and data-minimisation protections apply to momentary data as they do to summative surveys.

Koji is designed to complement the summative end-of-term evaluation with a valid in-semester signal, not to replace considered retrospective judgment. The same AI-moderated interview engine drives Koji's core research platform at koji.so, where in-the-moment customer and product feedback follows the identical experience-sampling logic.

Related Resources

References

  • Zirkel, S., Garcia, J. A., & Murphy, M. C. (2015). Experience-sampling research methods and their potential for education research. Educational Researcher, 44(1), 7–16. https://doi.org/10.3102/0013189X14566879
  • Csikszentmihalyi, M., & Larson, R. (1987). Validity and reliability of the experience-sampling method. Journal of Nervous and Mental Disease, 175(9), 526–536. https://doi.org/10.1097/00005053-198709000-00004
  • Goetz, T., Bieg, M., Lüdtke, O., Pekrun, R., & Hall, N. C. (2013). Do girls really experience more anxiety in mathematics? Psychological Science, 24(10), 2079–2087. https://doi.org/10.1177/0956797613486989
  • Hektner, J. M., Schmidt, J. A., & Csikszentmihalyi, M. (2007). Experience Sampling Method: Measuring the Quality of Everyday Life. Thousand Oaks, CA: Sage.