New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Graduate outcomes9 min read

Do Graduates Rate Their Courses Differently Years Later? What Delayed Evaluation Can and Can't Fix

It is tempting to think alumni, with hindsight and a career behind them, would rate their courses more wisely than final-year students. The evidence says the ratings barely move — which tells you what delayed evaluation is not for, and what it is uniquely good at.

Koji Education Team

Product · July 12, 2026

Bottom line up front: A popular hope in quality assurance is that if you wait — ask graduates a year, five years, a decade after they leave — you will get wiser, less biased course ratings than the end-of-term rush produces. The research is clear and slightly deflating: overall ratings are remarkably stable over time. Alumni and current students largely agree on who taught well. So delayed evaluation is not a correction for the biases in end-of-course scores. But that stability is exactly why the new thing graduates can tell you — what actually transferred to work and life — is so valuable. Collect retrospective feedback for applied usefulness, not as a validity fix.

The intuition — and why it is mostly wrong

The intuition is seductive. Final-year students rate a course while sleep-deprived, grade-anxious, and unable to know which parts will matter. A graduate three years into a career has perspective: they know which modules turned out to be foundational and which were forgettable. Surely their judgment is better calibrated?

On overall teaching quality, the evidence says it barely differs. Philip Overall and Herbert Marsh's longitudinal study (1980, Students' evaluations of instruction: A longitudinal study of their stability) compared the same students' end-of-course ratings with their ratings of the same courses at least a year later and found substantial agreement — the stability-based reliability was actually higher than conventional internal-consistency estimates. Howard, Conway and Maxwell (1985) found high positive correlations between current-student and alumni ratings of instructor effectiveness: students and alumni of five years' standing largely agreed on who had been effective. Marsh's broader programme of work — including a 13-year stability analysis across thousands of classes — concluded that student ratings are, under appropriate conditions, multidimensional, reliable, and stable over time (Marsh & Roche, 1997).

The implication is uncomfortable for the "wait and they will see clearly" hypothesis: if a course scored poorly at the end of term, it will most likely score poorly from alumni too — and if it was inflated by charm or leniency, hindsight does not reliably deflate it. Delayed evaluation does not launder out end-of-course bias. The ranking mostly survives.

So is there any point asking graduates? Yes — a different question

Here is the pivot, and it is where the honest reading gets constructive. Stability of overall ratings does not mean graduates have nothing new to say. It means you should stop asking them the same question and start asking the one only they can answer: what actually transferred?

End-of-course evaluation is trapped at the lowest rung of Kirkpatrick's hierarchy — reaction (see Beyond Reaction). A final-year student can tell you whether a course felt useful; they cannot tell you whether the statistics module actually held up when they had to run an analysis at work, or whether the "employability" content survived contact with a real job. That is a question about transfer of training — and transfer, especially far transfer to a different context, can only be observed after the context arrives. Graduates are the sole witnesses to it.

This reframes the value proposition. Alumni feedback is weak as a re-rating of teaching quality (stable, redundant with end-of-course scores) but strong as a lagging indicator of applied value (unique, unobtainable any other way). It is the natural complement to destination data from graduate tracer studies: tracer studies tell you where graduates ended up; retrospective course feedback tells you which parts of the education they credit for getting there — and which they now see as gaps.

The counterargument to the counterargument: memory is not a clean instrument

We should not oversell retrospective feedback either, and integrity requires naming its own weakness. Human memory of an educational experience is reconstructive and subject to well-documented distortions — the peak-end rule and hindsight bias among them. A graduate who is now thriving may over-credit their alma mater; one who struggled may under-credit it. Career outcomes contaminate recollection: success feels like it must have come from somewhere, and the degree is a convenient somewhere. And there is a lag problem — by the time alumni can report on transfer, the course they are describing may have been redesigned twice, making the feedback a verdict on a version that no longer exists.

None of this makes retrospective feedback useless. It makes it specific: valuable for surfacing applied usefulness and curricular gaps at the programme level, unreliable for re-litigating an individual lecturer's performance or for real-time course correction. Knowing which job each tool is for is the whole discipline.

What defensible practice looks like

  1. Don't use alumni surveys to re-rate teaching. The ratings won't move enough to justify it, and you already have that signal from end-of-course evaluation.
  2. Do use them to measure transfer and programme-level value — which skills held up, which felt missing, what graduates would tell their first-year selves to take seriously.
  3. Ask at the programme, not module, level. Memory of individual modules decays; memory of what the whole education did for you is more robust and more relevant to accreditation and curriculum review.
  4. Triangulate with destination data, so self-report is anchored against what graduates actually went on to do.

Retrospective feedback and accreditation

The programme level is also where retrospective feedback earns its regulatory keep. Accreditation and periodic programme review increasingly ask not just whether students were satisfied, but whether a programme produced the intended outcomes and prepared graduates for what came next — questions that end-of-module satisfaction cannot answer and that only people who have left can address. Graduate tracer data supplies the destinations; alumni course feedback supplies the interpretation — which parts of the curriculum graduates now credit, which competences they found missing when the workplace demanded them, and where the programme's own stated outcomes turned out to be aspirational rather than real.

Used this way, retrospective feedback becomes a curriculum-review instrument rather than a teacher-rating one, and that reframing dissolves most of its methodological problems. The hindsight and peak-end distortions that would make an individual lecturer's delayed rating unreliable matter far less when the unit of analysis is a whole programme and the question is "what held up in practice." Aggregated across a cohort and triangulated against destination data, even imperfect individual recollections converge on patterns a programme committee can trust: the recurring gap, the consistently-valued core, the module everyone remembers as decisive. That is evidence an accreditation panel recognises, and it is evidence the standard end-of-term form structurally cannot produce, because the form asks the wrong people at the wrong time about the wrong thing.

Where Koji fits

Retrospective, transfer-focused feedback is exactly the kind of evidence a static Likert form handles worst. "Rate the usefulness of your degree, 1–5" from an alumnus produces a stable, near-meaningless number; the value is entirely in the why, and legacy tools throw the why away.

Koji for Education is built for that why. Its AI-moderated conversational interviews can ask a graduate what they actually use from a programme, follow up on vague answers, and probe how a skill transferred to their work — the depth that turns "it was useful" into a specific, curriculum-relevant insight. Automatic thematic analysis then aggregates hundreds of alumni conversations into patterns a programme committee can act on: recurring skill gaps, consistently-credited modules, competences that graduates say the workplace demanded but the curriculum skipped. Because it runs as a formative, programme-level instrument rather than a summative teacher-rating exercise, it fits the job retrospective feedback is genuinely good for, and its closing-the-loop action tracking connects what graduates report back to curriculum revision. Koji does not pretend hindsight is unbiased — we have named the peak-end and hindsight problems above — but it captures the rich, probeable detail that lets you weigh those accounts intelligently instead of reducing them to a stable, uninformative mean.

The same conversational interview engine powers the main Koji platform, where longitudinal customer research faces the identical challenge: what people remember valuing, probed for the reasons behind it.

If your alumni survey is just an end-of-course form sent late, you are collecting the one thing hindsight leaves unchanged and missing the one thing only graduates can tell you. See how Koji captures graduate-level programme feedback.