Beyond Reaction: Why Course Evaluation Rarely Escapes Kirkpatrick's First Level
Most university course evaluation measures student reaction and treats it as evidence of teaching quality. Kirkpatrick's four-level model explains why that is the weakest possible inference — and how to climb toward learning, behaviour and results.
Koji Education Team
Product ·
Bottom line up front: Donald Kirkpatrick's four-level evaluation model — reaction, learning, behaviour, results — is the clearest diagnostic we have for what a course-evaluation system actually measures. Judged against it, the typical end-of-semester satisfaction survey never gets past Level 1. That matters because the empirical relationship between Level 1 reaction and Level 2 learning is close to zero. If your quality-assurance evidence stops at "students were satisfied," you are reporting the level the research says predicts learning least.
The model, briefly, and why it travels to higher education
Kirkpatrick first set out his four levels in a 1959 series of articles for the American Society for Training and Development, later consolidated in Evaluating Training Programs: The Four Levels (1994). It was built for corporate training, but its logic maps cleanly onto higher education:
- Level 1 — Reaction: Did participants find the experience favourable, engaging and relevant? In a university this is the end-of-module satisfaction survey: "The teaching on this module was good," rated 1–5.
- Level 2 — Learning: Did they acquire the intended knowledge, skills and attitudes? This is assessment of learning outcomes — exams, projects, competency demonstrations.
- Level 3 — Behaviour: Do they apply what they learned in a new setting? In higher education this is the placement, the capstone, the first graduate job.
- Level 4 — Results: What is the downstream impact — on employability, on the discipline, on society?
The model is hierarchical by design: each level is harder to measure and more meaningful than the one below it. Kirkpatrick's own framing was that organisations gravitate to Level 1 because it is cheap and immediate, and then quietly treat it as a proxy for the levels they never measure.
The inconvenient evidence: reaction barely predicts learning
This is where a PhD audience should sit up. The assumption underpinning satisfaction-based quality assurance — that happier students learned more — does not survive contact with the meta-analytic record.
In organisational psychology, Alliger and colleagues' 1997 meta-analysis in Personnel Psychology (34 studies, 115 correlations) found that affective reactions correlated with learning at roughly ρ = .02 to .07 — essentially nothing. Only "utility" reactions (whether participants judged the content useful) showed a meaningful relationship (ρ ≈ .26). Sitzmann and colleagues' later, larger meta-analysis reached a compatible conclusion: reactions are strongly tied to post-training motivation and self-efficacy, but weakly to actual learning.
Higher education's own literature points the same way. Uttl, White and Gonzalez's 2017 meta-analysis in Studies in Educational Evaluation re-analysed the multisection studies long used to defend student evaluations of teaching and found that once small-sample bias is accounted for, the SET–learning correlation is about r = .08 for rating averages, and near zero (r ≈ −.02) once prior ability is controlled. Their blunt summary: student rating averages explain at most 1% of the variance in how much students actually learn.
Read together, these literatures deliver one message. Level 1 reaction is a real and worthwhile construct — but it is a measure of experience and satisfaction, not a measure of learning. Treating a 4.2/5 satisfaction average as evidence of teaching effectiveness is a Level-1 inference dressed up as a Level-2 conclusion.
"But doesn't this just prove evaluation is pointless?" — the strongest counterargument
Critics of the Kirkpatrick framing make three fair objections, and an honest case has to meet them.
First: reaction still matters. It does. Disengaged, alienated students disenrol, disengage and stop showing up — and Level 1 is the right place to catch that. The problem is not measuring reaction; it is only measuring reaction and mislabelling it.
Second: Kirkpatrick is dated and the levels are not strictly causal. Also true. Holton (1996) and others argued the model is more a taxonomy than a validated causal chain — higher levels do not automatically follow from lower ones. We agree, and that is precisely why we use it as a diagnostic of what you measure, not as a promise that good reactions cause good results.
Third: Levels 2–4 are expensive and confounded. Learning gain is hard to isolate; graduate outcomes are shaped by the labour market, not just the curriculum. This is the real constraint. The answer is not to retreat to Level 1, but to gather better Level-1+ evidence — proximal indicators that are honestly labelled and triangulated — rather than pretending satisfaction is something it is not.
Climbing the levels without pretending you have eliminated confounds
For quality-assurance officers and programme directors, the practical move is to stop collecting a single Level-1 number and start collecting evidence that reaches toward Levels 2 and 3:
- Capture utility, not just affect. Alliger's finding is actionable: ask whether students found the course useful and applicable, not just whether they liked it. Utility judgements track learning better than satisfaction does.
- Probe self-reported learning carefully, and label it honestly. Self-rated learning is a perception, not a measurement — but a well-structured probe ("Describe something you can now do that you could not before this course") yields richer, more falsifiable evidence than a Likert item.
- Triangulate. Pair student feedback with assessment data, peer observation and, for Level 3, graduate and placement feedback. See our piece on triangulating teaching evaluation across multiple evidence sources.
- Measure learning gain where you can. Our argument for measuring learning gain rather than satisfaction is the Level-2 companion to this essay.
Where Koji fits
Koji for Education does not claim to magic away the confounds in Levels 2–4 — no instrument can. What it does is make the climb from Level 1 to richer evidence practical and standardised.
Instead of a single satisfaction Likert item, Koji runs AI-moderated conversational interviews that probe why a student responded as they did and what they can now do — eliciting the utility and applied-learning signals that the meta-analytic record says actually matter, rather than the affective signal that does not. Its automatic thematic analysis turns hundreds of open-text responses into structured themes, so a programme team can see whether students describe genuine capability gains (a Level-2/3 signal) or merely report enjoyment (Level 1). Six structured question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) let you mix a fast satisfaction pulse with deeper applied-learning probes in one instrument. And closing-the-loop action tracking means the evidence feeds decisions rather than a filing cabinet — the point Kirkpatrick made about Level 4 being where evaluation justifies itself.
Because the AI moderator is bias-aware and standardised, you avoid the human-moderator inconsistency that plagues qualitative follow-up at scale — the same conversational interview engine that powers customer and user research on the main Koji platform, applied to the classroom.
The honest framing: Koji helps you collect Level-1-plus evidence that reaches toward learning and behaviour. It surfaces and structures; it does not eliminate the confounds that make Levels 2–4 genuinely hard. That distinction is the whole point of taking Kirkpatrick seriously.
How to report levels honestly to a committee
The practical discipline that follows from Kirkpatrick is simple but rarely observed: label every figure you present with the level it actually measures. When a teaching-quality committee sees a 4.2/5, it should be captioned "Level 1, reaction" — not "teaching effectiveness." That single relabelling changes the conversation, because it forces the obvious next question: what is our Level-2 evidence?
A mature reporting template pairs each module's reaction score with at least one indicator from a higher level — pass rates and assessment performance for Level 2, placement or capstone evidence for Level 3 — and flags explicitly where higher-level data is absent. The gaps are themselves a finding. An institution that cannot produce any evidence above Level 1 for a flagship programme has learned something important about its own quality-assurance system. Naming the level is not pedantry; it is the difference between evidence and decoration.
The takeaway
Kirkpatrick's model is forty years old and imperfect, and it remains the sharpest question you can ask of your evaluation system: which level am I actually measuring, and which level am I claiming? For most European universities the honest answer is "we measure Level 1 and report it as Level 2." The fix is not more satisfaction surveys. It is evidence that climbs.
Ready to move your course evaluation beyond reaction? See how Koji for Education turns satisfaction surveys into structured, learning-aware evidence.