New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods9 min read

Course Evaluations Measure Reaction, Not Learning: What Kirkpatrick's Four Levels Reveal

Kirkpatrick's four-level model explains why an end-of-term course evaluation is a Level-1 'reaction' measure — and why decades of meta-analytic evidence show reaction correlates almost nothing with actual learning.

Koji Education Team

Product

Answer box (BLUF)

A standard end-of-term course evaluation sits at Level 1 — "reaction" in Donald Kirkpatrick's four-level model of evaluation. It records whether students liked a course, not whether they learned from it. The best meta-analytic evidence — Alliger and Janak's re-examination of the model, and Alliger, Tannenbaum, Bennett, Traver and Shotland's 1997 meta-analysis — found that affective reaction correlates with immediate learning at only about r ≈ .02–.07. In plain terms: knowing how much students enjoyed a course tells you almost nothing about how much they learned. A defensible quality-assurance process treats satisfaction and learning as separate constructs, measured separately, and never uses one as a proxy for the other.

What the research says

The model. Kirkpatrick's framework, first sketched in a series of 1959 articles and consolidated in Evaluating Training Programs (1994), proposes four ascending levels of evaluation evidence: (1) Reaction — did participants like it?; (2) Learning — did knowledge, skills or attitudes change?; (3) Behavior — did they apply it?; (4) Results — did outcomes improve? Kirkpatrick's original presentation implied the levels are causally linked and positively intercorrelated, so that a good reaction is a leading indicator of learning.

The empirical test. Alliger and Janak (1989), in Personnel Psychology, subjected those assumptions to scrutiny and found them largely unsupported. They identified three "problematic assumptions": that higher levels yield more information, that the levels are causally linked, and that they are positively intercorrelated. The intercorrelation assumption is the one that matters most for course evaluation — and it failed. The follow-up meta-analysis by Alliger, Tannenbaum, Bennett, Traver and Shotland (1997), pooling correlations across many studies, reported a correlation of roughly .02 between affective reactions (whether people enjoyed the experience) and immediate learning. Even "utility" reactions — judgements that the material was useful — correlated only modestly with learning (around .26). Reaction and learning are, empirically, near-orthogonal.

Corroborating and contrasting evidence. This is not a lone finding. Sitzmann, Brown, Casper, Ely and Zimmerman (2008) meta-analysed self-assessments of knowledge and found that self-reported learning is only weakly related to objectively measured learning (correlations around .34 for post-training self-assessment, and much lower once cognitive load and motivation are partialled out) — students are poor judges of their own knowledge gains. In higher education specifically, Uttl, White and Gonzalez (2017) re-analysed the multisection validity literature and concluded that when studies are properly weighted, student ratings are essentially unrelated to learning — a direct echo of the Kirkpatrick-level problem inside the SET literature. Contrastingly, some Level-1 defenders (e.g. Brown, 2005) argue that well-designed utility-focused reaction items carry more signal than generic "I enjoyed it" items — a nuance we return to below.

Why it matters for course evaluation in practice

Almost every institutional course-evaluation instrument is, structurally, a Level-1 reaction survey: "The lecturer explained concepts clearly," "I would recommend this course," "Overall satisfaction." These items measure the student's affective and perceptual response. The Kirkpatrick evidence tells a quality-assurance office three uncomfortable things:

  1. You cannot infer learning from satisfaction. A course with a 4.6/5 mean is not, on that basis, a course where students learned more than one at 4.1/5. Ranking programmes or instructors on reaction scores and calling it a measure of educational quality is a category error.
  2. The levels require different instruments. Learning (Level 2) is measured with assessment data, pre/post tests, or validated learning-gains instruments — not with a satisfaction Likert scale. Behavior (Level 3) needs follow-up: are graduates applying the competency? Results (Level 4) needs programme-outcome and employability data. A single end-of-term survey cannot reach any level above 1.
  3. Reaction data still has value — for the right question. Level-1 data is a legitimate and important measure of student experience: engagement, belonging, perceived support, workload tolerability, whether the course felt fair. Those are real quality dimensions worth protecting. The error is not collecting reaction data; it is relabelling it as evidence of learning.

For accreditation, this maps directly onto the distinction European frameworks draw between student satisfaction evidence and learning-outcome achievement evidence. An ENQA/ESG-aligned self-evaluation that leans only on satisfaction scores to demonstrate that programme learning outcomes are met is making exactly the inference Alliger and Janak showed is unwarranted.

Limitations and honest caveats

A rigorous reader should push back on several points, and we should concede them:

  • The evidence is largely from corporate training, not university courses. Kirkpatrick's model originates in workplace training evaluation. The near-zero reaction–learning correlation is best established there; higher-education replications (via the SET multisection literature) point the same way but with more heterogeneity and fierce methodological dispute (Uttl et al. vs. defenders of the earlier Cohen synthesis).
  • "Learning" is measured imperfectly too. Immediate post-test learning is itself a narrow criterion; it can miss durable, transferable, or delayed learning. A low reaction–learning correlation partly reflects the crudeness of both measures, not just the uselessness of reaction.
  • Reaction is not worthless. Utility-type reactions correlate more strongly with learning than affective ones, and reaction data predicts persistence and re-enrolment, which matter institutionally. The claim is precise: reaction is a weak proxy for learning, not a meaningless number.
  • Causal direction is unsettled. Some argue enjoyment can enable learning (motivation, attendance) even if the raw correlation is low because of confounds. The model's failure is about naïve intercorrelation assumptions, not a proof that experience is irrelevant to learning.

Stating these caveats is not hedging — it is the honest position. The defensible conclusion is modest and robust: do not treat a reaction score as a learning score.

How Koji incorporates this

Koji for Education is built around the premise that a single Likert reaction number is thin evidence, so the platform is designed to widen the aperture rather than pretend a satisfaction score measures learning:

  • AI-moderated conversational interviews probe the "why" behind the number. Where a Likert item captures a Level-1 reaction, Koji's conversational evaluation follows up — "You said the assessment felt unclear; what specifically were you unsure how to do?" — surfacing the process signals (clarity of tasks, alignment of assessment to teaching) that are more diagnostically useful than an affective rating. This does not turn reaction into learning data, but it makes the reaction data actionable.
  • Structured question types keep constructs separate. Koji supports scale, single_choice, multiple_choice, ranking, yes_no, and open_ended items, so an institution can deliberately design a reaction block distinct from any self-reported learning-gains block, and report them as different things — never averaging enjoyment and perceived learning into one "quality" figure.
  • Retrospective and self-reported learning items are flagged as self-report. Because Koji's reporting is bias-aware, self-assessed learning gains are presented as perception data (with the Sitzmann caveat) rather than as objective learning evidence, and the platform encourages triangulation with assessment outcomes held elsewhere.
  • Mid-cycle / formative collection targets the levels reaction can inform. Early-term conversational check-ins are designed to catch fixable experience problems (pace, workload, unclear expectations) — the legitimate use of Level-1 data — rather than to serve as a summative verdict on teaching quality.
  • Closing-the-loop action tracking records what changed in response, so reaction evidence feeds improvement (its proper job) instead of being over-interpreted as a learning metric.

Koji is designed to mitigate the misuse of reaction data, not to eliminate the underlying limitation — no instrument that asks students how a course felt can, by itself, measure how much they learned. Teams running broader research beyond the classroom can apply the same AI-moderated interview engine to product and customer research through Koji's core platform at koji.so.

Related Resources

Putting it into practice: a two-track evaluation

The cleanest operational response to the Kirkpatrick evidence is to stop running one instrument that pretends to do two jobs, and instead run two explicit tracks that are analysed and reported separately.

Track 1 — Experience (Level 1, done honestly). A short, well-designed reaction instrument that asks about the things students are genuinely positioned to judge: whether expectations were clear, whether the workload was manageable, whether they felt supported and included, whether assessment felt fair. These are valid, decision-relevant experience measures. Reported as experience, they are unimpeachable; the trouble only ever comes from relabelling them as evidence of learning.

Track 2 — Learning (Levels 2–4, with the right instruments). Achievement data, pre/post gains on validated instruments, progression and — for professional programmes — evidence that graduates apply the competency. This track owns the "did they learn?" question. It is slower, harder and more expensive, which is exactly why institutions are tempted to let a satisfaction mean stand in for it. The Kirkpatrick evidence is the argument for resisting that temptation.

When a review panel or accreditor asks "how do you know this programme is effective?", the two-track structure lets you answer without conflation: here is what students experienced, and separately, here is the evidence they learned. The near-zero reaction–learning correlation is not a reason to collect less feedback — it is a reason to be precise about which question each number answers, and never to let a friendly Level-1 average quietly do a job it was never able to do.

References

  • Kirkpatrick, D. L. (1994). Evaluating Training Programs: The Four Levels. Berrett-Koehler.
  • Alliger, G. M., & Janak, E. A. (1989). Kirkpatrick's levels of training criteria: Thirty years later. Personnel Psychology, 42(2), 331–342. https://doi.org/10.1111/j.1744-6570.1989.tb00661.x
  • Alliger, G. M., Tannenbaum, S. I., Bennett, W., Traver, H., & Shotland, A. (1997). A meta-analysis of the relations among training criteria. Personnel Psychology, 50(2), 341–358. https://doi.org/10.1111/j.1744-6570.1997.tb00911.x
  • Sitzmann, T., Ely, K., Brown, K. G., & Bauer, K. N. (2010). Self-assessment of knowledge: A cognitive learning or affective measure? Academy of Management Learning & Education, 9(2), 169–191. https://doi.org/10.5465/amle.9.2.zqr169
  • Uttl, B., White, C. A., & Gonzalez, D. W. (2017). Meta-analysis of faculty's teaching effectiveness: Student evaluation of teaching ratings and student learning are not related. Studies in Educational Evaluation, 54, 22–42. https://doi.org/10.1016/j.stueduc.2016.08.007