New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Graduate outcomes10 min read

Measuring Learning Gain, Not Satisfaction: What the UK Pilots Teach Course Evaluation

England spent over £4 million trying to measure how much students actually learn, and concluded there is no silver-bullet metric. That failure is the most useful thing course evaluation can learn from — because it shows exactly why satisfaction is the wrong proxy for value.

Koji for Education

Research & Editorial Team · June 24, 2026

Bottom line up front: Between 2015 and 2018, England ran the most ambitious attempt yet to measure learning gain — how much students actually develop over a degree — funding 13 pilot projects across more than 70 institutions with over £4 million. The headline conclusion of the final evaluation was sobering: there is no single, "silver-bullet" metric that measures learning gain comparably across subjects and contexts. For anyone responsible for course evaluation, this is not a counsel of despair — it is the clearest possible evidence that student satisfaction, the thing we routinely measure, is a poor proxy for the thing we actually care about: whether a programme develops graduates. This post draws out what the learning-gain experiment teaches evaluation practice, and how to act on it without waiting for a metric that may never arrive.

Satisfaction is not learning, and we have known it for a while

The unstated promise of most course evaluation is that a well-rated course is a good course — that satisfaction stands in for educational value. The evidence has been straining against that promise for years.

The decisive finding is that ratings barely track learning. Uttl, White & Gonzalez (2017), re-analysing the multisection studies, found the student-rating–learning correlation is essentially zero once prior ability is controlled. In the United States, Arum and Roksa's Academically Adrift (2011) used the Collegiate Learning Assessment to estimate that a large share of students showed no statistically significant gains in critical thinking and reasoning in their first two years — while satisfaction and confidence remained high. Students can feel they are learning, and report a course favourably, while objective gain is modest. That gap is exactly why the "feeling of learning" can move in the opposite direction to actual learning.

If satisfaction and learning come apart, then a quality system built only on satisfaction is optimising the wrong variable — and, through Goodhart's law, will tend to produce courses that are pleasant rather than demanding.

What the learning-gain pilots tried — and what they found

The English programme, funded by HEFCE and concluded under the Office for Students, tested three broad approaches to measuring gain: standardised tests of generic skills (critical thinking, reasoning), self-reported surveys of perceived development, and other proxies such as grades and engagement data. The intent was rigorous: to find a measure of value added that could sit alongside satisfaction in the sector's evidence base.

The final evaluation and the earlier annual reports reached a consistent verdict. Standardised tests struggled with student motivation (low-stakes tests yield low effort) and cross-disciplinary validity. Self-report measures were cheap and scalable but vulnerable to the same response-shift and social-desirability biases that plague satisfaction surveys. Grade-based proxies were confounded by inconsistent marking standards across institutions. The programme's honest summary was that no single metric captures learning gain comparably across all subjects and contexts.

That is not a failure of effort; it is a finding about the construct. Learning gain is multidimensional, discipline-specific, and entangled with prior attainment. Anyone selling you a tidy number for it is overpromising.

The lessons course evaluation should actually take

The pilots did not conclude "give up." They concluded "stop looking for one number, and triangulate instead." Four practical lessons follow.

1. Treat satisfaction as one signal, not the verdict. A favourable evaluation tells you students had a good experience. It does not tell you they developed. Keep measuring experience, but stop letting it stand in for value. (This is the same argument we make in why averaging a satisfaction score misleads.)

2. Ask about development, not just contentment. Self-report has limits, but well-designed self-report of specific skills — "I can now construct a counter-argument I couldn't before," "I can read a regression table" — is more informative than "I was satisfied." This is the move from satisfaction to capability that measuring skills rather than satisfaction is about, and it aligns evaluation with the graduate-outcomes and employability evidence that accreditation increasingly demands.

3. Triangulate across sources. Because no single measure works, the defensible approach combines experience data, self-reported development, assessment evidence, and — where available — graduate-outcome and employer signals. (More on triangulation.) Programme-level evaluation, not just course-level satisfaction, is where value added becomes visible — and where it connects to graduate tracer studies and employability.

4. Respect what self-report can and cannot do. Course evaluation cannot directly measure learning gain — and it shouldn't pretend to. What it can do is capture the proximal, perceived markers of development that, combined with harder evidence, make a credible case. We are deliberately careful about this boundary in why course evaluation cannot measure employability directly.

"But isn't self-reported learning gain just as biased as satisfaction?"

This is the sharpest objection, and the learning-gain pilots themselves raised it. Self-reported development suffers from response-shift bias (students recalibrate what "good" means as they learn, so a naive pre/post comparison understates gain), social desirability, and the same recency effects that distort satisfaction. If self-report is biased, why prefer "I developed" over "I was satisfied"?

Three responses. First, specificity beats globality. A vague "I learned a lot" is weak; "I can now do X, which I could not do in week one" is a concrete, checkable claim that a respondent finds harder to confabulate. Second, design can mitigate response shift — for example the retrospective pre-test ("then-now") method asks students to rate their prior ability after instruction, reducing recalibration error. Third, and most importantly, self-report is one leg of a triangle, not the whole structure. The pilots' conclusion was not "self-report is useless" but "no single source suffices." Perceived development earns its place precisely when it is cross-checked against assessment and outcome data — not when it replaces them.

So the objection is correct that self-reported gain is imperfect. It is wrong that this makes it equivalent to satisfaction. A specific, well-designed, triangulated measure of development is a meaningfully better proxy for value than a global contentment score — which is exactly why the sector spent £4 million trying to build one.

Where Koji fits

If the lesson is "ask about specific development, design against bias, and triangulate," then the instrument has to be capable of more than averaging a Likert item. That is the gap Koji for Education is built for.

Koji's AI-moderated conversational interviews can probe perceived development concretely — following up "I improved" with "improved at what, specifically, and how do you know?" — surfacing the checkable, skill-level claims that generic satisfaction items never reach. Its six structured question types (open-ended, scale, single-choice, multiple-choice, ranking, yes/no) support retrospective pre-test designs and skill self-assessments alongside experience measures. Automatic thematic analysis maps open-text responses onto skills and learning outcomes, turning narrative into programme-level evidence of development that can be triangulated with assessment and graduate-outcome data. And because Koji reports at programme and institution level, it captures value added where it actually shows up — across a degree, not a single lecture.

The same engine powers skills and outcomes research on the main Koji platform, where teams measure capability change rather than momentary satisfaction. Education is asking the same question every serious research function asks: did this actually develop the person?

To be precise: Koji does not measure learning gain in the standardised-test sense, and we would distrust any tool that claimed to. It captures the perceived, proximal markers of development — specifically and at scale — that belong in the triangulated evidence base the UK pilots said we need.

The takeaway

England spent over £4 million and concluded there is no single metric for how much students learn. The right response is not to keep pretending satisfaction is that metric. It is to demote satisfaction to one signal, ask students specifically about what they can now do, design against the biases of self-report, and triangulate with harder evidence. That is a more honest — and more useful — foundation for course evaluation than any average of a five-point scale.

Want course feedback that asks about development, not just satisfaction — and reports it at programme level? Explore Koji for Education.