What the UK's National Student Survey Reform Teaches the Rest of Europe
In 2023 the UK regulator did something most quality systems still flinch at: it deleted the overall-satisfaction question. The reasoning behind that decision is the most useful thing to happen to course evaluation in a decade — and it applies far beyond Britain.
Koji for Education
Research & Editorial Team ·
Bottom line: In 2023 the Office for Students removed the single "overall satisfaction" question from England's National Student Survey, arguing it was too consumerist and distracted from more diagnostic measures. The decision was contested and imperfect, but the underlying critique is sound and exportable: a summative satisfaction score is a popularity signal, not a quality measure, and optimising for it can actively distort teaching. Every European institution still anchoring its course evaluation on "overall satisfaction" should read the NSS reform as a warning about its own instrument.
A regulator blinked at its own headline number
The National Student Survey is one of the largest student-feedback exercises in the world: a census of final-year undergraduates across the UK, running since 2005, with results that feed league tables and recruitment. Its most cited output was always Question 27 — overall satisfaction with the course.
In 2023, after consultation, the Office for Students removed that overall-satisfaction question for providers in England, moved to more direct question wording and a four-point scale, and added items on mental-wellbeing services and — in England — freedom of expression. The stated rationale was striking for a regulator: the satisfaction question was judged too consumerist in nature, and to detract from the survey's more useful, diagnostic findings.
This was not universally welcomed. The Quality Assurance Agency warned it removed the only item that directly asked about overall course quality, and the large majority of consultation respondents opposed the change. Those objections are real and we will take them seriously below. But strip away the politics and the core methodological claim is one the research has supported for years.
Why "overall satisfaction" is the wrong anchor
Three problems sit inside that one tidy number.
It measures contentment, not learning. Satisfaction correlates with things institutions should not be optimising — entertainment value, easy grading, light workload — and barely correlates with how much students actually learn. The most rigorous meta-analysis of the multisection literature, Uttl, White and Gonzalez (2017), found the relationship between average ratings and learning is essentially zero once small-sample bias is accounted for. A course can be deeply satisfying and pedagogically thin, or demanding, uncomfortable and excellent. A satisfaction anchor cannot tell these apart.
It invites Goodhart's law. When a measure becomes a target, it stops being a good measure. Make overall satisfaction the headline that drives league tables and funding, and you create a direct incentive to raise it by the cheapest available means — lighter assessment, grade leniency, crowd-pleasing content — rather than by better teaching. The OfS's "too consumerist" framing is really a Goodhart argument: the number had become a thing to be gamed.
It is a single global judgement — the most bias-exposed quantity you can collect. A lone overall rating is exactly where halo effects, recency, and stereotype-driven bias concentrate, because it asks for an undifferentiated gut impression rather than a specific, checkable observation. Reducing a 12-week intellectual experience to one number on one scale throws away nearly all the diagnostic information and keeps the part most vulnerable to distortion.
The half-lesson and the full lesson
The NSS reform teaches a real lesson, but the OfS only learned half of it.
The half-lesson it got right: do not anchor a quality system on a summative popularity score. Removing Question 27 and moving toward more direct, specific items is, methodologically, a step toward measuring something more diagnostic.
The full lesson it missed: deleting the bad question is not the same as collecting good evidence. Replacing one global Likert item with a handful of more specific Likert items is still a static, low-bandwidth instrument. It still cannot ask a follow-up. It still cannot tell you why "assessment and feedback" scored low, or what students would change, or whether a complaint is widely shared or one loud voice. The deeper failure of satisfaction surveys is not the word "satisfaction" — it is the survey form itself, which can only ever harvest pre-coded reactions to questions written before anyone knew what the cohort would actually care about.
"But students deserve a say in overall quality — and you can't run league tables on themes"
This is the strongest objection, and it is the QAA's: removing the overall-quality question disenfranchises prospective students who used it to compare courses, and replaces a clean comparable metric with something messier.
It deserves a fair answer. Yes, prospective students need comparable information, and yes, a single number is convenient. But convenience is precisely the trap. A comparable-but-invalid number is worse than no number, because it launders a popularity signal into an apparently authoritative quality signal and sends applicants toward the most satisfying courses rather than the best-taught ones. The honest response to "students deserve a say in overall quality" is not to keep a question that does not measure quality — it is to build feedback that actually captures quality and can still be summarised for comparison. That is a harder engineering problem, not an argument for the old number.
A second fair objection: isn't dropping satisfaction just a way for providers to dodge accountability? It can be read that way, which is why the reform was contested. But accountability built on a gameable metric is fragile accountability. A measure that rewards grade inflation is not holding anyone to account for teaching; it is holding them to account for being agreeable.
What "after satisfaction" should actually look like
If the lesson is "stop anchoring on a global satisfaction score," the constructive question is what replaces it. Not nothing — and not just more Likert items. The replacement has to recover the diagnostic bandwidth a survey throws away.
This is the gap a conversational, AI-moderated approach is built for. Instead of asking every student to compress a semester into one number, Koji for Education runs a standardised AI-moderated interview that adapts to each student — probing why a view is held, asking the follow-up a paper form cannot, and using six structured question types (open-ended, scale, single- and multiple-choice, ranking, yes/no) so responses are both specific and analysable. Automatic thematic analysis then turns thousands of conversations into ranked, prevalence-weighted themes you can summarise and compare across cohorts — recovering the comparability the QAA rightly wanted, but built on what students actually said rather than how satisfied they declared themselves. Because the AI moderator is consistent and bias-aware across every respondent, it avoids both the human-moderator inconsistency of focus groups and the halo-prone single-number problem of surveys, and it does so with GDPR/AVG-compliant handling suited to European regulation. It does not pretend to eliminate bias or to make quality a single tidy digit — it surfaces and mitigates the former and replaces the false tidiness of the latter with something you can act on.
The same interview engine underpins the main Koji platform for general user research, where "our NPS is fine but we don't know why" is the identical failure mode dressed in commercial clothes.
The reframe for European quality offices
You do not need a UK regulator to make this decision for you. If your course evaluation still opens — or worse, ends — with "Overall, how satisfied were you with this module?", the NSS reform is a free natural experiment telling you what that number is and is not worth. The lesson is not merely "delete the satisfaction question." It is: stop asking a single static score to measure quality at all, and start collecting feedback rich enough to tell you what to actually do.
Rethinking what your evaluation anchors on? See how Koji for Education works.