New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Sector trends8 min read

Your Course Evaluations Now Feed the League Tables. That Quietly Changes What They Measure.

Student-satisfaction data collected to improve teaching is repurposed into public league tables that shape applications and budgets. The moment a diagnostic number becomes a ranking input, it stops being safe to read at face value — and the incentives around it change.

Koji Education Team

Product ·

Bottom line up front: Course-evaluation and student-satisfaction data are increasingly repurposed as inputs to public university rankings and league tables. The evidence shows those rankings measurably move applications — so the satisfaction number stops being a private diagnostic and becomes a high-stakes competitive metric. That transformation, not the survey instrument itself, is what most distorts course evaluation. If you want feedback that still tells you the truth, you have to keep the improvement channel separate from the ranking channel.

From diagnostic to scoreboard

Most course evaluation was designed to answer a local, formative question: how is this module going, and what should we change? But the aggregate of that data rarely stays local. In the UK, responses to the National Student Survey (NSS) are folded directly into the major public league tables. The Complete University Guide, for example, builds its "student satisfaction" measure by averaging responses to the survey's teaching-and-feedback questions, adjusting for subject mix, and then dividing the result — an explicit attempt to stop satisfaction from dominating the overall rank. The Guardian University Guide weights student experience even more heavily. Pan-European tools like U-Multirank likewise lean on self-reported student-satisfaction indicators.

The point is not that any single methodology is wrong. It is that a number gathered to improve a course is now doing a second, very different job: ranking the institution against its competitors.

Rankings are not passive — they move behaviour

It is tempting to dismiss league tables as vanity metrics that no one acts on. The evidence says otherwise. In the most rigorous study of the question, Gibbons, Neumayer and Perkins analysed NSS scores and undergraduate applications across British universities and found that satisfaction scores have a small but statistically significant effect on applications — and that the effect operates primarily through a university's position in the published league tables, not through the raw score itself (Gibbons, Neumayer & Perkins, 2015, LSE/CEP). The impact was strongest for more able applicants and for mid-to-upper-tier universities facing the most competition.

So the causal chain is real: satisfaction → league-table position → applications → tuition revenue and prestige. Once that chain exists, the satisfaction number is no longer a neutral thermometer. It is a lever attached to the institution's budget.

Goodhart's law, at the scale of a whole sector

When a measure becomes a target, it stops being a good measure. That is Goodhart's law, and course evaluation has always been exposed to it at the level of the individual lecturer — the temptation to teach to the survey, grade leniently, or bring cookies on evaluation day. (We have written before about how the measure quietly becomes the target.)

Rankings raise the same dynamic to the level of the whole institution. When a fraction of a satisfaction point can shift a university several places in a table, the rational institutional response is to manage the number, not just the teaching behind it: run "boost the NSS" campaigns in the weeks before the survey, coach students on how their answers affect the university's reputation, quietly discourage negative framing. None of this improves a single seminar. It improves the score. And every euro and hour spent defending the ranking is diverted from the closing-the-loop work that would actually make the course better.

There is a deeper measurement problem underneath the incentive one. Satisfaction is a weak proxy for learning. Uttl, White and Gonzalez's meta-analysis of the multisection literature found that, once prior ability is accounted for, the correlation between student-evaluation ratings and how much students actually learn is effectively zero (Uttl, White & Gonzalez, 2017, Studies in Educational Evaluation). Ranking universities on satisfaction therefore risks sorting them on something that is only loosely connected to educational quality — and then letting that ranking reshape where students apply.

"But doesn't public reporting improve accountability?"

This is the strongest counterargument, and it deserves a fair hearing. Transparency has genuine value. The European Standards and Guidelines (ESG) explicitly expect institutions to publish information about their programmes, and prospective students have a right to comparable data. Hiding satisfaction results would be its own kind of failure. Accountability pressure has, in places, pushed institutions to take feedback more seriously than they otherwise would.

Three qualifications keep that argument honest. First, accountability and improvement are different jobs, and the same instrument struggles to serve both — a tension long documented in the programme-evaluation literature and one we treat at length in our piece on the dual-purpose problem. A number that can end up in a league table is a number students and staff learn to manage rather than answer honestly. Second, ranking on a single, averaged satisfaction score is precisely the practice the Leiden Manifesto and the Metric Tide warn against: measure with a portfolio, account for context, and never let one indicator carry decisions it cannot bear. Third, transparency does not require reducing a programme to one comparable digit; it requires giving prospective students rich, honest information — which a mean rating conspicuously fails to do.

So the response to "rankings improve accountability" is not to abolish public reporting. It is to stop pretending that one satisfaction average, ranked to two decimal places, is either a fair or an informative basis for it.

What to do instead: separate the channels

The practical fix is structural, not technical. Keep two clearly distinct feedback channels and never let the second one contaminate the first.

  1. A protected improvement channel — formative, mid-cycle, genuinely low-stakes, and explicitly not fed into any ranking. This is where you get honest signal.
  2. A reporting channel — summative, public where required, but presented as distributions, context and qualitative evidence rather than a single leaderboard digit.

The moment students believe their honest mid-term criticism will be aggregated into a number that hurts "their" university's standing, the MUM effect kicks in and the useful signal disappears.

The European dimension

This is not only a British story. The UK's NSS-driven league tables are the most developed example, but the same repurposing is spreading wherever satisfaction data is standardised and made comparable — pan-European instruments such as U-Multirank rank on self-reported student satisfaction, and national quality-assurance regimes increasingly publish comparable programme-level indicators. The European Standards and Guidelines (ESG) rightly expect student feedback to inform quality assurance, but they say nothing that requires compressing a programme into a single rankable score. Institutions across the European Higher Education Area are therefore free to meet their ESG obligations with rich, contextual public information while keeping their internal improvement data out of the ranking machinery altogether. The distinction between reporting for accountability and ranking for competition is one that national QA bodies and universities can still choose to preserve — and the argument of this piece is that they should.

Where Koji fits

Koji for Education is built for the protected improvement channel that league-table culture keeps eroding. Instead of one averaged satisfaction score waiting to be ranked, Koji runs AI-moderated conversational interviews that probe why behind a rating — surfacing the specific, actionable detail a mean can never carry. Its automatic thematic analysis of open-text feedback turns hundreds of student comments into structured themes, so a programme team can act on what students actually said rather than defend a decimal place. Formative, mid-cycle collection gives you signal while the course is still running and before any summative reporting is in play, and closing-the-loop action tracking keeps the emphasis on change rather than score-management. Because the moderation is standardised and bias-aware, and the data handling is GDPR/AVG-compliant and EU-hosted, the improvement channel stays trustworthy even in a sector obsessed with rankings.

Many teaching-and-learning centres also run broader user and stakeholder research; the same AI interview engine powers the main Koji platform for that work, so the methodology stays consistent across course evaluation and wider research.

None of this eliminates the pull of the league tables — nothing a vendor sells can. But it lets you keep at least one channel of feedback honest, which is the precondition for improving the teaching the rankings claim to measure.

Ready to protect your improvement channel from ranking pressure? See how Koji for Education helps European institutions gather feedback that stays honest — and act on it.