New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Sector trends9 min read

Everyone Knows the Satisfaction Number Is Flawed. Why Does Every University Still Use It?

Decades of evidence show a single course-satisfaction mean measures little about teaching. So why is it still the near-universal instrument across European higher education? The answer is not ignorance — it is institutional isomorphism. And naming the real cause points to the real fix.

Koji Education Team

Product · July 15, 2026

BLUF: There is a genuine puzzle at the heart of course evaluation. The research consensus against relying on a single satisfaction or "overall effectiveness" mean is unusually strong — a 2017 meta-analysis found student-rating scores essentially unrelated to how much students learn — and yet the flawed number remains the near-universal instrument across European universities. The usual explanations (inertia, laziness, ignorance) do not survive contact with the fact that the people running these systems are often methodologically sophisticated and privately sceptical. A better explanation comes from organisational sociology: DiMaggio and Powell's theory of institutional isomorphism. Universities keep the satisfaction number not because they believe it works, but because the same three forces that make organisations in a field converge on identical practices make abandoning it individually irrational. Understanding why it persists is the precondition for replacing it well.

The evidence problem is not the bottleneck

Start with what we know. In their 2017 meta-analysis of the multisection literature — the cleanest design available, comparing sections of the same course — Uttl, White, and Gonzalez concluded that student-evaluation ratings and student learning "are not related," and that earlier positive findings were largely artefacts of small samples and publication bias (Uttl, White & Gonzalez, 2017, Studies in Educational Evaluation). Add the bias literature — Boring, Ottoboni, and Stark found ratings "more sensitive to students' gender bias and grade expectations than they are to teaching effectiveness" (2016, ScienceOpen Research) — and the case against the single mean as a measure of teaching quality is about as settled as social science gets.

If evidence drove practice, the satisfaction number would already be gone or heavily qualified. It is not. That gap between knowledge and behaviour is the phenomenon to explain — and it is a sociological phenomenon, not a statistical one.

Institutional isomorphism: why organisations converge on the same flawed practice

In their classic 1983 paper "The Iron Cage Revisited," Paul DiMaggio and Walter Powell asked why organisations in a field become so strikingly similar to one another, and identified three mechanisms (DiMaggio & Powell, 1983, American Sociological Review). All three are visibly at work in course evaluation:

  • Coercive isomorphism — pressure from bodies an organisation depends on. External quality assurance under the European Standards and Guidelines (ESG), accreditation reviews, and national QA agencies expect to see systematic student feedback as evidence. A university that cannot produce standard evaluation numbers looks non-compliant, regardless of whether the numbers are meaningful. The instrument is retained as an audit artefact.
  • Mimetic isomorphism — copying peers under uncertainty. Because "good teaching" is genuinely hard to measure, universities reduce their uncertainty by doing what respected peers do. When every comparable institution runs an end-of-term Likert survey, adopting one is the safe, defensible choice — and deviating invites the question "why are you not doing what everyone else does?"
  • Normative isomorphism — professional norms and shared training. Quality-assurance and institutional-research professionals move between institutions, attend the same conferences, and carry the same templates. The satisfaction survey is embedded in the profession's toolkit, so it reproduces itself through the people who administer it.

The crucial implication: no individual actor in this system needs to believe the number is valid for it to persist. Each is responding rationally to coercive, mimetic, and normative pressure. The flawed instrument is an equilibrium, not a mistake — which is exactly why pointing out the evidence, again, does not move it.

This reframes the debate

Most critiques of course evaluation — including several we have written, on why averaging Likert scores misleads and on the reliability-versus-validity confusion — implicitly assume the problem is a knowledge deficit that better arguments can fix. Isomorphism says the problem is structural. It also explains adjacent puzzles: why the student-as-consumer satisfaction framing proved so sticky, and why national exercises persist even after their designers acknowledge the flaws, as the UK's National Student Survey reform shows.

It connects, too, to Goodhart's Law: once the satisfaction number is institutionalised as the metric of teaching quality, it becomes a target, and its already-weak validity degrades further. Isomorphism is how a measure gets institutionalised in the first place; Goodhart is what happens to it afterward.

The counterargument: "the standardisation is a feature, not a bug"

The strongest defence of the status quo turns isomorphism into a virtue. Standardisation, the argument runs, is precisely what makes evaluation useful for accountability: a common instrument lets you compare across modules, departments, and institutions, and gives external reviewers a legible, auditable signal. A perfectly valid but idiosyncratic per-course instrument would be useless for quality assurance. Comparability is worth some validity.

This is a real trade-off and deserves respect. But it proves less than it claims. First, comparability of an invalid measure is comparability of noise — being able to rank departments on a number that does not track teaching quality is not an accountability win, it is a precise way to be systematically unfair, a point sharpened by the bias findings above. Second, and more importantly, the trade-off is a false binary. It is entirely possible to have a standardised process — consistent, auditable, comparable across a programme — that collects richer evidence than a single mean. The reason we settle for the mean is not that richer standardised evidence is impossible; it is that, historically, the only thing cheap enough to standardise at scale was a Likert average. That constraint is exactly what has now changed.

Where Koji fits

If the satisfaction number survives because it was the only thing that could be standardised and audited cheaply at scale, then the escape route is a method that is equally standardised and auditable but captures far more. That is Koji for Education. Its AI-moderated conversational interviews apply the same bias-aware, consistent moderation to every student — satisfying the coercive and normative demand for a systematic, comparable, defensible process — while gathering probed, qualitative evidence and structured responses across six question types instead of a lone average. Automatic thematic analysis and quality scoring turn that evidence into programme- and institution-level reporting that reviewers can audit, and GDPR/AVG-compliant handling meets the compliance expectations that drive coercive isomorphism in the first place. In DiMaggio and Powell's terms, Koji lets an institution satisfy every isomorphic pressure — look compliant, look like (better than) its peers, fit the profession's norms — without accepting the validity cost of the single mean. You do not have to choose between standardisation and rigour anymore.

The same logic explains why organisations beyond higher education adopt the main Koji platform for user and customer research: they need evidence that is standardised enough to trust across a company but rich enough to actually act on.

A prediction the theory makes

If isomorphism, not evidence, sustains the satisfaction number, then reform will spread the way isomorphism spreads — not paper by paper, but once a critical mass of legitimate peers, and one or two accreditation bodies, signal that richer evidence is the new expected standard. Change of this kind tends to look slow, then sudden: a new practice becomes safe to adopt only when not adopting it starts to look like the deviation. That is why the most useful move for a reform-minded institution is not to publish another critique of the mean, but to become the visible, credible peer whose better method others can safely copy — converting mimetic pressure from a force that preserves the status quo into one that dismantles it.

The takeaway

Stop treating the persistence of the satisfaction number as a failure of intelligence to be corrected with one more citation. It persists because coercive, mimetic, and normative pressures make it the rational individual choice even for people who know it is weak. That diagnosis is liberating: it means the lever is not more evidence but a replacement that satisfies the same institutional pressures at lower validity cost. The moment a standardised, auditable, peer-legitimate method captures more than a Likert mean, the iron cage opens — because keeping the old number stops being the safe choice and starts being the indefensible one.

Want a course-evaluation method that is standardised, auditable, and actually valid? Explore Koji for Education.