New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Sector trends10 min read

Goodhart's Law and Course Evaluations: What Happens When a Measure Becomes a Target

"When a measure becomes a target, it ceases to be a good measure." Tie tenure and contracts to student-evaluation scores and you do not get better teaching — you get grade inflation, easier courses, and metric-gaming. Here is the evidence, and what responsible-metrics frameworks say to do instead.

Koji for Education

Research & Editorial Team · June 22, 2026

Bottom line up front: The moment a course-evaluation score stops describing teaching and starts deciding careers, it begins to corrupt the thing it was meant to measure. This is Goodhart's Law, and in higher education its signature is unmistakable: grade inflation, lightened workloads, teaching-to-the-survey, and instructors optimising for satisfaction rather than learning. The problem is not that student feedback is worthless — it is invaluable as formative signal. The problem is what happens when you make it a summative target. The remedy is not to stop listening to students; it is to stop using a single satisfaction number as a high-stakes lever.

The law, stated precisely

The economist Charles Goodhart's original 1975 observation is usually paraphrased by the anthropologist Marilyn Strathern: "When a measure becomes a target, it ceases to be a good measure" (Goodhart's Law, Wikipedia summary). Its close cousin is Campbell's Law (Donald Campbell, 1979): "The more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort and corrupt the social processes it is intended to monitor."

Both describe the same failure mode. An indicator works as a measure precisely because the people being measured are not yet acting on it. Attach stakes — tenure, promotion, contract renewal, departmental funding — and behaviour reorganises around the indicator. The number keeps rising; the underlying quality it once tracked drifts away from it. Student evaluations of teaching (SET) are a textbook case, because the people who set the score (students) and the people who can most cheaply move it (instructors, via grades and workload) are the same population in a repeated game.

The evidence that SET-as-target distorts behaviour

This is not theory. The distortions show up in the data.

  • Grade inflation tracks the incentive. A recent meta-analysis of the SET literature reports that departments tying contract renewal to minimum SET thresholds exhibited roughly a 0.27 GPA-point rise relative to matched controls — a direct fingerprint of grades being used to buy ratings (Student Evaluations of Teaching Fail to Predict Learning, 2025 preprint).
  • Satisfaction is decoupled from learning. The most rigorous synthesis, Uttl, White and Gonzalez (2017) in Studies in Educational Evaluation, re-analysed 97 multisection studies and found that once you correct for small-sample and publication bias, the correlation between SET ratings and actual learning is essentially nil — and negative (r = -.06) in the studies that adjusted for prior ability. If ratings do not track learning, optimising ratings cannot be optimising learning.
  • The incentive runs the wrong way. Stroebe (2020) argues that SET, used summatively, actively encourages poor teaching: the fastest route to a higher score is often an easier course and more generous marking, not deeper learning. The mechanism is Goodhart's Law operating on a yearly cycle.

Put together: the indicator that institutions most often elevate to a target is one that does not measure learning, and responds to gaming in a direction that harms it. That is close to a worst case for Campbell's Law.

The legal system has noticed

The corruption argument is no longer confined to academic journals. In 2018, an arbitrator ruled in a dispute between Ryerson University and its faculty association that SET scores must not be used to measure teaching effectiveness for promotion and tenure decisions, finding the surveys "imperfect at best and downright biased and unreliable at worst." Notably, the remedy ordered was not "ignore students" but better measurement hygiene: replace the misleading numerical ranking, report results as frequency distributions with response rates, and train evaluators on the inherent biases (Inside Higher Ed, 2018). When a labour arbitrator independently arrives at "stop using the headline number as a target," the institutional risk of ignoring Goodhart's Law has become concrete.

"But we have to evaluate teaching somehow"

The strongest objection is pragmatic: institutions are accountable, ENQA's Standards and Guidelines (ESG) require evidence that teaching is monitored, and "abolish the metric" is not an option for a quality office. This is correct — and it is exactly why Goodhart's Law is so often misread as an argument for nihilism. It is not. The law does not say do not measure; it says do not let one measure become the target. The escape is well charted by the responsible-metrics movement that grew up around research assessment and applies directly to teaching:

  • The Leiden Manifesto (Hicks, Wouters et al., Nature, 2015) sets ten principles, the first of which is that quantitative evaluation should support, not replace, expert qualitative judgement, and that indicators must be used in context (Leiden Manifesto).
  • The Metric Tide (Wilsdon et al., 2015) frames responsible metrics around robustness, humility, transparency, diversity and reflexivity.
  • DORA insists that assessment rest on the substance of the work, not on a single proxy indicator.

Translated to course evaluation, these say: keep collecting student feedback, but use it as one contextualised input to a human judgement, never as a standalone threshold that triggers a consequence. That is the difference between a measure that informs and a target that corrupts.

A second objection: isn't formative use just as gameable?

A sharp reader will note that even formative feedback can be gamed if instructors know it will eventually surface. True — but the pressure is far weaker when no single number gates a career, when feedback is collected mid-cycle to be acted on rather than end-of-cycle to be judged, and when the evidence base is plural. The point of Goodhart's Law is about the concentration of stakes on one proxy, not the existence of measurement. Diffuse the stakes across triangulated evidence (peer review, learning outcomes, self-reflection, student voice) and the cheapest gaming strategy — inflate grades, lighten the load — stops paying off, because no one of those sources can be moved by it alone.

How Koji is designed against Goodhart's Law

Koji for Education is built on the premise that student feedback is most valuable when it is formative and rich, not summative and thin. Several design choices follow directly from Goodhart's Law:

  • Beyond the single number. Koji's AI-moderated conversational interviews probe why behind any rating, and its automatic thematic analysis turns open-text feedback into structured themes — so the evidence is hard to compress into one gameable threshold.
  • Mid-cycle by default. Formative collection during the course, with closing-the-loop action tracking, repositions evaluation as something done for improvement rather than to the instructor — the opposite of the end-of-term high-stakes survey.
  • Quality scoring and standardised moderation. A consistent, bias-aware AI moderator reduces the human-moderator inconsistency that makes naive scores noisy, while quality scoring helps distinguish substantive feedback from satisficing.
  • Reporting that resists the league table. Programme- and institution-level reporting is designed to present distributions and themes, in the spirit of the Ryerson remedy and the Leiden Manifesto, rather than a single rankable mean.

To be precise: no tool can repeal Goodhart's Law. If an institution insists on turning any Koji output into a sole high-stakes threshold, the corruption pressure returns. What Koji does is make the responsible path — formative, plural, contextualised feedback — the easy one, and the gameable single number harder to manufacture. The governance decision to use feedback formatively still belongs to the institution.

Teams doing general customer or employee research meet the same trap whenever an NPS or CSAT number becomes a bonus target; the main Koji platform applies the same interview engine and the same anti-Goodhart philosophy there.

The takeaway

Goodhart's Law is not a reason to stop evaluating teaching — it is a reason to stop targeting a single evaluation score. The evidence is clear that high-stakes SET drives grade inflation and easier courses while failing to track learning, and arbitrators and responsible-metrics frameworks alike now point to the same fix: collect rich student feedback, use it formatively, triangulate it with other evidence, and never let one number become the thing people optimise. The measure stays good precisely as long as it never becomes the target.


Koji for Education is built for formative, evidence-rich student feedback that informs human judgement instead of replacing it with a gameable number. See how Koji approaches responsible evaluation.