New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
best-practices10 min read

When a Course Evaluation Becomes a Target, It Stops Being a Measure: Campbell's Law and SET

Campbell's Law predicts that the more a course-evaluation score drives high-stakes decisions, the more it will be gamed and the less it will measure teaching. The evidence — grade inflation, watered-down rigor — and what a corruption-resistant evaluation looks like.

Koji Education Team

Product

Answer box (BLUF)

Donald Campbell's 1979 law states that "the more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort and corrupt the social processes it is intended to monitor." Applied to student evaluation of teaching (SET), it predicts something specific and testable: as course-evaluation scores are tied to promotion, tenure, contract renewal and league tables, instructors face rational pressure to raise the score by means that have nothing to do with teaching better — grading more leniently, reducing workload, teaching to the evaluation. There is real evidence this happens. The implication is not "abolish evaluation" but "design evaluation so the score is hard to game and never the sole high-stakes lever" — because a metric under maximum pressure measures the pressure, not the teaching.

What the research says

The law. Campbell articulated the principle in "Assessing the Impact of Planned Social Change" (1979), drawing on examples from policing statistics to standardized testing. Its twin, Goodhart's Law — "when a measure becomes a target, it ceases to be a good measure" (as later paraphrased) — makes the same point from economics. Both describe a mechanism: once actors know a number governs their fate, they optimise the number, and the cheapest route to a higher number is rarely the intended behaviour.

The gaming route is grades and rigor. The SET literature supplies the empirical mechanism Campbell predicts. Stroebe (2016, 2020) argues that SET creates a structural incentive for grade inflation: because expected grades and leniency are among the strongest correlates of ratings, an instructor can raise evaluations by grading more generously and demanding less — a direct corruption of the process SET is meant to monitor. Experimental and quasi-experimental work supports the incentive: the classic study by Greenwald and Gillmore (1997) found that relative grades correlate with ratings in a way consistent with a leniency effect, and concluded grading leniency is a "removable contaminant" of SET. The natural experiment by Braga, Paccagnella and Pellizzari (2014) at Bocconi is the sharpest evidence: instructors who produced the best downstream student performance received the lowest contemporaneous evaluations — the metric rewarded the opposite of the outcome it claimed to track.

High-stakes use amplifies the distortion. The corruption is a function of stakes, exactly as Campbell said. Where SET is decisive for personnel decisions, the pressure is maximal; the Ryerson arbitration case (2018) in Canada found SET scores so contaminated by bias that they could not fairly be used for tenure and promotion — an institutional recognition that a high-stakes measure had stopped measuring teaching.

Why it matters for course evaluation in practice

For a quality-assurance office or a dean, Campbell's Law reframes a familiar anxiety into a design principle:

  1. The score's validity degrades with its stakes. The same instrument can be a reasonable formative signal (low stakes, used by the instructor to improve) and a corrupted summative weapon (high stakes, used to fire or promote). It is not that the questions change — it is that attaching consequences changes the behaviour being measured.
  2. Grade inflation is the tell. If your institution sees creeping grades alongside rising evaluations, Campbell predicts you are watching the metric corrupt the process. Monitoring the relationship between grades and ratings is a way to detect gaming.
  3. Single-metric accountability is the failure mode. The corruption pressure is strongest when one number carries all the weight. Distributing the judgement across multiple, harder-to-game sources of evidence dilutes the incentive to game any one of them.
  4. Rigor is the casualty you cannot see in the score. A watered-down course can post excellent evaluations. The score looks healthy precisely while educational quality erodes — the most dangerous form of Campbell corruption because it is invisible on the dashboard.

Limitations and honest caveats

A sophisticated reader should resist over-reading the law, and we should mark its limits:

  • Campbell's Law is a hypothesis, not a proven constant. It describes a pressure, not a deterministic outcome. Many instructors do not game their evaluations, and professional norms, intrinsic motivation and institutional culture can blunt the incentive. The law tells you which way the wind blows, not that every ship is blown over.
  • The grade–rating correlation has innocent explanations. Students who learn more may both earn higher grades and rate the course higher — a genuine-quality pathway, not gaming. Greenwald and Gillmore's "removable contaminant" interpretation is contested precisely because leniency and learning are entangled. Detecting corruption specifically requires evidence that rigor fell, not just that grades and ratings move together.
  • Formative use is largely exempt. The law's force scales with stakes. Low-stakes, instructor-owned feedback carries little corruption pressure and remains valuable — abolishing all evaluation because high-stakes use corrupts it would be an overcorrection.
  • "Hard to game" can trade off against "easy to answer." Elaborate anti-gaming instrument design can raise respondent burden and lower response rates, creating a different validity threat. There is no free lunch.

How Koji incorporates this

Koji for Education cannot repeal Campbell's Law — no vendor can make a high-stakes number un-gameable. What the platform is designed to do is reduce the corruption surface and make gaming visible:

  • Conversational, open-ended evidence is costlier to fake than a Likert number. A single global rating is trivially inflatable; a set of AI-moderated interview answers about what specifically helped or hindered learning is not. By shifting weight from a gameable summary score toward substantive, probed accounts, Koji makes the "raise the number cheaply" route less available.
  • Triangulation is built into the model. Koji supports collecting evidence across cohorts and time and combining structured items (scale, single_choice, ranking, yes_no) with open-text and thematic analysis, so no single number carries all the decision weight — the multi-source dilution that Campbell's Law recommends.
  • Bias-aware and rigor-aware reporting. Because Koji surfaces themes rather than only means, a pattern of "easy, low-workload, entertaining" praise co-occurring with weak learning signals can be seen, rather than hidden inside a healthy-looking average — helping QA teams detect the invisible rigor casualty.
  • Formative, mid-cycle collection is treated as first-class. Koji is designed to support low-stakes, instructor-owned formative check-ins distinct from any summative use, keeping the highest-value, lowest-corruption use of evaluation front and centre and reducing reliance on a single end-of-term high-stakes score.
  • Closing-the-loop action tracking reframes evaluation as improvement evidence rather than a personnel verdict, lowering the stakes that drive gaming in the first place.

The framing is deliberately modest: Koji is designed to mitigate the corruption pressure Campbell identified, not to eliminate it. Any institution that makes one evaluation number decisive for careers will feel that pressure regardless of tooling; the defensible posture is multi-evidence, low-stakes-where-possible, and transparent. Teams running customer and product research beyond the classroom can apply the same AI-moderated interview engine through Koji's core platform at koji.so.

Related Resources

Designing a corruption-resistant evaluation

Campbell's Law is not a counsel of despair; it is a design brief. An evaluation system built with the law in mind minimises corruption pressure at each point where the pressure acts:

  • Separate formative from summative, and default to formative. The corruption pressure scales with stakes, so keep the largest share of evaluation low-stakes and instructor-owned. Reserve high-stakes summative use for the rare decisions that genuinely need it, and never let a single end-of-term mean be the deciding factor.
  • Never make one number decisive. Combine student feedback with peer observation, teaching portfolios, and — critically — evidence of learning and rigor. When no single indicator can settle a personnel decision alone, the incentive to game any one of them collapses.
  • Monitor the tell-tales. Track the grade–rating relationship and workload trends over time. Rising evaluations alongside softening grades or shrinking reading lists is the signature of Campbell corruption, and it is only visible if you look for it.
  • Weight substance over summary. A global "overall satisfaction" score is the most gameable object in the system. Shift decision weight toward specific, evidenced accounts of what helped or hindered learning, which are far harder to inflate cheaply.
  • Protect rigor explicitly. Because a watered-down course posts healthy scores, build rigor safeguards — external examiners, moderation, curriculum review — that operate independently of the satisfaction metric, so the invisible casualty becomes visible.

None of these repeals the law. They lower the pressure and raise the cost of gaming, which is the most any system can do. The institutions that get this right treat their evaluation number the way an engineer treats a load-bearing measurement under stress: with corroboration, tolerances, and a healthy suspicion of any reading that is too convenient.

References

Related articles

best-practices

Does Closing the Feedback Loop Actually Matter? The Evidence on Acting on Student Evaluations

Universities are good at collecting student feedback and bad at acting on it visibly. The research — Watson (2003), Leckey & Neill (2001), Shah et al. (2017) — shows that failing to close the loop drives the scepticism and declining response rates that quietly destroy your evaluation data.

best-practices

Should Student Evaluations Decide Tenure? The Ryerson Arbitration and the Limits of High-Stakes SET

A landmark 2018 Canadian arbitration ruled that student evaluations of teaching should not be used to measure teaching effectiveness for promotion and tenure. This guide explains the decision, the evidence behind it, and what it means for governing the high-stakes use of course-evaluation data.

research-methods

The Student-as-Consumer Effect: What a Consumer Mindset Does to Course Evaluations

When students see themselves as paying customers, what happens to how they rate courses — and how they learn? A look at Bunce, Baird & Jones (2017) and why the consumer frame quietly distorts the meaning of satisfaction scores.

research-methods

Does Grading Leniency Inflate Student Evaluations? What the Evidence Actually Shows

A research-grounded look at the grading-leniency hypothesis in student evaluations of teaching: what Marsh and Roche, Greenwald and Gillmore, and Stroebe found, why the grades-ratings correlation is genuinely ambiguous, and how to design evaluation so leniency cannot quietly buy better scores.