New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods9 min read

Do Student Evaluations Encourage Grade Inflation? The Incentive Problem

Wolfgang Stroebe argues that using student evaluations for high-stakes personnel decisions creates an incentive structure that rewards lenient grading and easy courses. We review the theory, the empirical evidence, the honest caveats, and what a defensible evaluation system should do instead.

Koji Education Team

Product

In brief: When student evaluations of teaching (SET) are used to make hiring, tenure, and merit decisions, they create an incentive for instructors to raise grades and reduce workload — because leniency reliably buys higher ratings. Wolfgang Stroebe (2016, 2020) calls this a systemic feedback loop that can contribute to grade inflation and reward exactly the teaching behaviours universities should discourage. The fix is not to abandon student feedback but to stop treating a single average as a high-stakes performance score, and to collect feedback that is harder to buy with a lenient grade.

The question this answers

Most debates about student evaluations focus on whether they are biased — by gender, accent, attractiveness, or discipline. Stroebe raises a different and arguably more uncomfortable question: even if a SET score were a perfect measurement, would the way universities use it corrupt the behaviour it measures? His answer is yes. This article summarises his argument, weighs the evidence for and against it, and explains how an evaluation programme can be designed so that good ratings cannot simply be purchased with easy grades.

What the research says

The anchor is Wolfgang Stroebe (2016), Why Good Teaching Evaluations May Reward Bad Teaching: On Grade Inflation and Other Unintended Consequences of Student Evaluations, published in Perspectives on Psychological Science (11(6), 800–816). Stroebe assembles a chain of well-replicated findings into a single causal argument:

  1. Students reward leniency. Across many studies, expected and actual grades correlate positively with SET ratings. Students who expect higher grades return higher evaluations of the same instructor and course.
  2. Leniency and low workload predict good ratings. Easier courses and lighter workloads are associated with more favourable evaluations, independent of how much students actually learn.
  3. Universities make leniency pay. Because administrators use SET averages for promotion, tenure, and merit pay, instructors face a rational incentive to grade leniently and lighten demands — the behaviours that most reliably raise their scores.
  4. The loop closes. Over time, this pressure contributes to grade inflation and can negatively relate ratings to learning, because the strategies that boost ratings (less challenge, more generous marks) are not the strategies that maximise long-run learning.

Stroebe expanded this into a fuller empirical treatment in Stroebe (2020), Student Evaluations of Teaching Encourages Poor Teaching and Contributes to Grade Inflation: A Theoretical and Empirical Analysis (Basic and Applied Social Psychology, 42(4), 276–294), marshalling longitudinal and cross-institutional data consistent with the incentive account.

The argument gains force from two adjacent bodies of work. Uttl, White & Gonzalez (2017), a multi-section re-analysis in Studies in Educational Evaluation, found that once prior-learning and section-size confounds are controlled, the correlation between SET ratings and actual learning is essentially zero — removing the central justification for treating ratings as a proxy for teaching effectiveness. And Boring, Ottoboni & Stark (2016), in ScienceOpen Research, showed in randomised and quasi-experimental data that SET ratings track instructor characteristics and student expectations better than they track learning. If ratings do not measure learning but do respond to grades, then rewarding high ratings rewards grade-giving, not teaching — exactly Stroebe's claim.

Why it matters for course evaluation in practice

For a quality-assurance office, the practical implication is sharp: the danger of SET is not only measurement error, it is the behaviour the measurement induces. A perfectly reliable instrument used as a high-stakes target will still distort teaching if the cheapest way to move the number is to grade more leniently. This is a textbook case of Goodhart's Law — when a measure becomes a target, it ceases to be a good measure.

Three concrete consequences follow:

  • Personnel decisions built on SET means are gameable. A department that ranks instructors on their average rating is, in effect, rewarding whoever was most generous with marks, holding teaching constant.
  • Grade inflation has a measurable feedback source. If raising grades raises ratings and ratings drive rewards, inflation is not a mystery of declining standards; it is a predictable response to the incentive.
  • The most demanding teaching is penalised. Instructors who set hard problems, mark strictly, and stretch students risk lower ratings — the opposite of what a programme committed to rigour wants to encourage.

None of this means student voice is worthless. It means student voice should inform formative improvement and be triangulated with other evidence, rather than being converted into a single summative score that personnel committees optimise against. See our companion analyses on grading leniency and on whether ratings measure learning for the underlying evidence.

Limitations and honest caveats

A PhD reader will rightly push back, and the objections matter:

  • Correlation is not proof of the causal loop. The grade–rating correlation is robust, but its interpretation is contested. The grade-leniency reading (students reward easy markers) competes with a validity reading (good teaching causes both more learning, hence higher grades, and higher satisfaction) and a student-characteristics reading (motivated students earn high grades and give high ratings). Stroebe argues the leniency component is large; others, including Marsh, argue the validity component is non-trivial. The honest position is that all three operate and their relative sizes vary by context.
  • Effect sizes are modest and heterogeneous. Grade–rating correlations typically sit in the small-to-moderate range and differ across disciplines and institutions. The incentive can be real without being the dominant driver of any single instructor's score.
  • Generalisability. Much of the foundational evidence is North American; European systems with external examiners, programme-level review, and different grading cultures may dampen or amplify the loop. The mechanism is plausible everywhere SET feeds promotion, but its magnitude is an empirical question per system.
  • Stroebe is a critic, not a neutral arbiter. His synthesis selects evidence to make a case. It should be read alongside defenders of SET validity (Benton & Cashin; Marsh & Roche) rather than as the last word.

Acknowledging these limits is itself part of good practice: the argument justifies not using SET means as high-stakes targets; it does not justify discarding student feedback.

How Koji incorporates this

Koji for Education is built on the premise that student feedback is most useful when it is hard to buy with a lenient grade and too rich to collapse into a single gameable number. Several mechanisms map directly onto the incentive problem Stroebe identifies:

  • Conversational, AI-moderated interviews instead of a satisfaction average. Rather than asking students to rate the instructor on a 1–5 scale, Koji conducts a structured conversational interview that probes what students learned, which activities helped, and where they struggled. A generous grade does not give a student more to say about how a concept was taught — so the evidence is less responsive to leniency than a global satisfaction item.
  • Structured question types that separate learning from satisfaction. Koji supports open_ended, scale, single_choice, multiple_choice, ranking, and yes_no questions, so an evaluation can distinguish "I enjoyed this course" from "this course made me work hard and I can now do X." Reporting these separately makes it visible when a high satisfaction score sits on top of low challenge — the signature of the leniency trade-off.
  • Automatic thematic analysis of open text. Koji clusters open responses into themes (clarity, workload, assessment fairness, perceived rigour) and surfaces them at programme level, giving committees behavioural evidence that is far harder to inflate than a mean.
  • Bias- and context-aware reporting designed to discourage ranking. Koji is designed to present distributions, themes, and confidence rather than a single decimal league-table figure, supporting the well-established recommendation that SET should not be used to rank instructors for personnel decisions. (See our note on why ranking instructors by scores fails.)
  • Formative, mid-cycle collection. Because Koji makes mid-semester feedback cheap to run, institutions can use evaluation primarily to improve a course in flight rather than to judge an instructor after the fact — shifting the purpose away from the high-stakes use that creates the incentive in the first place.

To be clear about what this does and does not do: Koji is designed to mitigate the incentive problem by making feedback richer and steering institutions away from single-number personnel use. It cannot eliminate the underlying pressure, which lives in how universities choose to use any feedback. Technology can make the gameable number less central; governance must do the rest.

Koji's core research platform at koji.so applies the same AI-moderated interview engine to product and customer research, where the equivalent failure mode — optimising a single NPS or CSAT figure — is just as well documented.

Related resources

References

  • Stroebe, W. (2016). Why good teaching evaluations may reward bad teaching: On grade inflation and other unintended consequences of student evaluations. Perspectives on Psychological Science, 11(6), 800–816. https://doi.org/10.1177/1745691616650284
  • Stroebe, W. (2020). Student evaluations of teaching encourages poor teaching and contributes to grade inflation: A theoretical and empirical analysis. Basic and Applied Social Psychology, 42(4), 276–294. https://doi.org/10.1080/01973533.2020.1756817
  • Uttl, B., White, C. A., & Gonzalez, D. W. (2017). Meta-analysis of faculty's teaching effectiveness: Student evaluation of teaching ratings and student learning are not related. Studies in Educational Evaluation, 54, 22–42. https://doi.org/10.1016/j.stueduc.2016.08.007
  • Boring, A., Ottoboni, K., & Stark, P. B. (2016). Student evaluations of teaching (mostly) do not measure teaching effectiveness. ScienceOpen Research. https://doi.org/10.14293/S2199-1006.1.SOR-EDU.AETBZC.v1