New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods9 min read

Does Grading Leniency Inflate Student Evaluations? What the Evidence Actually Shows

A research-grounded look at the grading-leniency hypothesis in student evaluations of teaching: what Marsh and Roche, Greenwald and Gillmore, and Stroebe found, why the grades-ratings correlation is genuinely ambiguous, and how to design evaluation so leniency cannot quietly buy better scores.

Koji Education Team

Product

In brief: Student evaluation scores are reliably, positively correlated with the grades students expect, but the cause of that correlation is contested. The leading "contamination" account (Greenwald & Gillmore, 1997; reinforced by Stroebe, 2020) holds that lenient grading inflates ratings; the leading "validity" account (Marsh & Roche, 2000) holds the correlation largely reflects genuine learning and prior motivation. Because the same data fit more than one story, no single average score should be read as bias-free. The defensible response is not to discard evaluations but to collect richer, reasoned evidence and to interpret numeric scores against expected-grade context.

The question, stated precisely

Across thousands of studies, one of the most stable empirical regularities in higher education is this: classes in which students expect higher grades tend to give their instructors higher evaluation scores. The correlation is modest but persistent, typically reported in the 0.10–0.30 range at the class level and sometimes higher.

What that correlation means is the entire debate. Three explanations compete:

  1. Grading-leniency (bias) hypothesis — instructors who grade easily are rewarded with inflated ratings that do not reflect better teaching. If true, evaluations create a perverse incentive to lower standards.
  2. Validity hypothesis — better teaching produces both more learning (hence higher grades) and more satisfied students (hence higher ratings). The correlation is then a feature, not a bug.
  3. Student-characteristics hypothesis — pre-existing motivation, interest, or ability drives both grades and ratings, so the link is confounded rather than causal in either direction.

Distinguishing these matters enormously for any university that uses evaluation scores in promotion, tenure, or programme review.

What the research says

Greenwald & Gillmore (1997) analysed roughly 600 classes across three consecutive quarters at the University of Washington and introduced additional data "markers" designed to discriminate among the competing theories. They concluded that the full pattern of markers was most consistent with grading leniency acting as a genuine, removable contaminant — meaning ratings could in principle be statistically adjusted for expected grade. Their framing, captured in the title "Grading leniency is a removable contaminant of student ratings" (American Psychologist, 52(11), 1209–1217), treated leniency as a real validity threat rather than a myth.

Marsh & Roche (2000) pushed back hard. Working with large multisection datasets and the multidimensional SEEQ instrument, they decomposed the grades-ratings relationship and argued that most of the correlation reflects two benign sources — prior student motivation and genuine learning — rather than leniency-driven bias. Their paper, pointedly subtitled "Popular Myth, Bias, Validity, or Innocent Bystanders?" (Journal of Educational Psychology, 92(1), 202–228), concluded that under appropriate conditions evaluations are multidimensional, reliable, relatively valid against external indicators of effective teaching, and "relatively unaffected" by grading leniency, class size, and workload. On this account, adjusting scores for expected grade would remove valid variance, not bias.

Stroebe (2020) revived and sharpened the bias case. In "Student Evaluations of Teaching Encourages Poor Teaching and Contributes to Grade Inflation: A Theoretical and Empirical Analysis" (Basic and Applied Social Psychology, 42(4), 276–294), he argued that the institutional use of evaluation scores in personnel decisions creates a feedback loop: students reward lenient, low-workload instructors with high ratings, instructors respond rationally to that incentive, and grade inflation follows. Stroebe's contribution is less a new dataset than a systems argument — the bias need not be large in any single class to distort behaviour across an institution over time.

Two decades of work, then, have not converged. The honest summary is that the grades-ratings correlation is real and robust, but its causal interpretation remains genuinely contested, and reasonable methodologists land on different sides depending on the instrument, the controls, and the institutional stakes they have in mind.

Why it matters for course evaluation in practice

For a quality-assurance office, the practical implications are sharper than the academic stalemate suggests.

  • A raw average is not interpretable on its own. If two instructors score 4.1 and 4.4 on a five-point scale but one taught a high-stakes gateway course with strict grading and the other an elective with generous marks, the numbers are not comparable. Expected-grade context is part of the measurement, not noise to ignore.
  • Stakes change behaviour. Stroebe's argument implies that the higher the consequences attached to evaluation scores, the stronger the incentive to teach to the rating rather than to the learning outcome. Using a single summative number for tenure decisions is precisely the condition under which leniency pressure is greatest.
  • Adjustment is not a clean fix. Greenwald and Gillmore's "removable contaminant" framing tempts institutions to statistically adjust scores for expected grade. But if Marsh and Roche are even partly right, such adjustment strips out valid signal. There is no purely statistical escape from the interpretation problem.

The methodologically defensible move is to stop treating the numeric average as a self-sufficient verdict and to triangulate it with evidence about what and how much students actually engaged with — evidence that a single Likert item cannot carry.

Limitations and honest caveats

A critical reader should hold several caveats firmly:

  • Correlation is not the effect size of bias. Even where leniency contributes to the grades-ratings link, the contaminating component may be small relative to the valid component. Pointing to a significant correlation does not establish that bias is large enough to change real decisions.
  • Most evidence is observational. The strongest causal evidence comes from rare randomised or quasi-experimental designs; much of the literature is correlational and cannot fully separate the three hypotheses. Both the bias camp and the validity camp are partly arguing over how to model confounds.
  • Generalisability is bounded. Greenwald and Gillmore studied one US institution; Marsh and Roche's instrument and conditions may not transfer to every European discipline or assessment culture. Findings about US letter-grade systems do not map cleanly onto, say, a Dutch 1–10 scale or a UK classification.
  • "Expected grade" is itself measured imperfectly. Students' grade expectations at the time of evaluation are noisy and may reflect optimism, anchoring, or recall effects, which weakens any adjustment built on them.

Showing these limits is not hedging — it is the reason no responsible QA process should let a single averaged score stand as proof of teaching quality or its absence.

How Koji incorporates this

Koji for Education is designed to mitigate the leniency-interpretation problem rather than pretend it away. Concretely:

  • It probes the reasons behind a number, not just the number. Koji's AI-moderated conversational interview follows a Likert rating with adaptive open-ended follow-ups — "what specifically made the course feel that way?" — so an evaluator can see whether a high score rests on genuine learning and clear feedback or on light workload and easy marks. This directly targets the ambiguity Marsh, Roche, Greenwald, and Gillmore argued over: the conversation surfaces the mechanism behind the score.
  • It separates workload and challenge from satisfaction. Using structured question types (scale, single_choice, open_ended), an evaluation can measure perceived workload, clarity of standards, and learning gain as distinct constructs, so analysts can inspect whether high ratings travel with low challenge — the exact pattern the leniency hypothesis predicts.
  • Its automatic thematic analysis of open text flags when positive sentiment clusters around "easy" or "relaxed grading" rather than "I learned a lot" or "feedback was useful," giving QA officers a qualitative signal that a numeric average cannot reveal.
  • Bias-aware reporting keeps expected-grade and workload context visible alongside scores instead of collapsing everything into one rank-ordered number, which is the institutional condition Stroebe warned against.
  • Formative, mid-cycle collection lets instructors improve teaching during the term rather than optimise for an end-of-term summative score, weakening the reward loop that drives leniency pressure.

Koji is built to mitigate — not eliminate — these biases: no instrument can resolve a causal debate that the field itself has not settled, but richer, reasoned evidence makes it far harder for leniency to quietly buy a better score. The same AI-moderated interview engine powers Koji's core research platform at koji.so for product and customer research, where the identical "ask why, not just how much" principle applies.

Related resources

Practical guidance for evaluation committees

If your institution uses evaluation scores in personnel or programme decisions, the grading-leniency debate translates into a few defensible operating rules. First, never read a single averaged score as a verdict on teaching quality; pair it with expected-grade and workload context so the committee can see whether a high rating travelled with genuine challenge or with light demands. Second, resist the temptation to apply a blanket statistical adjustment for expected grade — because the correlation is partly valid (Marsh & Roche) and partly contaminating (Greenwald & Gillmore), any one-size correction will either over- or under-correct, and the safer path is richer qualitative evidence. Third, lower the stakes attached to any single number: Stroebe's systems argument predicts that the more weight a committee places on the summative score, the stronger the incentive for instructors to teach to the rating. Fourth, separate the constructs you measure — ask about learning gain, clarity of standards, and workload as distinct items so leniency-driven satisfaction is visible rather than baked into one global score. These habits do not resolve the underlying causal question, but they make it far harder for leniency to quietly distort a high-stakes decision.

References

  • Greenwald, A. G., & Gillmore, G. M. (1997). Grading leniency is a removable contaminant of student ratings. American Psychologist, 52(11), 1209–1217. https://doi.org/10.1037/0003-066X.52.11.1209
  • Marsh, H. W., & Roche, L. A. (2000). Effects of grading leniency and low workload on students' evaluations of teaching: Popular myth, bias, validity, or innocent bystanders? Journal of Educational Psychology, 92(1), 202–228. https://doi.org/10.1037/0022-0663.92.1.202
  • Stroebe, W. (2020). Student evaluations of teaching encourage poor teaching and contribute to grade inflation: A theoretical and empirical analysis. Basic and Applied Social Psychology, 42(4), 276–294. https://doi.org/10.1080/01973533.2020.1756817