New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Evaluation bias8 min read

Do Easy Graders Get Better Course Evaluations? The Grading-Leniency Debate, Read Honestly

Students who expect higher grades tend to give higher ratings. That correlation is real and well-replicated — but what causes it is one of the most genuinely contested questions in evaluation research. Here is what the evidence actually supports.

Koji Education Team

Product · June 21, 2026

Short answer: Yes, there is a positive correlation between the grades students expect and the ratings they give their instructors — this is one of the most replicated findings in the student-evaluation literature. But the correlation is modest, and its cause is genuinely disputed. Some of it reflects grading leniency (a bias), some reflects better teaching producing both better learning and better ratings (validity), and some reflects pre-existing student motivation (a spurious third variable). A single average score cannot tell you which mechanism is operating in your data — and that is the real problem.

The claim, stated fairly

Walk into almost any faculty common room and you will hear some version of it: "If you want good evaluations, grade easy." The worry is the grading-leniency hypothesis — the idea that instructors can buy higher student evaluations of teaching (SET) by relaxing standards, and that students reward generous grading with generous ratings.

It is a serious charge. If true, it means SET scores are partly a measure of how lenient an instructor is, not how well they teach — and that every promotion or tenure decision resting on those scores is quietly importing that bias. So it deserves to be examined with the same rigor a quality-assurance officer would demand of any other measurement claim.

What the evidence actually shows

Start with the part that is not in dispute. Students who expect higher grades do, on average, give higher ratings. Researchers typically use expected grade rather than actual grade because, at the point students complete most evaluations, they do not yet know their final mark. The positive association is robust and appears across decades of studies.

The disagreement is entirely about why. There are at least three competing explanations, and they are not mutually exclusive:

  1. Grading leniency (bias). Lenient grading directly inflates ratings, independent of teaching quality. This is the version that should worry us.
  2. Validity. Better teaching produces both more learning (hence higher grades) and more satisfied students (hence higher ratings). On this account the correlation is a feature, not a bug — it is exactly what you would expect if SET measured something real.
  3. Student characteristics (a spurious link). More motivated, better-prepared, or more interested students both work harder (earning higher grades) and rate courses more favorably. The grade-rating link is then driven by a third variable neither caused.

Here is where serious researchers part company. Anthony Greenwald and Gerald Gillmore argued, in work published in the late 1990s, that the pattern of evidence is most consistent with grading leniency operating as a genuine biasing influence — and provocatively titled one paper "Grading Leniency Is a Removable Contaminant of Student Ratings."

Herbert Marsh and Lawrence Roche pushed back hard. In their analysis "Effects of Grading Leniency and Low Workload on Students' Evaluations of Teaching: Popular Myth, Bias, Validity, or Innocent Bystanders?" (Journal of Educational Psychology, 2000), they concluded that while a grading-leniency effect may produce some bias in SET, the support for it is weak and the size of any such effect is likely to be insubstantial. In Marsh's broader reviews, the leniency story is treated as a popular myth that survives more on intuition than on data.

So the honest summary is not "easy grading buys ratings" and it is not "grades are irrelevant." It is: a modest correlation exists, leniency probably contributes a little to it, but leniency does not come close to explaining most of the variance in evaluation scores. Anyone who tells you the question is settled in either direction is overselling.

Why a single number cannot resolve this

Notice what the debate is really about. Everyone is looking at the same correlation and disagreeing about its causal structure. That is not a failure of the researchers — it is a structural limitation of the instrument. A Likert mean is a single scalar. It compresses teaching quality, grading generosity, student motivation, course difficulty, and a dozen other things into one number, and then asks that number to adjudicate between explanations it can never distinguish.

This connects to a wider methodological point we have made before: averaging Likert scores discards exactly the information you need to interpret them. If a course scores 3.8, the leniency debate means you cannot even be sure whether a higher score would be good news. The number is causally ambiguous by construction.

It also means the leniency worry is not really an argument against gathering student voice. It is an argument against gathering it in a form so thin that you can never tell what it is measuring.

"But doesn't this just prove evaluations are worthless?"

This is the strongest counterargument from the other direction, and it is worth taking seriously. Some critics — pointing to work like Bob Uttl and colleagues' 2017 meta-analysis, which found that once you correct for small-sample bias, SET ratings are essentially unrelated to how much students actually learn — conclude that the whole enterprise should be scrapped.

We think that goes too far. "Ratings do not validly measure learning outcomes" and "ratings measure nothing useful" are different claims. Students are excellent witnesses to their own experience: whether they could follow the explanations, whether feedback arrived in time to act on it, whether the workload was survivable, whether they felt able to ask questions. Those are real, decision-relevant facts that no exam score captures. The mistake is asking a satisfaction number to stand in for a learning-effectiveness measure it was never built to be — and then making high-stakes decisions on it. The fix is to use student voice for what it is genuinely valid for, and to triangulate it with other evidence. (We make that case in detail in Triangulation in Teaching Evaluation.)

What to actually do about it

If grading leniency contaminates ratings even a little, three practical responses follow:

  • Stop treating the mean as self-explanatory. A score without context — without knowing the grade distribution, the cohort, the difficulty — is uninterpretable. Report distributions and context, not just averages. (See why benchmarking a 4.1 against a 4.4 is a trap.)
  • Ask questions leniency cannot easily inflate. "I expect a good grade, so I rate everything highly" is a halo a vague satisfaction item invites. Specific, behavioral questions — Did you receive feedback in time to improve before the next assessment? Can you give an example of something that helped you understand a difficult concept? — are much harder to answer generously out of grade optimism.
  • Probe the "why" rather than logging the "what." The single most useful thing you can do is ask a follow-up. When a student rates a course highly, the next question — what specifically made it work for you? — separates "the teaching was genuinely clear" from "the grading was generous." A static survey never asks. A conversation does.

How Koji addresses this

This is precisely the gap Koji for Education is built to close. Instead of a one-shot Likert form that produces a causally ambiguous average, Koji runs an AI-moderated conversational interview that adapts to each student. When a student gives a high rating, the AI probes why — surfacing whether the praise is about clarity, structure, and feedback (signals of real teaching quality) or about easy marking (a leniency signal you would want to discount). Its six structured question types let you combine scales with open-ended, example-seeking prompts that grade optimism cannot easily inflate, and the automatic thematic analysis tags the reasons behind a score, not just its magnitude.

Because the moderation is standardized and bias-aware, every student gets the same careful, neutral probing — no human interviewer drifting into leading questions — and the platform reports at programme and institution level with the grade and cohort context that makes a score interpretable in the first place. Koji does not pretend to eliminate leniency effects; no instrument can. It surfaces them, so committees can read a score knowing what is behind it.

The same conversational interview engine powers the main Koji platform for general customer and user research — wherever a single satisfaction number hides the reason behind it.

The grading-leniency debate has run for forty years because a Likert average can never settle it. The way forward is not to win the argument. It is to collect evidence rich enough that the argument no longer has to be guessed at.

Ready to move beyond ambiguous averages? See how Koji for Education evaluates courses.