Do Students Who Expect Higher Grades Rate Courses Higher? Expected-Grade Bias, Read Honestly
Students who anticipate a good grade tend to give better course evaluations — a correlation that has survived decades of research. But is it bias, or evidence that good teaching produces both? The distinction decides whether your data means anything.
Koji Education Team
Product ·
Yes, students who expect higher grades tend to give higher course evaluations — the correlation is one of the most replicated findings in the entire student-ratings literature. What is genuinely contested is what it means. If the expected grade colours the rating regardless of teaching quality, your evaluation data is contaminated at the source. If good teaching simply produces both more learning and higher grades, the correlation is a sign of validity, not bias. Universities routinely act as though the answer is settled. It is not — and the honest reading changes how much weight any evaluation score can bear.
This is a distinct problem from the more familiar grading-leniency question. Grading leniency is about the instructor's behaviour — do easy graders buy better ratings? Expected-grade bias is about the student's anticipation — does the grade a student thinks they will get shift their rating, before any objective learning is measured? The two are related but not identical, and conflating them muddies the debate.
What the evidence actually shows
The positive correlation between expected or received grades and course ratings is old and robust. Kenneth Feldman's foundational reviews in the 1970s established it; subsequent meta-analyses have found average correlations commonly in the range of roughly 0.2 to 0.3 — modest but persistent, and rarely absent. On this, proponents and critics of student evaluations largely agree. The fight is over the causal story behind the number.
Three explanations compete, and they are not mutually exclusive:
- The validity hypothesis. Better teaching causes both more learning and higher grades, so the grade-rating correlation reflects real teaching quality. Herbert Marsh and colleagues have long defended a version of this: student ratings, they argue, are substantially valid despite the correlation.
- The grading-leniency hypothesis. Instructors buy ratings with easy grades; students reward the leniency. A landmark 1997 special issue of American Psychologist — including work by Anthony Greenwald and Gerald Gillmore — assembled evidence that grading leniency is a genuine contaminant, not merely an artefact.
- The student-characteristics hypothesis. Abler or more motivated students both learn more and rate more generously, so a third variable drives both.
The most damaging evidence for anyone who wants to treat ratings as a clean measure of teaching came later. Bob Uttl, Carmela White and Daniela Gonzalez's 2017 meta-analysis in Studies in Educational Evaluation re-analysed the multisection studies — the strongest design, where different instructors teach the same course with a common final exam — and found that once study artefacts and small-sample effects are accounted for, student evaluations do not meaningfully predict how much students actually learn. If ratings barely track learning, then the grade-rating correlation cannot be comfortably explained as "both reflect good teaching". Something other than learning is doing the work.
Why expected grade is the more troubling version
The reason to separate expected grade from grading leniency is timing. Many students complete evaluations before the final grade is released — so what moves the rating is not the grade itself but the grade the student anticipates. That anticipation is shaped by mid-term marks, by how well the last assignment went, and by the peak-end impression of recent performance.
This makes expected grade a close cousin of the expectation-disconfirmation account of evaluation: students rate their satisfaction relative to what they expected, and an anticipated good grade primes a generous judgment before the term is even scored. It also interacts with the Dr. Fox effect and the halo effect: a confident expectation of success can radiate into ratings of clarity, organisation, and fairness that have nothing to do with the grade.
The practical danger is a perverse incentive. If anticipated grades lift ratings, and ratings feed tenure and promotion, the rational move for a rating-maximising instructor is to inflate students' sense of how well they are doing — which is the opposite of the honest, challenging feedback that produces the active learning students often rate lower in the moment.
"So just statistically control for grades" — the counterargument
A sophisticated objection says: fine, the correlation exists, but it is small (0.2–0.3), and we can simply regress it out — adjust each rating for the student's expected grade and read the residual as clean.
This is more attractive than it is safe, for three reasons. First, if the validity hypothesis is even partly true, controlling for grade removes real signal — you would be partialling out exactly the learning that good teaching produced, under-crediting effective instructors. Second, a correlation of 0.3 is an average, not a constant; it varies by discipline, level, and cohort, so a single blanket adjustment mis-corrects most courses. Third, and most fundamental, you rarely have a clean measure of "expected grade" to control for — self-reported expected grades are themselves shaped by the same halo you are trying to remove. Statistical control assumes you can cleanly separate variables that are, in practice, mutually contaminated.
The defensible conclusion is deflationary rather than corrective: expected-grade bias is a reason to hold evaluation scores more loosely, especially near decision thresholds, not a coefficient you can subtract your way out of. It is one more argument for triangulation — reading student ratings alongside peer review, learning evidence, and self-reflection — rather than for a cleverer single number.
Where Koji fits
A number cannot tell you why a student rated a course the way they did. That is the root of the expected-grade problem: a static form records a satisfaction score with no visibility into whether an anticipated grade, rather than the teaching, drove it. You are left inferring motive from an average.
Koji's AI-moderated conversational interviews attack the problem at the level of reasons. When a student rates a course highly, the AI moderator can probe what made it good — a specific concept they can now use, a piece of feedback that changed their work — surfacing whether the judgment rests on learning or merely on how the term felt gradewise. Its automatic thematic analysis then separates "I learned X" reasons from "I am doing well / it was easy" reasons across a whole cohort, giving an evaluation committee something a mean never can: the grounds of the score, not just its height. Koji's standardized, bias-aware moderation applies the same probing to every student, and its formative, mid-cycle collection lets you gather feedback decoupled from end-of-term grade anticipation.
To be precise about the claim: Koji does not eliminate expected-grade bias — no instrument can read a student's mind, and anticipation will always tint judgment. What it does is make the reasoning visible enough that a reader can weigh it, rather than treating a contaminated average as ground truth. Legacy SET tools give you the number and leave the contamination invisible; that is the difference.
The same confound — respondents rating an experience through the lens of how well they did — shows up in customer and product research, where satisfaction scores are quietly driven by the respondent's own outcomes. The shared AI interview engine behind koji.so probes the same "why" for that general research work.
If your evaluation scores feed real decisions, the question is not just how high they are, but what is holding them up. See how Koji for Education surfaces the reasons behind a rating.
FAQ
Do students who expect higher grades give higher course evaluations? Yes. The positive correlation between expected or received grades and course ratings is one of the most replicated findings in the literature, with average correlations commonly around 0.2 to 0.3. What is debated is whether this reflects genuine teaching quality or a bias.
Is expected-grade bias the same as grading leniency? No. Grading leniency concerns instructor behaviour — whether easy graders earn better ratings. Expected-grade bias concerns the student's anticipation — whether the grade a student expects shifts their rating, often before any grade is even released. They are related but distinct.
Does the grade-rating correlation prove course evaluations are invalid? Not on its own. Under the validity hypothesis, good teaching produces both learning and grades, so some correlation is expected. But Uttl and colleagues' 2017 meta-analysis found ratings do not meaningfully predict learning, which weakens the innocent explanation and strengthens the bias reading.
Can you just statistically control for expected grade? It is risky. If part of the correlation is genuine teaching signal, controlling for grade removes real information; the correlation varies across contexts so a blanket adjustment mis-corrects; and self-reported expected grades are themselves contaminated by the same halo. Holding scores loosely and triangulating is safer than a single correction.
How does Koji address expected-grade bias? Koji's conversational interviews probe the reasons behind a rating and its thematic analysis separates learning-based justifications from grade-anticipation ones, making the contamination visible. Mid-cycle collection also decouples feedback from end-of-term grade anticipation. Koji mitigates and surfaces the bias; it does not claim to eliminate it.