Can Instructors Buy Better Ratings With Easy Grades? Instrumental Variables and the Grade–Evaluation Endogeneity Problem
Grades and ratings are jointly determined, so a naive regression overstates the grade effect. How Krautmann & Sander (1999) used two-stage least squares to recover the causal effect of expected grades on evaluations.
Koji Education Team
Product
Answer first. Yes — the evidence indicates instructors can, to a degree, "buy" better ratings through lenient grading, and a naive regression of ratings on grades overstates the effect because grades and ratings are jointly determined. Krautmann and Sander (1999) tackled this simultaneity with two-stage least squares (2SLS), an instrumental-variables method, and still found that expected grades causally raise evaluations. The lesson for quality assurance is methodological: when the thing you are using to "explain" a rating is itself shaped by the same forces that produce the rating, ordinary correlations are biased, and you need an identification strategy — an instrument, a natural experiment, or fixed effects — to recover a causal estimate.
What the research says
The correlation between the grade a student expects and the rating they give is one of the most replicated findings in the course-evaluation literature. The hard question is causal: does lenient grading cause higher ratings (a validity threat), or do good instructors simply produce both more learning — hence higher grades — and more satisfaction (a benign explanation)? The two stories imply opposite policy responses, and a raw correlation cannot separate them, because expected grade is endogenous: it is co-determined with the evaluation by unobserved factors such as instructor quality, course difficulty, and student ability.
Krautmann and Sander (1999), in Economics of Education Review, confronted this directly. They modelled evaluations allowing expected grades to be simultaneously determined with ratings, and estimated the system with both ordinary least squares (OLS) and two-stage least squares (2SLS). In 2SLS, the endogenous regressor (expected grade) is first predicted from instrumental variables — variables that affect grades but do not directly affect the rating except through grades — and that predicted, "cleaned" version is used in the second stage. Their result survived the correction: expected grades still had a positive causal effect on evaluations, supporting the interpretation that instructors can obtain better evaluations via more lenient grading. The endogeneity correction did not explain the effect away; it sharpened it.
The finding has been reinforced with different identification strategies. Ewing (2012), also in Economics of Education Review, used data from the University of Washington containing each student's relative expected grade and applied both a 2SLS/IV approach and instructor fixed effects to strip out unobserved instructor characteristics. He too found an incentive for lenient grading, and — importantly — showed that failing to account for instructor fixed effects biased the estimated grade effect upward, a direct demonstration that the identification strategy matters for the magnitude you report. Nowell (2007) examined relative grade expectations and reached compatible conclusions. Together these studies illustrate the methodological point cleanly: three different attempts to handle endogeneity (instruments, fixed effects, relative-grade measures) converge on a real but carefully-bounded causal effect.
The instrumental-variables logic itself is standard econometrics. Angrist and Krueger (2001) give the canonical exposition: a valid instrument must be relevant (strongly correlated with the endogenous regressor) and excludable (affect the outcome only through that regressor). Bound, Jaeger, and Baker (1995) warned that weak instruments — those only loosely correlated with the endogenous variable — can produce estimates more biased than the OLS they were meant to fix, a caution that applies squarely to the messy, observational world of course evaluations.
Why it matters for course evaluation in practice
Naive "controlling for grades" is not enough. A QA office that regresses ratings on expected grade and reports the coefficient is measuring a biased quantity, because grade is endogenous. The bias can go either way depending on what is omitted, so the direction is not even predictable without an identification strategy. Ewing's finding — that adding instructor fixed effects shrinks the grade effect — is a concrete warning that headline "grade-inflation-drives-ratings" numbers are often overstated.
High-stakes use amplifies the incentive. If evaluations feed promotion, tenure, or contract renewal, and lenient grading raises evaluations, the system creates a documented incentive to inflate grades — a perverse-incentive problem for any institution using SET summatively. Understanding the causal magnitude, not the inflated correlation, is what lets a committee judge how seriously to weight ratings against this contamination.
Identification thinking generalises. The grade–rating case is the vivid example, but the same endogeneity structure recurs whenever a "control variable" in an evaluation analysis is itself an outcome of teaching: attendance, engagement, self-reported learning. Course-evaluation analysts who internalise the instrumental-variables discipline stop treating regression coefficients as causal by default and start asking "what would identify this effect?" — a habit that improves every adjusted comparison a QA office makes.
Because clean instruments are genuinely scarce in this setting, the more common and often more credible identification routes are natural experiments (random assignment of students to sections, as in the value-added literature), regression discontinuity around grade or admission cutoffs, and fixed-effects designs. IV is one tool in a family; the unifying idea is that a causal claim about ratings requires a design that breaks the simultaneity, not just more covariates.
Limitations and honest caveats
- Valid instruments are hard to find and impossible to fully verify. The exclusion restriction — that the instrument affects ratings only through grades — is untestable and often implausible. Any instrument correlated with unobserved instructor quality violates it. Reasonable economists disagree about whether the instruments used in this literature are clean, and a sceptical reader should treat point estimates as bounded arguments, not settled facts.
- Weak-instrument bias is a real trap. Following Bound, Jaeger, and Baker (1995), a weak instrument can deliver 2SLS estimates that are more biased and less stable than OLS, with misleadingly tight-looking inference. First-stage strength must be reported and scrutinised.
- The effect sizes are modest and context-bound. "Instructors can buy better evaluations" does not mean grades are the dominant driver of ratings; the documented causal effects are real but bounded, and they come largely from US institutions decades ago. Generalisation to contemporary European quality-assurance systems, with different grading cultures and evaluation designs, should be cautious.
- Endogeneity correction improves internal validity, not construct validity. Even a perfectly identified causal effect of grades on ratings tells you the rating is contaminated by grading leniency; it does not tell you what the rating should measure. The remedy for a contaminated instrument is triangulation, not a cleverer regression.
- This is an analytical lens, not a routine reporting method. Few QA offices will run 2SLS on every course. The practical value is conceptual — knowing why raw grade-adjusted comparisons mislead — plus knowing when to commission a properly identified study before making a high-stakes claim.
How Koji incorporates this
Koji cannot manufacture an instrumental variable, and it does not pretend to. What it does is reduce the reliance on a single contaminated number and make the confounds visible.
- Separating satisfaction from learning signals. Because grade-driven satisfaction is a core threat, Koji's AI-moderated interviews are designed to probe specific teaching behaviours and learning experiences ("what did you find hard, and how did the instructor help?") rather than resting on a global satisfaction number that lenient grading can inflate. Behaviour-specific evidence is harder to buy with an easy grade than a summary rating is.
- Triangulation against non-rating evidence. Koji is built to combine student feedback with other quality signals so that a rating contaminated by grade expectations is not read alone — the same design principle that the endogeneity literature ultimately points to.
- Structured metadata for defensible adjustment. Koji captures course context (level, cohort, discipline) as structured fields, giving analysts the covariates a credible adjusted comparison requires — while the platform's reporting is framed to avoid presenting a raw grade-adjusted coefficient as if it were causal.
- Bias-aware framing of summative use. Consistent with the perverse-incentive concern, Koji's reporting is designed to support formative, behaviour-level interpretation rather than to encourage ranking instructors on a single number that grading leniency can move. These are framed as mechanisms designed to mitigate the grade-contamination threat, not to eliminate an incentive that is ultimately structural.
Product and customer-research teams face the same endogeneity trap — "did the feature cause satisfaction, or did already-happy users adopt it?" — and Koji's core platform at koji.so brings the same behaviour-specific, triangulated approach to those studies.
Related resources
- Propensity Score Matching for Confounds
- Grading Leniency and Student Evaluations
- Gender Bias: Natural-Experiment Causal Evidence
- Regression Discontinuity and Cutoff Effects
- The E-Value for Unmeasured Confounding
- SET Incentives and Grade Inflation (Stroebe)
Frequently asked questions
Do easy grades really cause better course evaluations?
The causal evidence says yes, to a bounded degree. Krautmann and Sander (1999) used two-stage least squares to account for grades and ratings being jointly determined and still found expected grades raise evaluations; Ewing (2012) reached the same conclusion with instrumental variables and instructor fixed effects. The effect is real but modest, and grades are far from the only driver of ratings.
What is endogeneity, and why does it bias the grade–rating estimate?
Endogeneity means an explanatory variable is correlated with the error term — here, expected grade is co-determined with the rating by unobserved factors like instructor quality and course difficulty. A plain regression of ratings on grades therefore mixes the causal effect with those confounds, and the bias can inflate or deflate the coefficient, so its raw value cannot be read as causal.
What is two-stage least squares (2SLS)?
2SLS is an instrumental-variables technique. In the first stage, the endogenous regressor (expected grade) is predicted from instruments — variables affecting grades but not ratings directly. In the second stage, that predicted, confound-free version replaces the original grade variable, yielding a causal estimate provided the instruments are strong and validly excluded.
Why can't we just add grade as a control variable?
Because grade is not a clean exogenous control — it is itself an outcome of teaching and shares causes with the rating. Ewing (2012) showed that omitting instructor fixed effects biased the grade effect upward. "Controlling for grades" with OLS produces a biased coefficient; recovering the causal effect requires an identification strategy such as IV, fixed effects, or a natural experiment.
Does this mean course evaluations are worthless?
No. It means the grade-related portion of a rating is a known contaminant that summative uses must account for. The constructive response is to weight ratings appropriately, probe specific teaching behaviours rather than global satisfaction, and triangulate with evidence that lenient grading cannot easily move — not to discard student feedback altogether.
References
- Krautmann, A. C., & Sander, W. (1999). Grades and student evaluations of teachers. Economics of Education Review, 18(1), 59–63. https://doi.org/10.1016/S0272-7757(98)00004-1
- Ewing, A. M. (2012). Estimating the impact of relative expected grade on student evaluations of teachers. Economics of Education Review, 31(1), 141–154. https://doi.org/10.1016/j.econedurev.2011.10.002
- Nowell, C. (2007). The impact of relative grade expectations on student evaluation of teaching. International Review of Economics Education, 6(2), 42–56. https://doi.org/10.1016/S1477-3880(15)30105-8
- Angrist, J. D., & Krueger, A. B. (2001). Instrumental variables and the search for identification. Journal of Economic Perspectives, 15(4), 69–85. https://doi.org/10.1257/jep.15.4.69
- Bound, J., Jaeger, D. A., & Baker, R. M. (1995). Problems with instrumental variables estimation when the correlation between the instruments and the endogenous explanatory variable is weak. Journal of the American Statistical Association, 90(430), 443–450. https://doi.org/10.1080/01621459.1995.10476536
Related articles
Propensity Score Matching for Course-Evaluation Confounds
How propensity score matching (PSM) adjusts student-evaluation-of-teaching comparisons for confounds like class size, course level, discipline, and grading — and where it honestly falls short.
How Strong Would the Hidden Confounder Have to Be? The E-Value for Adjusted Course-Evaluation Comparisons
Every adjusted course-evaluation claim invites the objection "but you did not control for X". The E-value, from epidemiology, quantifies exactly how strong that unmeasured X would have to be to explain away your finding — turning a vague worry into a number.
Does Grading Leniency Inflate Student Evaluations? What the Evidence Actually Shows
A research-grounded look at the grading-leniency hypothesis in student evaluations of teaching: what Marsh and Roche, Greenwald and Gillmore, and Stroebe found, why the grades-ratings correlation is genuinely ambiguous, and how to design evaluation so leniency cannot quietly buy better scores.
Do Student Evaluations Encourage Grade Inflation? The Incentive Problem
Wolfgang Stroebe argues that using student evaluations for high-stakes personnel decisions creates an incentive structure that rewards lenient grading and easy courses. We review the theory, the empirical evidence, the honest caveats, and what a defensible evaluation system should do instead.