New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Evaluation bias10 min read

Should You Adjust Course-Evaluation Scores for Bias? What Statistical Correction Can and Cannot Fix

If class size, discipline, and expected grade bias student ratings, why not just regress them out and report a fair, adjusted score? Because "controlling for bias" quietly assumes you measured it correctly — and adjusting for the wrong variable can manufacture bias rather than remove it.

Koji Education Team

Product ·

Answer up front: Once you accept that student ratings are biased by things like class size, discipline, course level, and expected grade, an appealing fix presents itself: build a regression model, statistically adjust each instructor's score for those factors, and report a "bias-corrected" number that compares like with like. Some institutions and researchers do exactly this. It is a legitimate tool — but a dangerous default. Statistical adjustment only removes a bias you have correctly identified, measured, and modelled; it silently assumes there is no other bias you missed; and, most treacherously, adjusting for the wrong variable — a mediator or a collider rather than a genuine confounder — does not clean the data, it introduces new bias. Before you trust an adjusted score, you have to defend a causal claim about every variable in the model. Most adjustment done in practice never does.

The reasonable idea behind adjustment

The evidence that student evaluations carry non-teaching bias is substantial. Larger classes tend to score lower; some disciplines run systematically below others; and a student's expected grade correlates with the rating they give. There are also demographic biases — the much-cited experiment by MacNell, Driscoll and Hunt (2015) had students rate the same online instructor higher when they believed the instructor was male, across all twelve traits measured, regardless of the instructor's actual gender (MacNell et al., Innovative Higher Education, 2015). If these factors distort the score, the reasoning goes, then a fair comparison should net them out — just as epidemiologists adjust for age when comparing death rates.

The instinct is sound in principle. Comparing a 4.1 in a 300-person engineering lecture with a 4.4 in a 12-person history seminar is, as we have argued, a benchmarking trap. Adjustment promises to defuse it. The question is whether the promise holds.

Problem one: you can only adjust for the bias you modelled

Every adjustment model is a list of variables you decided to include. Its output is only as good as that list. The bias literature itself warns that the confounds are numerous and interacting — class time, size, whether the course is required or elective, difficulty, discipline, level, format, classroom, instructor characteristics, and student interest — and that, in the words of one review, it is simply "not possible to adjust for all of these", so no simple formula can completely correct biased ratings (review of bias control in SET). An adjusted score carries an air of precision — it looks corrected — while quietly omitting every bias you did not think to include. Unadjusted bias at least announces itself; adjusted-but-incompletely bias hides behind a clean number.

Problem two: adjusting for the wrong variable manufactures bias

This is the part practitioners most often get wrong, and it comes straight from causal inference. Not every variable that correlates with the rating is a legitimate thing to control for. Statisticians distinguish confounders (common causes of both the treatment and the outcome — good to adjust for) from mediators and colliders (bad to adjust for). Conditioning on a "bad control" introduces a spurious association where none existed, or erases a real one (Wikipedia, "Bad control", drawing on Pearl's causal framework).

Expected grade is the cautionary example. Institutions frequently want to adjust ratings for expected grade, on the theory that lenient grading buys good reviews. But consider the alternative causal path: good teaching produces more learning, more learning produces higher grades, and good teaching independently produces higher ratings. On that path, expected grade is a mediator of teaching quality, not a confounder. Control for it, and you subtract out exactly the legitimate signal you were trying to measure — penalising the excellent teacher whose students learned more and therefore expected better grades. Whether expected grade is a confounder (bias to remove) or a mediator (signal to keep) depends on an untestable causal story, and the same variable can be both at once. Adjust blindly and you may be manufacturing the unfairness you set out to cure. To make it concrete: picture two lecturers of identical true quality, one whose students expect a 70 and one whose students expect a 55 because that second cohort entered weaker. If the grade gap reflects the cohorts rather than lenient marking, then "adjusting for expected grade" drags the first lecturer's score down for a difference they did not create — the correction becomes the bias.

Problem three: adjustment can launder a construct problem

Suppose your instrument does not really measure teaching effectiveness in the first place — it measures a blend of entertainment, warmth, and workload, what we have called construct-irrelevant variance. No regression adjustment fixes that. You will produce a beautifully adjusted measure of the wrong thing. Worse, the adjustment ritual lends the number a credibility it has not earned: a committee that would have treated a raw score sceptically may take an "adjusted, bias-controlled" score at face value. Statistical sophistication applied to an invalid instrument does not rescue it; it disguises the problem.

But isn't an imperfect adjustment better than none?

This is the strongest defence of adjustment, and it is sometimes right. If a bias is large, well-established, plausibly a genuine confounder (not a mediator), and reliably measured, then adjusting for it can make a comparison fairer than leaving it raw — class size is often a reasonable candidate, since it plausibly biases ratings without being caused by teaching quality. The honest position is not "never adjust" but "adjust only when you can defend the causal role of each variable, quantify your uncertainty, and show your working." That means pre-specifying the model rather than fishing for one, reporting adjusted and unadjusted scores side by side, propagating the uncertainty (an adjusted point estimate with a wide interval is not a precise verdict), and — as with any single analytic choice — remembering that the result can flip depending on how you analysed it. Adjustment is a scalpel, not a laundry. Used with a defensible causal model it sharpens comparison; used as a reflexive clean-up step it does harm while looking rigorous.

It is also worth being honest that the entire premise — a stable, transportable set of bias coefficients — is shaky. Bias magnitudes vary by context and may not transfer cleanly from the US studies to European settings, so borrowing an adjustment factor from the literature rather than estimating it in your own data is its own error.

Where Koji fits

Koji for Education is built on a different theory of the problem: the most reliable way to reduce bias is not to model it out after the fact, but to collect better evidence in the first place. Its standardised, bias-aware AI moderation applies the same probing, neutral interview to every student, removing the human-moderator inconsistency and some of the framing effects that inflate variance before any adjustment is needed. Its conversational interviews gather the reasons behind a rating, so a low score can be read as "the assessment was unfair" or "the pace was too fast" rather than left as a bare number begging to be adjusted. And because Koji surfaces evidence for triangulation rather than a single league-table figure, it reduces the pressure to manufacture a falsely precise "corrected" score in the first place. Koji does not claim to eliminate bias — we are careful never to say "eliminate" — and where genuine confounders remain, transparent adjustment with a defensible causal model still has its place. The point is to need less post-hoc correction because the collection was better designed. (The same interview engine powers general user and customer research at koji.so.)

Adjustment is not a substitute for judgement about what is biasing your scores and why. Regress out the wrong variable and you do not remove bias — you invent it, and hand it the authority of a corrected number.

Closing the loop

If you would rather reduce bias at the source than launder it afterwards, see how Koji for Education standardises collection and surfaces the reasons behind every rating. A "bias-corrected" score is only as trustworthy as the causal story behind the correction.