New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Methodology10 min read

Construct-Irrelevant Variance: The Validity Problem Underneath Course-Evaluation Bias

We argue endlessly about whether course evaluations are "biased". The more precise question, from measurement theory, is whether they are valid — and Messick's framework names exactly why so many of them are not.

Koji for Education

Research & Editorial Team · June 24, 2026

Bottom line up front: The debate over student evaluations usually gets framed as a fight about bias. Measurement theory offers a sharper and more useful frame: validity. Samuel Messick's unified theory names two threats that explain almost everything people dislike about student evaluation of teaching (SET) — construct-irrelevant variance (the score moves for reasons unrelated to teaching quality) and construct underrepresentation (the score fails to capture much of what good teaching actually is). Seen this way, "bias" is not a political accusation but a specific, diagnosable property of an instrument. And once you diagnose it, you can design feedback that has less of both.

From "is it biased?" to "is it valid?"

Validity is not a property of a test; it is a property of the interpretations and uses of a test's scores. That is Messick's central move, formalised in his 1995 American Psychologist paper and embedded in the Standards for Educational and Psychological Testing (AERA/APA/NCME, 2014). The question is never "is this number biased?" in the abstract. It is: "is the inference we are drawing from this number — this instructor teaches well — supported by evidence?"

That reframing matters because it tells you what evidence to look for and where instruments fail. Messick identified two principal threats to a valid interpretation, and SET scores are unusually exposed to both.

Threat one: construct-irrelevant variance

Construct-irrelevant variance (CIV) is variation in scores caused by something other than the construct you mean to measure. If "teaching effectiveness" is the construct, then any systematic movement in the score driven by something else is contamination.

The SET literature is essentially a catalogue of CIV:

  • Grading leniency and expected grade. Students who expect higher grades give higher ratings — a robust association across decades of studies, and a direct incentive problem we explore in our grading-leniency deep-dive.
  • Instructor characteristics unrelated to teaching. Perceived gender, age, accent, and physical attractiveness all move ratings for materially identical instruction. The MacNell et al. (2015) gender natural experiment is the cleanest demonstration: same online teaching, different perceived gender, different scores.
  • Course structure. Electivity, discipline, level and class size shift ratings regardless of who teaches — situational confounds we treat as a fundamental attribution error.
  • Delivery fluency. The Dr Fox effect — confident, expressive delivery inflating ratings independent of content — is CIV in its purest form.

Each of these is variance the score should not contain if it is to mean "teaching quality." The decisive empirical evidence is that the score barely correlates with learning at all: Uttl, White & Gonzalez's (2017) meta-analysis found the SET–learning relationship is near zero (r ≈ −0.02) once prior ability and small-sample bias are accounted for. A measure that does not track its target construct, but does track grades, gender and delivery style, is a measure dominated by construct-irrelevant variance.

Threat two: construct underrepresentation

The second threat is subtler and arguably more damaging. Construct underrepresentation (CUR) is the degree to which an instrument fails to capture the construct it claims to measure. Even a perfectly unbiased five-item Likert survey underrepresents teaching, because teaching is not five dimensions of student satisfaction.

Good teaching includes appropriate intellectual challenge, alignment between learning outcomes and assessment, the development of independent thinking, and effects that only appear in later courses. A questionnaire that asks "the lecturer explained clearly" and "I would recommend this course" samples a thin slice of that universe — and worse, it over-samples the comfortable, immediately likeable end of it. This is why active, effortful teaching can score lower: the instrument captures the feeling of learning, not learning, and the two diverge precisely where good pedagogy lives.

CUR is why "we average a few Likert items" is not a neutral methodological choice. It defines teaching as whatever those items happen to measure, and then quietly forgets everything they leave out. (We make the statistical version of this argument in why averaging Likert scores misleads.)

Why this framing is more useful than "bias"

Three reasons.

First, it is decision-relevant. Messick insisted validity be judged against use. A SET instrument might be valid enough for low-stakes formative reflection by an instructor and simultaneously invalid for ranking colleagues for promotion. The same number, two different validity verdicts, because the inference and the stakes differ. This dissolves a lot of unproductive argument: the question is not "are evaluations good or bad?" but "valid for what?"

First-rate quality assurance already implies this. The European Standards and Guidelines (ESG 2015) expect institutions to gather evidence and act on it — but using a construct-underrepresenting, CIV-laden number for high-stakes personnel decisions is exactly the misuse Messick warns against.

Second, it tells you what to fix. Reducing CIV means controlling or surfacing the confounds — benchmarking within comparable contexts, separating grade expectations from teaching feedback, probing whether a complaint is about the room or the explanation. Reducing CUR means widening the construct sample — asking about challenge, alignment and skill development, and capturing the open-ended detail that closed items cannot.

Third, it sets a fair standard for any replacement. New methods, including AI-assisted ones, must clear the same bar. The right question to ask of any course-feedback tool — including ours — is: does it reduce construct-irrelevant variance, and does it represent more of the construct?

"But every measure has some construct-irrelevant variance"

True, and worth conceding plainly. No instrument is perfectly valid; Messick's framework is about degree and consequence, not a pass/fail gate. Physical sciences tolerate measurement error all the time.

The honest response is twofold. First, the amount matters: when CIV is large enough that the score correlates more with grades and gender than with learning, the instrument has crossed from "imperfect" to "unfit for high-stakes use." Second, Messick added a consequential dimension to validity — the consequences of score use are part of the validity argument. An imperfect thermometer is fine; an imperfect thermometer used to deny someone tenure is a validity failure with a victim. So "all measures are imperfect" is not a defence of using a bad one for a serious decision. It is an argument for matching the instrument to the stakes and for improving the instrument where you can.

How Koji reduces both threats

Koji for Education is designed around exactly these two failure modes.

To reduce construct-irrelevant variance: Koji's standardized, bias-aware AI moderation removes the human-moderator inconsistency that adds noise to qualitative evaluation, and its conversational probing distinguishes a complaint about grades or scheduling from a complaint about teaching — letting the irrelevant variance be identified and set aside rather than silently absorbed into a mean. Programme- and institution-level reporting benchmarks within comparable contexts instead of across confounded ones.

To reduce construct underrepresentation: Koji uses six structured question types (open-ended, scale, single-choice, multiple-choice, ranking, yes/no) plus AI-moderated conversational interviews that follow up in the student's own words — sampling far more of the teaching construct than a fixed five-item form. Automatic thematic analysis turns that open-text breadth into structured evidence on challenge, alignment and skill development, not just satisfaction. Quality scoring flags thin or low-information responses so they do not masquerade as signal.

The same conversational engine powers general user research on the main Koji platform — because reducing CIV and widening construct coverage is simply what valid measurement looks like, in any domain.

None of this eliminates bias — no instrument can, and we are careful not to claim it. It mitigates construct-irrelevant variance and reduces construct underrepresentation, which is the most any honest measurement programme should promise.

The takeaway

"Biased" is a blunt word for a precise problem. Messick gives us the precise version: course evaluations are exposed to construct-irrelevant variance and construct underrepresentation, and their validity depends entirely on the inference and the stakes. Name the two threats, match the instrument to the decision, and design feedback that reduces both — that is the difference between arguing about evaluations and improving them.

Want course feedback engineered to reduce construct-irrelevant variance and capture more of what teaching actually is? Explore Koji for Education.