Construct Underrepresentation: The Course-Evaluation Validity Threat That Matters More Than Bias
The bias debate has swallowed the course-evaluation conversation. But the deeper validity problem is not that a five-point scale is biased — it is that it samples only a sliver of what "teaching quality" actually means. Messick named this fifty years ago, and it changes what a good instrument has to do.
Koji Education Team
Product · July 14, 2026
The short answer: Validity theory identifies two great threats to any measurement: construct-irrelevant variance (the instrument picks up things it should not — bias, mood, an instructor's looks) and construct underrepresentation (the instrument fails to capture important parts of what it claims to measure). The course-evaluation debate is almost entirely about the first. The more damaging and less-discussed threat is the second. A short satisfaction survey does not just risk contamination; it structurally cannot sample most of the construct of teaching quality — course design, assessment validity, inclusion, durable learning, intellectual challenge. Fixing bias in a five-item Likert form still leaves you measuring a fraction of the thing. This is the validity argument that should reframe how institutions evaluate teaching.
Messick's two threats, and why we only argue about one
When the psychometrician Samuel Messick set out the modern, unified theory of validity, he named two principal threats to the meaningfulness of a test score. The first is construct-irrelevant variance: score inflation or deflation caused by systematic factors unrelated to the attribute you intend to measure — the classic example being reading load on a mathematics test, which lowers scores for reasons that have nothing to do with mathematical ability (Messick, 1995). The second is construct underrepresentation: the degree to which a measure 'fails to capture important aspects of the construct', eliciting too narrow a sample of behaviour relative to the universe of behaviours that would actually indicate the trait (see this validity primer for discipline-based education researchers).
Almost the entire public argument about student evaluations of teaching (SET) is a fight about construct-irrelevant variance. Does gender bias contaminate scores? Does physical attractiveness? Grading leniency? Accent? These are real and important — we have written about several of them — and they are all, in Messick's terms, forms of contamination: the scale is picking up signal it should not. But notice what the entire bias literature quietly concedes: it assumes that if only we could scrub out the irrelevant variance, the remaining score would be a valid measure of teaching quality. Construct underrepresentation says that assumption is false. Even a perfectly unbiased five-item satisfaction survey would still measure only a sliver of the construct.
What a satisfaction survey structurally cannot see
What is 'teaching quality'? Any serious account includes at least: the coherence of course design and constructive alignment; the validity and fairness of assessment; the intellectual challenge and appropriate difficulty of the material; whether the course builds durable, transferable knowledge rather than exam-week cramming; whether it is inclusive across a diverse cohort; and whether students developed as independent thinkers. A typical end-of-term instrument asks students to agree or disagree, on five points, that the instructor was clear, well-organised, and available. Those items are not wrong — but they sample the presentation facet of teaching and almost nothing else. The construct is a continent; the instrument photographs one city.
The most striking evidence for this under-sampling is the relationship between SET scores and actual learning. The most rigorous synthesis available — Uttl, White and Gonzalez's 2017 meta-analysis in Studies in Educational Evaluation — re-analysed the multisection studies that had long been cited as proof that ratings track learning. Once prior ability was properly controlled, the correlation between SET ratings and student learning was essentially zero; the authors concluded that 'students do not learn more from professors with higher SET ratings.' Read through the lens of construct underrepresentation, that is not a scandal — it is the predicted result. If your instrument samples satisfaction and presentation but not learning, it will correlate with satisfaction and presentation but not learning. The zero correlation is the under-representation showing up in the data.
Why under-representation is the more dangerous threat
Construct-irrelevant variance is, in principle, correctable. You can warn raters, adjust for known biases, use anchoring vignettes, redesign items. The contamination is a leak you can try to plug. Construct underrepresentation is different: you cannot fix an omission by cleaning up the items you already have. A survey that never asks about assessment fairness will never measure it, no matter how unbiased its existing questions are. Under-representation is a ceiling on what the instrument can ever tell you, and no amount of de-biasing raises that ceiling.
There is a second, subtler danger. Because the number that comes out of a narrow instrument looks authoritative, institutions treat it as if it represented the whole construct. A 4.1 becomes 'good teaching' in a promotion file, when it strictly means 'students, on average, agreed the presentation was clear.' This is the jingle fallacy — assuming that because we labelled the scale 'teaching effectiveness', it measures teaching effectiveness. Under-representation plus an authoritative-looking score is how a satisfaction measure gets quietly promoted into a teaching-quality verdict.
But isn't a narrow, reliable measure better than a broad, fuzzy one?
The strongest counterargument comes straight from measurement theory: narrowing a construct improves reliability. A tightly focused set of items about clarity and organisation will have a high Cronbach's alpha; broaden the instrument to cover design, assessment, inclusion and challenge and you introduce multidimensionality and measurement noise. Isn't a precise measure of a small thing preferable to a vague measure of a big thing?
Only if the small thing is the thing you actually care about. Messick's whole point is that reliability is necessary but not sufficient — a perfectly reliable measure of the wrong (or a too-narrow) construct is still invalid for the intended use. A bathroom scale reliably measures weight; it is an invalid measure of health, however consistent its readings. The answer is not to abandon rigour for a sprawling questionnaire nobody finishes. It is to represent the construct adequately with methods suited to each facet — some quantitative, some qualitative, some direct measures of learning, some evidence of course design — and to be honest that a single five-point mean cannot carry the weight institutions place on it. Triangulation across sources, not a longer Likert form, is the route out; see our piece on triangulating teaching evaluation across multiple evidence sources.
How Koji widens the sample of the construct
Koji for Education is designed to represent more of the construct, not to polish a narrow one. Its AI-moderated conversational interviews do what a fixed form cannot: they follow up. When a student says a course was 'hard', the interview probes whether that meant productive challenge or poor scaffolding — the difference between a course that under-serves students and one that does exactly what a rigorous course should. That probing samples facets (challenge, assessment fairness, coherence of design) that a static scale never reaches. Koji's six structured question types (open-ended, scale, single- and multiple-choice, ranking, yes/no) let an instrument cover breadth without collapsing everything into one averaged number, and its automatic thematic analysis turns the resulting open text into structured evidence about design, inclusion and learning rather than leaving it as unread free text.
Crucially, Koji does not claim to measure learning outcomes directly — no self-report instrument can, and pretending otherwise would be its own validity error. What it does is stop under-representing the construct: it collects evidence across many facets of teaching quality and reports them distinctly, so a programme sees a portrait rather than a single misread pixel. Institutions can then triangulate that richer evidence with direct measures of learning and peer review. The same conversational engine underpins Koji's main research platform, where the identical principle applies: a Net Promoter Score tells you a customer's temperature, not why they feel it.
The takeaway
The bias debate is important but incomplete. Even if every source of construct-irrelevant variance were eliminated tomorrow, a five-item satisfaction survey would still fail the more fundamental test: it does not sample most of what teaching quality means. Construct underrepresentation is the quieter, deeper, less-fixable threat — and recognising it changes the goal. The aim is not a cleaner number for a narrow construct. It is an instrument, and an evaluation system, that represents the construct honestly enough to deserve the decisions built on it.