Are Course Evaluations Valid? Reliability vs Validity, and Why the Difference Matters
Student evaluations of teaching are remarkably reliable - and that is exactly what lulls institutions into trusting them. Reliability is not validity. A measure can be perfectly consistent and still measure the wrong thing. Here is why that distinction is the most important - and most ignored - fact about course evaluations.
Koji for Education
Research & Editorial Team ·
Bottom line up front: Student evaluations of teaching (SET) are reliable — consistent across students, stable over time, internally coherent. They are not, on the strongest available evidence, valid measures of teaching effectiveness: the largest meta-analysis finds SET ratings essentially unrelated to how much students learn, and SET respond more to instructor gender and students' grade expectations than to teaching quality. Confusing the two — treating a consistent number as a true one — is the central measurement error in how universities use evaluations. A bathroom scale that always reads five kilograms heavy is perfectly reliable and perfectly wrong.
Reliability and validity are not the same thing
In measurement theory, reliability is consistency: would you get the same result on repeat measurement? Validity is accuracy: does the instrument measure the thing it claims to measure? The two are independent. A measure can be highly reliable and completely invalid — the miscalibrated scale, a clock that is consistently ten minutes fast, a thermometer that always reads two degrees low. Consistency tells you the instrument is stable; it tells you nothing about whether it is pointed at the right target.
This matters because the case for SET leans heavily on reliability evidence, and the case against rests on validity evidence — and the two sides are often talking past each other. When defenders say SET are "among the most studied and psychometrically sound measures in higher education," they are usually citing reliability: high inter-rater agreement, stability across administrations, sensible factor structure. All true. None of it establishes that a higher score means better teaching.
The reliability case is genuinely strong
Give it its due. Decades of research — much of it by Herbert Marsh using the SEEQ instrument — show that well-constructed SET are multidimensional, reasonably stable across time, and consistent across cohorts. Students in the same class tend to agree; the same instructor tends to receive similar ratings year to year. By the standards of psychometric reliability, good SET instruments perform respectably. If reliability were sufficient, the debate would be over.
It is not sufficient, because reliability is satisfied by any stable signal — including a stable bias. If students consistently rate male instructors higher, or consistently reward easy grading, those biases will show up reliably, term after term. High reliability in the presence of systematic bias does not rescue validity; it entrenches it. A consistent thumb on the scale is still a thumb on the scale.
The validity case is where it falls apart
Validity asks the question that matters: do higher SET scores track better teaching, as measured by what students actually learn? The most comprehensive answer comes from Uttl, White and Gonzalez (2017), a meta-analysis of the multisection studies that are the gold standard for this question (same course, multiple sections, common final exam). Re-analysing the prior literature with proper attention to sample size, they found that SET ratings are not related to student learning — earlier positive correlations were artefacts of small studies, and once that is corrected, SET explain at most around 1% of the variance in learning.
Why did the old consensus say otherwise? Because the apparent SET-learning correlation shrinks as study quality rises — a pattern consistent with small-study bias rather than a true effect. The headline "students learn more from highly-rated teachers" rested largely on underpowered studies.
Worse, SET are demonstrably sensitive to things that are not teaching. Boring, Ottoboni and Stark (2016) showed, in a natural experiment, that SET respond more strongly to instructor gender and to students' grade expectations than to teaching effectiveness — and that the gender bias can be large enough to make a more effective instructor score lower than a less effective one. When an instrument moves reliably in response to the wrong variables, its consistency is not reassuring; it is the problem.
But surely students can judge their own experience?
Here is the strongest steel-man of SET, and it is partly right. Students are genuinely expert witnesses to their own experience: whether the lectures were clear to them, whether feedback was timely, whether they felt supported, whether the workload was manageable. On these experiential constructs, SET have a reasonable claim to validity — students are the only people who can report them.
The fallacy is the leap from "students validly report their experience" to "SET validly measure teaching effectiveness." Those are different constructs. Perceived clarity is not learning. Satisfaction is not effectiveness. A demanding course that produces deep learning may feel harder and rate lower; an entertaining, undemanding course may feel great and rate high. The Dr. Fox effect — where an expressive lecturer delivering little content earns high ratings — is the canonical demonstration that students can rate engagement they enjoyed without learning much. SET are valid for what students can see; they are invalid when treated as a proxy for what they cannot.
This is the resolution: SET are a reliable, partly-valid measure of the student experience and an invalid measure of teaching effectiveness. The error is not collecting them — it is mislabelling the construct and then attaching tenure decisions to it. (We have written separately on why student evaluations should not, on their own, decide tenure and promotion.)
What this means for practice
Three implications follow:
- Stop using SET as a standalone measure of teaching effectiveness. They cannot bear that weight. Triangulate with peer observation, teaching portfolios, and learning evidence — student ratings are necessary input, not sufficient proof.
- Measure the construct you actually have. SET are a window onto student experience. Design the instrument to capture that well — clarity, support, workload, belonging — rather than pretending it measures pedagogical quality in the abstract.
- Attack the bias, since reliability won't. Because consistent bias passes the reliability test untouched, validity has to be defended at the level of instrument design and analysis: bias-aware question design, attention to grade-expectation and demographic confounds, and reporting that resists over-precision.
This is where the format of the instrument matters. A Likert form is reliable precisely because it is crude — a 1–5 scale is easy to reproduce and easy to bias. Koji for Education is built around the experiential constructs students can validly report, using AI-moderated conversational interviews that probe what a number cannot: why a course felt unclear, what specifically helped learning, whether a low rating reflects the teaching or the grade expectation. Its standardized, bias-aware AI moderation applies the same probing to every student — removing the human-moderator inconsistency that erodes reliability — while its thematic analysis and quality scoring surface the difference between "I enjoyed this" and "I learned from this." Koji does not claim to make student ratings a valid measure of effectiveness — no instrument can, because students cannot observe effectiveness directly. It claims to measure the experiential reality students can report, accurately and at scale, and to keep that evidence in its proper place. (The same conversational interview engine powers experience research beyond the classroom on the main Koji platform.)
A reliable measure of the wrong thing is worse than an honest measure of the right thing, because reliability lends it false authority. The first step to using course evaluations well is admitting what they are valid for — and what they are not.
Key takeaways
- Reliability (consistency) and validity (accuracy) are independent; a measure can be perfectly reliable and still measure the wrong thing.
- SET are genuinely reliable - stable, internally coherent, consistent across cohorts - which is why institutions over-trust them.
- On validity, the strongest evidence (Uttl et al., 2017) finds SET essentially unrelated to learning, and SET respond more to gender and grade expectations than to teaching (Boring et al., 2016).
- Consistent bias passes the reliability test, so high reliability can entrench bias rather than reassure.
- SET are valid for student experience, invalid as a proxy for teaching effectiveness; measure the construct you actually have, triangulate the rest, and attack bias at the level of design.