Are Student Ratings Just a Mood? The Longitudinal Stability Evidence
A common objection to course evaluations is that students cannot judge teaching until years later. Overall and Marsh (1980) tested this directly by re-surveying the same students after at least a year. We review the stability evidence, its limits, and what it means for how feedback should be timed and used.
Koji Education Team
Product
In brief: One of the oldest objections to student evaluations is that end-of-course ratings reflect a fleeting mood and that students can only judge teaching with hindsight. Overall and Marsh (1980) tested this by re-surveying the same students at least a year after the course: retrospective ratings correlated strongly with the original end-of-term ratings (class-average stability coefficients around r = 0.8). Students' summative judgements of a course are not a transient mood — but stability is not the same as validity, and it does not rescue ratings from grade or bias effects baked in at both time points.
The question this answers
Faculty resistance to student evaluations often rests on an intuition: "Students don't know what was good for them until much later. Ask them at the end of term and you measure how they felt, not how well they were taught." If that intuition were correct, end-of-course ratings would be near-worthless — a snapshot of exam-week emotion that bears little relation to a considered judgement. Overall and Marsh (1980) put the intuition to a direct empirical test, and the answer has shaped how the measurement community thinks about timing ever since.
What the research says
The anchor study is J. U. Overall and Herbert W. Marsh (1980), Students' Evaluations of Instruction: A Longitudinal Study of Their Stability, Journal of Educational Psychology, 72(3), 321–325. The design is simple and powerful: students who had evaluated a course at the end of term were asked to evaluate the same course and instructor again at least one year later — in some cases as alumni, after they had moved on and could reflect with the benefit of distance and subsequent coursework.
The key findings:
- High stability. End-of-term and retrospective ratings were strongly and significantly correlated, whether analysed at the level of individual students or class averages. Class-average stability coefficients were high (in the region of r = 0.8), indicating that a course rated highly at the end of term was still rated highly a year or more later.
- No systematic reversal. Retrospective judgements did not contradict the original ones. Students did not, on reflection, decide that the demanding course they disliked at the time had actually been excellent. The rank order of courses was largely preserved.
- The "wait and see" critique fails on its own terms. If end-of-course ratings were dominated by transient mood or by not-yet-appreciating-the-rigour, retrospective ratings would diverge sharply. They do not.
This result is consistent with the broader Marsh research programme. Marsh (1984), Students' Evaluations of University Teaching: Dimensionality, Reliability, Validity, Potential Biases, and Utility (Journal of Educational Psychology, 76(5), 707–754), synthesised dozens of studies and concluded that well-constructed multidimensional ratings (the SEEQ) are reliable and reasonably stable. Marsh and Roche (1997), in American Psychologist (52(11), 1187–1197), reaffirmed that ratings show good stability and that the multidimensional structure of teaching is preserved across time and setting. Later replications of mean-rating stability over long periods (e.g. multi-year tracking of the same instructors) reach the same conclusion: aggregate ratings of a given teacher are remarkably steady.
Why it matters for course evaluation in practice
For a quality office, the stability evidence carries three practical messages:
- End-of-course timing is defensible. You do not need to wait until students graduate to get a considered judgement. The summative end-of-term evaluation captures something durable, not just exam-week relief or frustration. This matters for the feasibility of any evaluation cycle — you can act on the data now.
- Stability strengthens the case for acting on feedback. If ratings were noise, closing the feedback loop would be theatre. Because the signal persists, a course flagged as weak is likely to be judged weak by the next cohort too — so intervention is warranted. See closing the feedback loop.
- Stability is necessary but not sufficient. A reliable measure can still be reliably wrong. If a grade-expectation effect or a gender bias is present at the end of term, it is plausibly present in the retrospective rating too — stability would then preserve the bias, not remove it. Stability answers the "mood" objection; it does not answer the "validity" objection, which is a separate question addressed by whether ratings measure learning.
The distinction between reliability and validity is the crux. Overall and Marsh show ratings are reliable over time. Whether they are valid — whether they track teaching quality rather than charisma, leniency, or first impressions — is a different question, and the honest answer there is more qualified.
Limitations and honest caveats
A critical reader should hold several reservations:
- Stability can come from shared bias, not shared truth. If both the original and retrospective ratings are coloured by the same grade, the same halo, or the same thin-slice first impression, the high correlation reflects a stable bias, not a stable measurement of teaching. Consistency over time is reassuring about reliability and silent about validity.
- Memory reconstruction. Retrospective ratings are not independent observations; a year later, students may partly recall (or reconstruct) the judgement they already made, inflating the apparent stability. A truly independent second measurement is hard to obtain.
- Era and generalisability. The anchor study is from 1980, in a North American setting, using the SEEQ instrument. Stability of well-constructed multidimensional instruments may not transfer to the short, ad-hoc global-satisfaction forms many institutions still use. The result is strongest for instruments built to Marsh's standards.
- Attrition. Students who respond to a follow-up a year later may be unrepresentative (more engaged, more positive), which can bias stability estimates.
These caveats do not overturn the finding — they bound it. The reasonable conclusion is: summative student judgements of a course are stable over time; treat them as a real signal, but do not mistake their durability for proof that they measure teaching effectiveness.
How Koji incorporates this
Koji for Education is designed to take the stability evidence seriously without overstating it — to capture the durable signal while continuously testing whether that signal is tracking learning or tracking bias.
- Multidimensional, structured capture rather than a single global item. Marsh's stability results are strongest for multidimensional instruments. Koji's evaluations use structured question types (
open_ended,scale,single_choice,multiple_choice,ranking,yes_no) to separate distinct dimensions — clarity, organisation, workload, assessment fairness, perceived learning — rather than collapsing everything into one rating whose stability is hard to interpret. - AI-moderated conversational interviews that probe beyond the number. Because a stable number can hide a stable bias, Koji's conversational interviews ask students why they rated a course as they did and what specifically helped or hindered learning. This produces evidence that a committee can scrutinise for validity, not just consistency.
- Mid-cycle and end-of-cycle collection. Stability across the term means formative mid-semester feedback genuinely forecasts the end-of-course judgement, so Koji's mid-cycle collection lets instructors intervene while it still helps the current cohort, rather than waiting for a verdict they cannot change.
- Triangulation and trend reporting. Koji aggregates results across cohorts and over time, so a quality office can distinguish a stable course-level signal (which warrants action) from year-to-year noise — and can watch whether an intervention actually shifts the durable pattern.
- Bias-aware reporting. Because stability preserves bias, Koji's reporting is designed to surface distributions and themes and to flag patterns consistent with known biases, supporting the principle that a consistent score still should not be used as a stand-alone high-stakes metric.
Koji is designed to mitigate the risk that a stable rating is mistaken for a valid one; it does not claim to certify validity, which always requires triangulation with peer review, learning evidence, and programme outcomes. The same engine powers general user and customer research at koji.so, where the reliability-versus-validity distinction is equally central to trustworthy insight.
The practical takeaway for a quality office
Treat the stability evidence as a permission and a warning at once. The permission: you may act on end-of-course data now, without apologising that students "cannot really know yet" — Overall and Marsh settled that the considered judgement and the end-of-term judgement largely coincide. The warning: never let a department cite stability as proof of validity. The two questions are independent, and a course that scores consistently well may be consistently benefiting from charisma, a generous grade, or a strong first impression rather than from teaching that produces learning. The disciplined posture is to use durable course-level signals to prioritise where to look, then triangulate with peer observation, assessment evidence, and programme outcomes before attaching any consequence to an instructor. Stability tells you the signal is real; only triangulation tells you what the signal is.
Related resources
- What Do Student Evaluations Actually Measure? Marsh and the SEEQ
- Generalizability Theory and the Reliability of Student Ratings
- Do Student Evaluations Measure Learning? The Uttl Meta-Analysis
- Do First Impressions Decide Your Course Evaluation? The Thin-Slice Evidence
- Interpreting and Reporting Student Ratings Responsibly
References
- Overall, J. U., & Marsh, H. W. (1980). Students' evaluations of instruction: A longitudinal study of their stability. Journal of Educational Psychology, 72(3), 321–325. https://doi.org/10.1037/0022-0663.72.3.321
- Marsh, H. W. (1984). Students' evaluations of university teaching: Dimensionality, reliability, validity, potential biases, and utility. Journal of Educational Psychology, 76(5), 707–754. https://doi.org/10.1037/0022-0663.76.5.707
- Marsh, H. W., & Roche, L. A. (1997). Making students' evaluations of teaching effectiveness effective: The critical issues of validity, bias, and utility. American Psychologist, 52(11), 1187–1197. https://doi.org/10.1037/0003-066X.52.11.1187
- Marsh, H. W. (2007). Students' evaluations of university teaching: Dimensionality, reliability, validity, potential biases and usefulness. In R. P. Perry & J. C. Smart (Eds.), The Scholarship of Teaching and Learning in Higher Education (pp. 319–383). Springer. https://doi.org/10.1007/1-4020-5742-3_9
Related articles
Generalizability Theory and the Reliability of Student Ratings: Why One Class Is Not Enough
Reliability is not one number. Generalizability theory (Gillmore, Kane & Naccarato 1978; Marsh 1984) decomposes the variance in student ratings into student, teacher, course and occasion components — and shows that a single class can be reliable for the course yet a poor estimate of the teacher. What that means for fair evaluation.
Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees
The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.
Do Student Evaluations Measure Learning? The Uttl Meta-Analysis Revisited
A research-grounded reading of the Uttl, White & Gonzalez (2017) meta-analysis on the SET–learning relationship, with implications for European course evaluation and quality-assurance policy.
What Do Student Evaluations Actually Measure? Marsh, the SEEQ, and the Case for Multidimensional Feedback
Herbert Marsh spent decades showing that well-built student evaluations are multidimensional, reliable, and stable — and that a single global score throws away most of what they can tell you. What his work establishes, where critics push back, and how to design feedback that is actually usable.