Should You Run Course Evaluations Before or After the Final Exam? The Timing Evidence
A research-grounded look at whether course evaluations should be administered before or after the final examination, anchored to Vehovar and Strlekar (2024), Arnold (2009), and Overall and Marsh (1980), with practical guidance for European quality assurance.
Koji Education Team
Product
Quick answer: Administer the main course-evaluation survey before the final examination, ideally once 70-85% of teaching is complete. The largest controlled study to date (Vehovar & Strlekar, 2024; 5,077 students answering the same items before and after their exam) found that post-exam surveys yield slightly higher average scores (about 0.06 points on a 1-5 scale, p < 0.001) and weaker test-retest reliability, because a post-exam rating is partly a reaction to how the student feels about their result. If a programme genuinely needs end-of-term data, the evidence favours a "before-plus-short-after" design over a single post-exam survey.
Why timing is a measurement decision, not an administrative convenience
Most universities treat the question "when do we open the evaluation window?" as a logistics issue, slotted around exam timetables and registry deadlines. It is actually a measurement decision that changes what the numbers mean. A student who rates a course the week before the exam is reporting on the teaching. A student who rates it the day after results are released is, to some degree, also reporting on the outcome — their grade, their relief or disappointment, and their post-hoc rationalisation of the effort the course demanded. Those are different constructs, and conflating them undermines the validity of any comparison across cohorts, instructors, or years.
What the research says
Vehovar & Strlekar (2024) provide the most direct evidence. In a within-subjects replication at a large European university, 5,077 students answered seven course-related items both before and after their final examination. Average scores were significantly higher in the after-examination survey, but the difference was small in absolute terms — roughly 0.06 on a 1-5 scale. More importantly, the test-retest correlation between the two administrations was relatively low, indicating that individual students were not simply giving stable answers shifted by a constant. The authors conclude that the before-examination survey better reflects students attitudes toward the teaching and yields higher data quality, while recommending that institutions needing post-exam information add a brief follow-up rather than relocating the whole instrument to after the exam.
Arnold (2009) adds a crucial moderator: the effect is concentrated among students who do badly. Using nearly 3,000 observations from the Erasmus School of Economics, where students could complete the online questionnaire in a window spanning one week before to one week after the final exam, Arnold found that among students who passed, pre- and post-exam ratings barely differed. Among students who failed, evaluation scores were significantly lower after the exam on several items. Arnold reads this as evidence of a self-serving bias — disappointed students revise their judgement of the course downward — rather than deliberate "revenge" on the instructor. Either way, the practical implication is the same: post-exam windows let grade outcomes leak into teaching ratings, and they do so unevenly across the grade distribution.
Overall & Marsh (1980) supply the reassuring counterpoint that bounds how alarmed we should be. In their longitudinal study, end-of-term ratings were highly correlated with ratings collected from the same students after graduation (median r ≈ 0.83). Student evaluations are not pure noise or fleeting mood; they carry stable signal about teaching that survives years. The timing problem is therefore best understood as a contamination at the margin — a small, systematic, grade-linked distortion layered on top of an otherwise reasonably stable judgement — not as proof that the whole exercise is worthless.
Taken together, the literature points to a clear ordering of preferences: a pre-exam window is the safest default; a post-exam window introduces a small upward bias overall and a larger downward bias among failing students; and the cleanest design for institutions that want both is to collect the substantive teaching evaluation before the exam and append a short outcome-focused survey afterward.
Why it matters for course evaluation in practice
For a quality-assurance office, three consequences follow.
-
Comparability is at stake. If some departments evaluate before exams and others after — or if the same department drifts year to year — then cross-unit benchmarking compares contaminated and uncontaminated numbers. A 0.1 difference in a programme dashboard could be a timing artefact rather than a teaching difference.
-
Fairness to instructors of demanding courses. Arnold's finding means that instructors of high-stakes, high-failure modules are disproportionately exposed to post-exam downgrading. Using post-exam scores in personnel or promotion contexts quietly penalises exactly the staff who teach the hardest, most consequential courses.
-
Response-rate trade-offs. Post-exam windows often coincide with students leaving campus and disengaging, depressing response rates and worsening non-response bias on top of the timing bias. A pre-exam window, opened while students are still attending, typically captures a more representative sample.
Limitations and honest caveats
A critical reader should hold several caveats in view. First, the headline effect in Vehovar & Strlekar (2024) is small — 0.06 points — and may be immaterial for formative purposes even if it matters for high-stakes benchmarking. Second, the strongest causal evidence (Arnold, 2009) comes from a single economics faculty, and self-serving bias may be more pronounced in disciplines with sharper pass/fail stakes; generalisation to, say, reflective humanities seminars is uncertain. Third, "before the exam" is not a single moment: a survey opened immediately after the last lecture differs from one opened during a revision week, and the literature does not finely resolve the optimum within the pre-exam window. Fourth, these studies measure average and aggregate effects; an individual instructor cannot infer from them whether their own scores were timing-distorted. Finally, none of this resolves the deeper validity debate about whether student ratings measure learning at all — timing is a refinement at the margin, not a fix for the construct-validity questions raised elsewhere in this knowledge base.
How Koji incorporates this
Koji is designed to make timing an explicit, defensible parameter rather than an accident of the registry calendar.
-
Mid-cycle and formative collection by design. Koji supports opening a conversational evaluation at a configurable point in the term, so programmes can run the substantive teaching evaluation before exams as a default, in line with the evidence. Formative mid-semester rounds — which sit well before any exam — can be scheduled separately, keeping diagnostic feedback cleanly upstream of grade contamination.
-
A "before-plus-short-after" pattern, made practical. Because Koji studies are quick to configure and re-issue, the design Vehovar & Strlekar recommend — a full pre-exam evaluation plus a brief post-exam follow-up — is straightforward to operationalise without doubling student burden. The short after-survey can target outcome-sensitive questions explicitly, so any grade-linked signal is captured separately rather than silently mixed into the teaching scores.
-
Bias-aware reporting and metadata. Koji timestamps every interview and can record the administration window relative to assessment, so reports can flag when a cohort was evaluated post-exam and treat those numbers with appropriate caution in benchmarking — rather than comparing contaminated and clean cohorts as if they were equivalent.
-
Probing beyond the grade reaction. Because Koji uses AI-moderated conversational interviews (with structured question types — open_ended, scale, single_choice, multiple_choice, ranking, yes_no) rather than a static Likert grid, follow-up questions can distinguish "I am unhappy with my mark" from "the teaching was unclear." Automatic thematic analysis of the open-text responses then surfaces whether negative sentiment clusters around assessment fairness versus instruction quality, helping committees separate the two even when timing is imperfect. This is designed to mitigate — not eliminate — self-serving bias.
Koji's core research platform at koji.so applies the same AI-moderated interview engine to product and customer research, where survey timing relative to a purchase or support event raises closely analogous contamination questions.
Frequently asked questions
When exactly should the pre-exam window open? A common, evidence-consistent practice is to open it once roughly 70-85% of teaching is complete and close it before study/revision week, so students have experienced most of the course but have not yet sat the exam.
Is the post-exam bias large enough to worry about? For formative use, usually not — 0.06 points is small. For high-stakes benchmarking, ranking, or personnel decisions, yes, because the bias is systematic and falls hardest on instructors of high-failure courses.
Do failing students "take revenge" on instructors after the exam? The evidence (Arnold, 2009) is more consistent with self-serving bias — disappointed students revise their judgement of the course downward — than with deliberate retaliation. The practical risk to the data is the same either way.
Should we just move everything to after the exam so students have seen the whole course? No. The marginal information gained is outweighed by grade contamination, lower reliability, and falling response rates. Prefer a full pre-exam evaluation plus an optional short post-exam follow-up.
Does timing affect open-text comments as well as scores? Likely yes — post-exam comments are more prone to referencing grades and assessment. Thematic analysis that separates assessment-related themes from teaching-related themes helps keep the distinction visible.
References
- Vehovar, V., & Strlekar, L. (2024). When to conduct student evaluation of teaching surveys: before or after the final examination? Assessment & Evaluation in Higher Education, 49(6), 767-780. https://doi.org/10.1080/02602938.2023.2298771
- Arnold, I. J. M. (2009). Do examinations influence student evaluations? International Journal of Educational Research, 48(4), 215-224. https://doi.org/10.1016/j.ijer.2009.09.001
- Overall, J. U., & Marsh, H. W. (1980). Students evaluations of instruction: A longitudinal study of their stability. Journal of Educational Psychology, 72(3), 321-325. https://doi.org/10.1037/0022-0663.72.3.321
Related resources
Related articles
What Actually Raises Course-Evaluation Response Rates? The Experimental Evidence
Online course evaluations chronically under-perform paper. We review the experimental evidence — Dommeyer''s grade-incentive trials and Nulty''s adequacy thresholds — on what genuinely lifts response rates, what it costs in data quality, and how to hit a defensible rate without coercion.
Selection Bias in Course Evaluations: What Goos and Salomons Found
A research-grounded reading of Goos & Salomons (2017) on selection bias in online course evaluations, with practical implications for response-rate policy and reporting at European universities.
Does Grading Leniency Inflate Student Evaluations? What the Evidence Actually Shows
A research-grounded look at the grading-leniency hypothesis in student evaluations of teaching: what Marsh and Roche, Greenwald and Gillmore, and Stroebe found, why the grades-ratings correlation is genuinely ambiguous, and how to design evaluation so leniency cannot quietly buy better scores.
Online vs. Paper Course Evaluations: Do Lower Response Rates Mean Worse Data?
Online course evaluations consistently draw lower response rates than in-class paper forms, but the research shows the resulting scores are largely equivalent. Here is what Dommeyer, Nulty, and Stowell actually found, and what an adequate response rate really requires.