Before or After the Grade? Why the Timing of Course Evaluation Is a Validity Decision, Not an Admin One
When you collect course evaluations — mid-term, in the last lecture, after the final exam, or once grades are posted — changes who responds and what they say. The choice looks like scheduling. It is really a decision about what your data can and cannot mean.
Koji for Education
Editorial Team · June 30, 2026
Answer up front: When you collect a course evaluation is a measurement decision, not a logistics one. Collect in the last lecture and your sample skews toward the students who still attend — generally higher achievers — inflating scores through non-response. Collect after the final exam or after grades are posted and you risk grade-driven reactions, where students who did badly mark the teaching down. The research does not support a single universally "correct" moment, but it is clear that timing changes both the composition of your respondents and the considerations driving their answers. Treating it as an afterthought quietly undermines comparability across courses that schedule evaluation differently.
Three things move when timing moves
Most evaluation policies pick a window — "open in the last two teaching weeks, close before exams" is common — and never revisit it. But three distinct things shift with that window.
1. Who responds (the sample). Robert Frisbie's classic study on the effect of timing on the validity of student ratings found no large difference in mean ratings between administering before a final exam, after it, or by mail once grades were known. But it found something more important: students present on the last days of class were higher achievers who gave higher ratings than the absentees. Evaluate in the final lecture and you systematically over-sample the satisfied, motivated students who still show up — a selection effect that no amount of careful question design corrects. This is the same non-response problem that haunts every voluntary evaluation, but timing controls how severe it is.
2. What students know (the information set). A student evaluating in week 10 has not yet sat the exam, received the final assignment back, or seen a grade. A student evaluating after results are posted has. The judgement they are forming is literally about a different object: "the course as I am experiencing it" versus "the course as it turned out for me." Neither is wrong, but they are not interchangeable, and aggregating them as if they were is a category error.
3. What drives the answer (the reaction). This is the contested one. A study in Studies in Educational Evaluation asking whether examinations influence student evaluations found that among students who passed, pre- and post-exam ratings barely differed — but among students who failed, post-exam scores dropped significantly on several items. In other words, the timing effect is not uniform; it is concentrated in the students who had a bad outcome and now have a grade to attribute it to.
The grade-attribution problem
The deeper issue behind post-grade timing is attribution. When a student receives a disappointing grade, the fundamental-attribution machinery makes "the teaching was poor" a more comfortable explanation than "I underprepared." Evaluating after grades are released hands every disappointed student a fresh, emotionally salient reason to mark the course down — and hands every pleased student a reason to mark it up. This is adjacent to, but distinct from, the well-documented grading-leniency association, where the expectation of a generous grade lifts ratings. Reviews consistently report a small-to-moderate positive correlation between the grades students expect and the ratings they give; post-grade timing converts that expectation into a confirmed fact, which can sharpen the reaction — especially at the negative tail.
It is precisely this risk that leads many institutions to a deliberate compromise. Stanford, for example, collects end-of-term feedback before final grades are assigned, and withholds results from instructors until after grades have been released — a two-sided firewall designed to prevent both grade-driven student retaliation and any suspicion that an instructor's marking was influenced by the feedback. The logic is sound: separate the act of evaluating from the knowledge of the grade, in both directions.
So when should you collect?
The honest answer is that it depends on what you want the data to do, and a recent review in Assessment & Evaluation in Higher Education on whether to survey before or after the final examination reaches the same nuanced conclusion rather than a slogan. A defensible framework:
- For summative, comparative ratings (the numbers that feed dashboards, benchmarking, and personnel decisions): collect before the exam and before grades are posted, to keep grade-attribution reactions out of the signal. Accept that you are measuring the experience, not the outcome.
- For formative improvement (what should change next time): the most useful moment is often mid-cycle, not the end at all. A formative, mid-semester check lets the current cohort benefit from changes, and it captures impressions before end-of-term fatigue and grade anxiety dominate. We have argued separately that in-the-moment experience sampling beats retrospective end-of-term recall for exactly this reason.
- Never mix timings within a comparison set. If half your programmes evaluate pre-exam and half post-grade, their averages are not comparable, and any league table you build from them encodes a timing artefact as if it were teaching quality.
"But isn't post-grade feedback actually more honest?"
This is the strongest counterargument, and it deserves a fair hearing. The case for waiting until after grades is that students then have the complete picture: they know whether the assessment was fair, whether the course prepared them, whether the promised learning materialised in their performance. A mid-term rating is arguably premature — the student is reviewing a film they have only half-watched. On this view, post-grade reactions are not "bias" but information: a student who failed and blames poor teaching might be right.
There is real force here, and it is why the answer is not "always pre-grade." Sometimes the failing students are the most important signal in the dataset — they may be the canaries for genuine problems with pace, scaffolding, or assessment design. The mistake is not collecting post-grade feedback; it is collecting a single number post-grade and treating it as an unconfounded measure of teaching quality. A score that blends "how good was the teaching" with "how did I do" cannot be cleanly interpreted as either. The resolution is not to pick a side but to separate the constructs: ask about the experience and the outcome as distinct things, and analyse disappointed students' feedback for its content rather than letting their lower numbers silently drag down a mean.
It is also worth keeping the magnitude honest. Frisbie's finding that mean ratings moved little with timing is a useful corrective against panic: timing is not the largest source of distortion in course evaluation. But the effect is real, it is concentrated where it does the most reputational damage (the negative tail), and it interacts with non-response in ways that a single average conceals.
How conversational evaluation changes the calculus
Most of the timing dilemma is forced by the one-shot, end-of-term survey: because you get a single snapshot, you have to gamble it on one moment and live with whatever confound that moment carries. An AI-moderated, conversational approach loosens that constraint in two ways.
First, timing becomes cheap and repeatable. Koji for Education is built for formative, mid-cycle collection as well as end-of-term — so you can run a light-touch interview in week 5 and again at the end, and watch how impressions evolve, rather than betting everything on a single post-exam survey. Second, and more importantly, the conversation can disentangle the confound that a Likert item fuses together. When a student signals dissatisfaction, the AI moderator can probe why — distinguishing "I struggled with the assessment and I think the teaching let me down" from "I did fine but the lectures were disorganised." Koji's automatic thematic analysis then surfaces grade-attribution patterns as a visible theme instead of an invisible bias buried in the average. Standardised AI moderation means this probing is consistent across every interview, with no human-moderator drift.
Koji is deliberately modest about the claim: changing how you collect does not abolish the fact that a student who failed feels differently about a course. It surfaces and contextualises that reaction so committees can weigh it, rather than letting it masquerade as an objective rating. The same adaptive interview engine underpins the main Koji platform for teams running customer and user research, where the timing of feedback relative to an experience is an equally live problem.
The takeaway
Timing is not a calendar question handed to an administrator. It determines who is in your sample, what they know, and what is driving their answer. Decide it on purpose: pre-grade for clean comparative ratings, mid-cycle for improvement, and never blend timings inside a comparison. And whenever you can, collect feedback rich enough to separate the experience from the outcome — because a single end-of-term number, taken after the grades go out, is one of the easiest figures in higher education to misread.
Koji for Education supports formative mid-cycle and end-of-term evaluation with AI moderation that probes the reasons behind a rating. Explore the platform.