Formative vs Summative Course Evaluation: The Mid-Cycle Feedback Most Universities Skip
Summative end-of-term surveys arrive too late to help the students who completed them. Formative, mid-cycle evaluation is the only kind with strong causal evidence that it improves teaching - yet most institutions still pour their effort into the version that cannot. Here is the distinction that matters and why it should reshape your evaluation calendar.
Koji for Education
Research & Editorial Team ·
Bottom line up front: Summative course evaluation (the end-of-term survey) is designed to judge a course after it is over; formative evaluation (mid-cycle feedback) is designed to improve it while there is still time to act. The evidence is lopsided: a classic meta-analysis found that giving instructors mid-semester feedback measurably raised their end-of-term ratings, and the effect grew when feedback was paired with consultation. End-of-term surveys, by contrast, have almost no demonstrated effect on the learning of the students who fill them in. If your institution invests heavily in summative data and treats formative feedback as optional, you are optimising the instrument with the weaker evidence base.
Two purposes that get conflated
Most universities run one evaluation event — a survey in the final weeks of term — and ask it to do two incompatible jobs. The summative job is accountability: a number for the programme review, the promotion file, the accreditation self-assessment. The formative job is development: actionable insight a teacher can use to change something. A single end-of-cycle instrument does the first job adequately and the second job not at all, because the feedback arrives after the course has ended. The students who supplied it have already left; the only people who could benefit are next year's cohort, and only if the lecturer remembers and acts.
The terms come from programme evaluation more broadly: Michael Scriven's distinction between formative evaluation (improving a programme during development) and summative evaluation (judging its overall worth) maps directly onto teaching. Formative is feedback for learning and teaching; summative is a verdict on it. Conflating them is how institutions end up with mountains of summative data and a chronic complaint from staff that "the evaluations never tell me anything I can use."
The evidence strongly favours formative feedback
Here is the part that should reorder priorities. The canonical evidence comes from Cohen's (1980) meta-analysis of intervention studies in which instructors received student ratings partway through a course. Across the studies, instructors who got mid-semester feedback ended the term with higher overall ratings than those who did not — a difference of roughly a fifth of a standard deviation, corresponding to a meaningful percentile gain. Crucially, the effect was substantially larger when the ratings were accompanied by consultation — when someone helped the instructor interpret the feedback and plan changes, rather than just handing over numbers.
More recent syntheses reach compatible conclusions. A meta-analysis of student-feedback intervention studies confirms that structured feedback to teachers during a course can improve teaching and class outcomes, with consultation amplifying the benefit. The mechanism is unsurprising: feedback you receive while you can still act on it changes behaviour; feedback you receive after the fact mostly produces a score.
Contrast this with the summative case. The best evidence on end-of-term ratings as a measure of effectiveness is sobering — Uttl, White and Gonzalez (2017) found SET ratings essentially unrelated to how much students learn. Summative scores are a backward-looking judgement of uncertain validity; formative feedback is a forward-looking intervention with demonstrated effect. That asymmetry is the whole argument.
Why mid-cycle feedback is also better feedback
Beyond timing, formative evaluation tends to produce richer data. Mid-course, students are still inside the experience; they remember the specific lecture that lost them and the assignment that finally made a concept click. End-of-term, those details have compressed into a vague global impression — which is exactly the condition under which halo effects, grade expectations, and recency dominate the rating. Asking "what should change for the rest of this term?" elicits concrete, fixable items. Asking "rate this course 1–5" at the end elicits a summary judgement contaminated by everything from the exam difficulty to the weather.
Formative feedback also rebuilds trust. When students see a mid-course change made in response to their input — "you said the pace was too fast, so here is a revised schedule" — they learn that feedback is acted on. That visible responsiveness is one of the few reliable ways to lift the chronically low response rates that plague end-of-term surveys, and it shifts students from passive respondents toward partners in the course.
But doesn't formative feedback just inflate end-of-term scores?
This is the strongest objection, and it deserves a straight answer. If mid-semester feedback raises end-of-term ratings, is that real improvement — or has the instructor simply learned to please the same students who will rate them, perhaps by easing demands? The honest answer: the Cohen effect cannot, by itself, fully distinguish genuine pedagogical improvement from rating-management. Some of the gain may be relational — students rate a teacher more warmly for having listened.
But three things blunt the objection. First, the effect is largest when paired with consultation focused on teaching practice, not on placating students — the improvement is mediated by deliberate pedagogical change, not mere responsiveness. Second, formative feedback should never be summative: mid-cycle data exists to inform the teacher, not to be scored or compared, which removes the incentive to game it. Third, even if part of the effect is relational, a teacher who listens and adjusts is, on most accounts, teaching better — responsiveness is a pedagogical virtue, not a confound to be scrubbed out. The risk of gaming is real but it is an argument for keeping formative feedback developmental and confidential, not for abandoning it.
What this means for your evaluation calendar
The practical implication is not to abolish summative evaluation — accreditation and quality assurance require it (see the European Standards and Guidelines expectations around acting on student feedback). It is to stop letting the summative survey crowd out the formative one. Concretely:
- Run a genuine mid-cycle evaluation in every substantial module, early enough that changes are still possible (typically around the one-third to halfway point).
- Keep formative data developmental and confidential — out of the promotion file — so it stays honest.
- Pair feedback with consultation wherever you can; the evidence says this is where the gains come from.
- Close the loop visibly — tell students what changed because of what they said.
This is where the design of the instrument matters. A static Likert form repeated at midterm is still a static Likert form. Koji for Education runs formative, mid-cycle collection through AI-moderated conversational interviews that ask students why and follow up in the moment — surfacing the specific, fixable issues that a 1–5 scale flattens away. Its automatic thematic analysis turns dozens of open-ended mid-course responses into a ranked list of what to change this term, and its closing-the-loop action tracking records what the teacher did about it. Because the moderation is standardized AI rather than an inconsistent human facilitator, every cohort gets the same probing quality — and because formative results are kept separate from summative reporting, the data stays developmental. (The same conversational interview engine underpins iterative product and customer research on the main Koji platform.)
Universities have spent decades perfecting the survey that judges teaching and neglecting the feedback that improves it. The evidence says the priorities are backwards.
One organisational caveat is worth naming. Formative evaluation only works if teachers are given the time and support to act on it; bolting a mid-cycle survey onto an unchanged workload, with no consultation and no slack in the schedule to revise anything, reproduces the summative problem in a new slot — data collected, nothing changed. The evidence is not that collecting mid-cycle feedback helps; it is that acting on it, with support helps. Institutions that want the Cohen effect have to fund the consultation, not just the survey.
Key takeaways
- Formative (mid-cycle) evaluation exists to improve a course in progress; summative (end-of-term) evaluation exists to judge it afterward — one instrument cannot do both well.
- Cohen's (1980) meta-analysis shows mid-semester feedback measurably raises end-of-term ratings, with larger gains when paired with consultation; end-of-term SET, by contrast, is essentially unrelated to learning.
- Mid-course feedback is richer because students still remember specifics, and visible mid-course changes rebuild trust and response rates.
- The "it just inflates scores" objection is real but limited: keep formative data confidential and developmental, and pair it with pedagogical consultation.
- Don't abolish summative evaluation — stop letting it crowd out the formative feedback that actually changes teaching.