Do Student Evaluations Actually Improve Teaching? The Feedback-Intervention Evidence
Student ratings handed back as a bare number barely change teaching (d ≈ 0.2). Paired with structured consultation, the same data produces moderate, durable improvement (d ≈ 0.6–0.7). What the intervention meta-analyses tell quality teams about closing the loop.
Koji Education Team
Product
In brief: Student evaluations can improve teaching, but only when the data is acted on, not merely collected. Cohen's (1980) meta-analysis found that simply returning ratings to instructors produced a small end-of-term improvement (d ≈ 0.20); when ratings were paired with personal consultation, the effect roughly tripled to d ≈ 0.64. Penny and Coe (2004) confirmed a moderate-to-large effect (d ≈ 0.69) for consultation-augmented feedback, and Marsh and Roche (1993) showed an individually targeted intervention raised teaching on the very dimensions instructors chose to work on. The decisive variable is not the survey — it is what happens after.
Universities spend enormous effort collecting course evaluations and remarkably little ensuring the results change anything. The implicit theory is that handing an instructor their scores will prompt reflection and improvement. The evidence says that theory is half right: feedback alone nudges teaching slightly, but the return on a bare number is small. The improvement that quality frameworks actually want — visible, sustained gains in teaching — appears only when feedback is wrapped in a structured, supportive process.
What the research says
The foundational synthesis is Cohen (1980), "Effectiveness of student-rating feedback for improving college instruction," published in Research in Higher Education. Cohen meta-analysed controlled intervention studies in which instructors were randomly assigned either to receive their mid-term student ratings or not, with end-of-term ratings as the outcome. Instructors who simply received their ratings improved by about d = 0.20 of a standard deviation on end-of-term scores — a real but modest effect. The pivotal finding was the moderator: when the feedback was accompanied by consultation — a trained colleague or developer helping the instructor interpret the data and plan changes — the average effect rose to roughly d = 0.64. In other words, the same information was three times as powerful when someone helped the instructor act on it.
Penny and Coe (2004), published in Review of Educational Research, focused specifically on the consultation ingredient. Their exploratory meta-analysis of intervention studies found an average effect of about d = 0.69 for student-ratings feedback augmented by consultation — squarely in the moderate-to-large range. Critically, they showed the approaches are not interchangeable: effectiveness depended on what the consultation actually did. Strategies that engaged the instructor actively (helping them interpret data, set concrete goals, examine teaching beliefs, and review progress) outperformed passive hand-over of a report. Consultation is a craft, not a checkbox.
Marsh and Roche (1993), in the American Educational Research Journal, ran one of the cleanest individual experiments. Ninety-two teachers were evaluated using a multidimensional instrument (an Australian version of the SEEQ), and randomly assigned groups received a feedback-plus-consultation intervention at different points, with a no-intervention control. Each teacher "targeted" specific dimensions of their teaching for improvement. The intervention produced gains, and — importantly for validity — the gains were concentrated on the dimensions teachers had chosen to work on, exactly what a genuine improvement effect (rather than a generic halo) should look like. This also leans on the multidimensional view of what ratings measure, developed in our note on Marsh's multidimensionality.
A consistent theme across newer syntheses of student-feedback intervention studies is the same shape: feedback is necessary but weak on its own, and the active ingredient is a structured process that turns data into a concrete, owned plan. This is the empirical backbone of "closing the loop," and it converges with our evidence review on closing the feedback loop and the meta-analytic case for mid-semester feedback consultation.
Why it matters for course evaluation in practice
The single most important implication for a quality-assurance office is uncomfortable: most institutional evaluation systems are optimised for the weak version of the effect. They collect ratings, generate a PDF, email it to the instructor, and file it for the next review. That workflow buys roughly the d = 0.20 outcome — the floor. The moderate-to-large gains require investment in the part almost everyone skips: helping instructors make sense of the data and plan a response.
Three practical consequences follow.
First, timing should favour formative use. The intervention studies typically used mid-term feedback so instructors could act while the same students were still in the room — closing the loop within the cycle rather than only between cycles. End-of-term summative ratings, by contrast, can only help the next cohort and arrive when motivation to change is lowest. This is the core of the formative-versus-summative distinction.
Second, consultation does not have to be expensive to be structured. Penny and Coe's results suggest the effective elements are specific behaviours — interpret the data with the instructor, identify two or three concrete changes, set goals, and follow up — not necessarily a lengthy programme. A short, well-designed conversation outperforms a long, vague one. Departments without a teaching-development unit can still build a light peer-consultation routine that captures most of the benefit.
Third, acting on feedback is itself an accreditation expectation. European quality frameworks (the ESG, and national agencies) ask not just whether feedback is collected but whether it informs improvement — the documented "you said, we did" cycle. The intervention literature gives QA teams the evidence to argue that the loop-closing step is where the educational value, and the auditable evidence of enhancement, actually lives. See our overview of student feedback as ESG accreditation evidence.
Limitations and honest caveats
The evidence is encouraging but should be read with care. Many of the primary studies are decades old and conducted in North American higher education; teaching contexts, student populations, and evaluation instruments have changed, and effect sizes from the 1970s–1990s may not transfer cleanly to a 2026 European multilingual classroom. Replication of these specific intervention effects in modern settings is thinner than the foundational citations suggest.
There is also a moderator-as-outcome subtlety. Cohen's headline contrast (0.20 vs 0.64) comes from comparing study subsets that differ in more than just consultation; the consultation studies may differ in instructor motivation, instrument quality, or institutional support. Penny and Coe's own conclusion — that consultation strategies vary widely in effectiveness — is a warning against treating "d ≈ 0.6" as a guarantee. A poorly run consultation may deliver little.
Two measurement caveats matter. The outcome in most studies is end-of-term student ratings, not an independent measure of learning. If an intervention teaches instructors to raise ratings without raising learning, the effect is partly an artefact — the same concern that runs through the SET-versus-learning literature. Marsh and Roche's targeting result is reassuring here (gains tracked chosen dimensions), but the field would benefit from more learning-based outcomes. Finally, selection and motivation confound real programmes: instructors who volunteer for consultation are precisely those most ready to change, so observational "our consultation works" claims overstate what a mandatory roll-out would achieve.
None of this undermines the core message; it bounds it. Feedback plus a genuine, well-executed, actively engaged consultation produces moderate improvement in measured teaching. Feedback alone produces little. Anyone promising large gains from a dashboard with no human follow-up is overselling.
How Koji incorporates this
Koji is built on the premise that the value of an evaluation is realised after collection, in the loop-closing step the research identifies as decisive.
- Richer, more actionable feedback to act on. Consultation works best when the instructor has something concrete to discuss. Koji's AI-moderated conversational interviews probe beyond a Likert number for specific, behaviour-level detail ("what would have helped you follow the lectures?"), and automatic thematic analysis distils open text into named themes. The result is feedback an instructor and a consultant can act on directly, rather than a column of means that invites defensiveness.
- Mid-cycle / formative collection. Because Koji studies can be run at any point, departments can collect mid-term feedback — the timing the intervention studies used — so instructors can close the loop with the same cohort, not just the next one.
- Targeted, multidimensional structure. Mirroring the Marsh and Roche "target a dimension" design, Koji's structured question types (
scale,ranking,single_choice,open_ended) let instructors focus an evaluation on the specific aspects of teaching they are trying to improve, then re-measure those same dimensions to see whether the change worked. - Action tracking that operationalises "closing the loop." Koji is designed to support recording what was changed in response to feedback and reporting it back, turning the consultation-and-action step into auditable quality-cycle evidence for ESG, NVAO, or national reviews.
- Triangulation. Because student feedback alone is one source, Koji's reporting is built to sit alongside other evidence (peer review, outcomes) rather than stand as a lone verdict — consistent with the caution that ratings predict ratings, not necessarily learning.
Koji does not claim that running a study improves teaching by itself — the research is explicit that the dashboard is the weak lever. What Koji does is make the strong lever cheaper to pull: better raw material for a consultation, the right timing, and a structured way to record and re-measure the changes. The same AI-moderated engine powers Koji's core research platform at koji.so for product and customer teams, where "collect, then actually act" is the identical discipline.
Related Resources
- Closing the Feedback Loop: The Evidence
- Mid-Semester Feedback and Consultation: A Meta-Analysis
- Formative vs Summative Course Evaluation
- What Do Student Evaluations Measure? Marsh's Multidimensionality
- Student Evaluations, Teaching, and Learning: The Meta-Analysis
- Student Feedback as ESG Accreditation Evidence
References
- Cohen, P. A. (1980). Effectiveness of student-rating feedback for improving college instruction: A meta-analysis of findings. Research in Higher Education, 13(4), 321–341. https://link.springer.com/article/10.1007/BF00976252
- Penny, A. R., & Coe, R. (2004). Effectiveness of consultation on student ratings feedback: A meta-analysis. Review of Educational Research, 74(2), 215–253. https://doi.org/10.3102/00346543074002215
- Marsh, H. W., & Roche, L. (1993). The use of students' evaluations and an individually structured intervention to enhance university teaching effectiveness. American Educational Research Journal, 30(1), 217–251. https://doi.org/10.3102/00028312030001217
Related articles
Turning Student Feedback into ESG / ENQA Accreditation Evidence
A buyer's guide mapping the ESG 2015 internal quality assurance standards to concrete, accreditation-ready evidence you can generate from student feedback — and how AI-moderated evaluation closes the loop.
Does Closing the Feedback Loop Actually Matter? The Evidence on Acting on Student Evaluations
Universities are good at collecting student feedback and bad at acting on it visibly. The research — Watson (2003), Leckey & Neill (2001), Shah et al. (2017) — shows that failing to close the loop drives the scepticism and declining response rates that quietly destroy your evaluation data.
What Do Student Evaluations Actually Measure? Marsh, the SEEQ, and the Case for Multidimensional Feedback
Herbert Marsh spent decades showing that well-built student evaluations are multidimensional, reliable, and stable — and that a single global score throws away most of what they can tell you. What his work establishes, where critics push back, and how to design feedback that is actually usable.
Do Student Evaluations Measure Learning? The Uttl Meta-Analysis Revisited
A research-grounded reading of the Uttl, White & Gonzalez (2017) meta-analysis on the SET–learning relationship, with implications for European course evaluation and quality-assurance policy.