New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Methodology9 min read

Does Course Evaluation Actually Improve Teaching? The Evidence Says: Only With Consultation

Handing an instructor their ratings produces a small effect. Handing them their ratings plus a structured conversation with a skilled colleague produces an effect roughly four times larger. Institutions spend their budget on the first half and almost nothing on the second.

Koji Education Team

Product ยท

Short answer: Yes, student feedback improves teaching - but the improvement comes overwhelmingly from what happens after the data arrives. Cohen's 1980 meta-analysis found that simply returning mid-semester ratings to instructors produced a modest gain of about 0.16 of a rating point. Penny and Coe's 2004 meta-analysis found that ratings augmented with consultation produced an average effect size of d = 0.69. The instrument is not the intervention. The conversation is.

This is the most actionable finding in the entire course-evaluation literature, and it is almost universally ignored in how institutions allocate their evaluation budget. Enormous effort goes into designing the questionnaire, chasing response rates and building dashboards. Comparatively nothing goes into the step the evidence says actually changes teaching.

What the evidence shows

Cohen (1980), in Research in Higher Education, integrated 22 comparisons of student-rating feedback at college level (Cohen, 1980). The design in these studies is typically clean: instructors are randomly assigned either to receive mid-semester ratings or not, and end-of-semester ratings are compared.

The result: instructors who received mid-semester feedback averaged 0.16 of a rating point higher on end-of-semester overall ratings than those who received none - roughly a third of a standard deviation, or about a 15-percentile gain. Cohen also reported that effects were accentuated when the ratings were augmented by consultation.

Penny and Coe (2004), in Review of Educational Research, followed that thread directly (Penny & Coe, 2004, RER 74(2): 215-253). Analysing 11 intervention studies in which student ratings feedback was combined with peer and expert consultation, they found an average effect size of d = 0.69.

The gap between those two numbers is the whole argument. Feedback alone: about d = 0.33. Feedback plus consultation: about d = 0.69. Roughly a doubling in standardised terms, from a modest effect to a substantial one - and in education research, where effects of d = 0.4 are considered respectable, d = 0.69 is a genuinely strong intervention.

Later work has broadly sustained the pattern. L'Hommedieu and colleagues (1990) and subsequent overviews of intervention studies continue to find positive but modest effects for feedback alone, with augmentation doing the heavy lifting. A 2024 hierarchical meta-analysis in Educational Assessment, Evaluation and Accountability extended the question to school settings and again found that the impact of student feedback on teaching quality depends heavily on how the feedback is processed rather than merely delivered.

Why feedback alone underperforms

Three mechanisms explain the gap, and each has a design implication.

1. Ratings say what, never why. An instructor who learns that "clarity of explanation" scored 3.1 knows there is a problem and nothing about its cause. Was the notation unexplained, the pace too fast, the prerequisite assumed? Without a diagnosis, the natural response is either a generic adjustment or none at all. A consultant's first job is to convert a score into a specific, actionable behaviour.

2. Defensive reception. Feedback that threatens professional identity tends to be discounted, especially when it is anonymous, aggregated and arrives without support. This is a well-documented dynamic and part of why faculty distrust course evaluations. A colleague who interprets the data with the instructor reframes it from verdict to evidence.

3. Timing. Cohen's effect came from mid-semester feedback, where the instructor could act on it with the same cohort. End-of-term ratings arrive after the students who gave them have gone. That is a structural argument for formative rather than purely summative evaluation - the studies that show improvement mostly used the formative timing that most institutional systems do not.

Notice how much of this is about interpretive capacity, not data volume. Collecting more feedback with no capacity to interpret it is a strictly negative-return investment: it adds student burden and survey fatigue without adding improvement.

Critics argue: the outcome measure is circular

The strongest objection is methodological, and it is serious. In most of these studies the outcome is later student ratings - the same instrument being evaluated. If feedback plus consultation raises subsequent ratings, that may mean teaching improved, or it may mean the instructor learned what students reward. Given evidence that ratings can be raised by leniency and expressiveness rather than substance, circularity is a real risk.

A second objection: consultation is expensive. An effect size of 0.69 obtained by pairing every instructor with a skilled educational developer is not a policy most institutions can fund at scale, and the studies are typically conducted with volunteer participants - who are precisely the instructors least likely to need the intervention. Selection into consultation may inflate the estimate.

Third, publication bias affects this literature as it does most others, and several of the primary studies are decades old and small.

These caveats narrow the claim without overturning it. The honest version is: feedback that is interpreted, discussed and acted upon changes teaching more than feedback that is merely delivered - and the effect on independently observed teaching behaviour is less well established than the effect on subsequent ratings. The consultation studies that use trained-observer or peer-observation outcomes remain the most valuable and the least numerous.

Even under that narrower reading, the practical implication is unchanged, because the alternative on offer is not "rigorous consultation" versus "cheap feedback". It is consultation versus a PDF nobody opens.

What institutions should change

  • Move budget from collection to interpretation. If the choice is between a more elaborate questionnaire and a half-post educational developer, the evidence favours the developer.
  • Collect mid-cycle. The improvement effect in the literature comes from feedback delivered while the course is still running.
  • Make the data diagnostic, not just descriptive. An instructor cannot act on a mean. They can act on "six students said the week-four jump in notation lost them".
  • Separate improvement from accountability. Feedback routed into a promotion file will be received defensively, which is exactly the condition under which it does not work. This is the dual-purpose problem.
  • Track whether anything changed. Most institutions cannot say what any evaluation cycle caused. That is the action gap.

Where Koji fits

Consultation works because a skilled person asks follow-up questions and turns a score into a specific, attributable cause. That is expensive precisely because it is labour. Koji for Education automates the part of it that is structured questioning - not the professional judgement, but the probing that produces something worth judging.

Koji runs AI-moderated conversational interviews that ask why behind every rating, so what reaches the instructor is already diagnostic rather than descriptive. Automatic thematic analysis of open-text responses turns hundreds of comments into named, quantified themes, which is the preparatory work an educational developer would otherwise do by hand before the conversation can even start - and the step that most often does not happen at all. Quality scoring identifies which responses carry real information, so limited consultation time is spent on signal.

Because Koji supports formative, mid-cycle collection, it delivers feedback at the point in the term where the literature says it changes outcomes, rather than after the cohort has dispersed. Closing-the-loop action tracking records what an instructor or programme committed to and whether it happened, which is the only way an institution can ever evaluate its own evaluation system. Programme- and institution-level reporting shows educational developers where scarce consultation capacity will do the most good. All of it is GDPR-compliant with EU-appropriate data handling, and moderation is standardised and bias-aware, so probing does not vary with who happens to be running it.

To be precise about the claim: Koji does not replace consultation, and the evidence above does not support anyone claiming it could. Skilled human interpretation remains the active ingredient. What Koji does is remove the preparation bottleneck that keeps consultation from scaling, and ensure the feedback arriving on the desk is specific enough to act on.

Colleagues doing general user or customer research use the same AI interview engine on the main Koji platform.


The questionnaire was never the intervention. See how Koji for Education makes student feedback diagnostic enough to act on.