New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods12 min read

Mid-Semester Feedback and the Power of Consultation: What the Meta-Analyses Show

Penny & Coe (2004) and Cohen (1980) on the formative effectiveness of mid-semester student feedback when combined with consultation, and the implications for European course-evaluation design.

Koji Education Team

Product

Answer in brief

Meta-analytic evidence — Penny and Coe (2004) and Cohen (1980) — shows that mid-semester student feedback on its own produces only a small improvement in end-of-term ratings, but mid-semester feedback combined with consultation (a brief conversation with a teaching consultant or peer about what the feedback means and how to act on it) produces a moderate-to-large effect (Cohen's d ≈ 0.69 in Penny & Coe's 11-study synthesis). For European quality-assurance teams, this is the strongest empirical evidence that course evaluation can improve teaching: not as a summative ranking but as a formative cycle in which evidence reaches the instructor with enough specificity, and enough support, to drive change.

What the research says

The anchor paper for this brief is Penny, A. R., & Coe, R. (2004). "Effectiveness of consultation on student ratings feedback: A meta-analysis." Review of Educational Research, 74(2), 215–253. Penny and Coe synthesised 11 well-controlled intervention studies in which instructors received mid-semester student ratings with or without an accompanying consultation, and the outcome was measured at end of term on the same instrument. The mean effect size for consultation-augmented feedback was Cohen's d ≈ 0.69 — a moderate-to-large effect by Cohen's benchmarks. Penny and Coe also conducted a moderator analysis: effects were larger when consultation was substantive (multiple meetings, structured discussion of evidence) and weaker when the consultation was perfunctory.

The paper builds on the earlier Cohen, P. A. (1980), "Effectiveness of student-rating feedback for improving college instruction: A meta-analysis of findings," Research in Higher Education (13(4), 321–341), which integrated 22 comparisons and concluded that mid-semester rating feedback alone produced a modest improvement — instructors receiving feedback ended the term about 0.16 of a rating point higher than controls, equivalent to roughly a third of a standard deviation or a 15-percentile gain. Cohen also noted that "augmentation" of the rating with consultation accentuated the effect, foreshadowing what Penny and Coe quantified more precisely two decades later.

More recent work corroborates and extends these findings. Marsh and Roche (1993) and Marsh (2007) report that targeted feedback paired with a structured "Students' Evaluations of Educational Quality" (SEEQ) consultation produced sustained improvements in instructor ratings across subsequent terms. A more recent hierarchical meta-analysis on student-to-teacher feedback in primary and secondary education by Schildkamp et al. (2024), published in Educational Assessment, Evaluation and Accountability (https://doi.org/10.1007/s11092-024-09450-9), finds positive effects on classroom-quality dimensions in school contexts, with consultation again moderating the effect. The mechanism repeats: feedback that reaches the teacher with enough specificity, and with enough scaffolded reflection on what the data means and what to do next, changes teaching behaviour. Feedback that arrives as a year-end rating, with no scaffold and no time to act, generally does not.

The finding is one of the clearest in the SET literature: the formative use case has measurable effects; the summative-ranking use case has weak evidence (see Do Student Evaluations Measure Learning? The Uttl Meta-Analysis Revisited). This is sometimes framed as an asymmetry — SET data is useful when it is fed back into a learning loop, much less useful when it is used purely to rank.

Why it matters for course evaluation in practice

For a Dutch, Belgian, German, or Nordic QA office, three implications flow from this evidence base.

First, timing matters. End-of-semester evaluations cannot influence the cohort that produced them; by definition, that group has finished. The evidence on instructional improvement is largely drawn from mid-cycle feedback, where the instructor still has time to act and the same students can experience the result. If your evaluation programme is exclusively end-of-term, you are operating in the part of the feedback design space where Cohen and Penny–Coe found weaker effects.

Second, consultation is not optional. The Penny–Coe effect size requires a consultation layer. This can be a discipline-based educational developer, a head of programme, a peer coach, or a structured AI conversation that helps the instructor read the evidence and choose a plan. Without that layer, you are at the lower-bound Cohen 1980 effect size (small but present) at best.

Third, the consultation is about interpretation, not about marking. Penny and Coe found that the effective consultations focused on what the feedback meant in context and what specific behaviours could change. Consultations framed as performance reviews — "your scores are below threshold; please improve" — do not produce the same effect and risk creating defensive responses that undermine the cycle. Under ESG 2015 standard 1.5 (teaching staff), the expectation is that institutions support the development of teaching staff; a formative feedback-plus-consultation loop is a direct implementation of that standard.

There is a parallel implication for programme-level QA. The same logic — specific evidence, scaffolded interpretation, time to act — applies to programme-committee review. Programme committees that receive a year-end PDF of mean scores cannot run the kind of structured interpretation cycle that produces change; programme committees that receive mid-cycle qualitative evidence and meet to discuss it can.

Limitations and honest caveats

The Penny and Coe (2004) meta-analysis draws on 11 studies — a modest evidence base, with most studies from North American higher education in the 1980s–1990s. Generalisation to contemporary European HE is plausible (the underlying mechanism — specific feedback + reflection → behaviour change — is well supported across many learning contexts) but not guaranteed. The outcome in these studies is end-of-term student ratings, which by the SET–learning evidence (Uttl et al. 2017) are an imperfect proxy for actual teaching quality. So strictly, what Penny–Coe demonstrates is that consultation-augmented feedback improves student-perceived course experience — not necessarily that it improves learning. That is still a worthwhile outcome and consistent with the ESG focus on student-centred learning, but it is worth labelling accurately.

The original studies also varied substantially in how "consultation" was operationalised, from single 30-minute meetings to multi-session structured discussions. The d ≈ 0.69 figure is an average across this heterogeneity; the practical effect at any particular institution will depend on the depth and skill of the consultation function the institution can sustain. Building consultant capacity is a real resource commitment.

Finally, formative mid-cycle feedback can interact badly with summative end-of-term feedback if students perceive the mid-cycle as a performance test. The literature recommends keeping formative cycles clearly separate from summative use and not aggregating mid-cycle data into year-end performance scores.

How Koji incorporates this

Koji for Education (edu.koji.so) is built around the formative cycle that the Penny–Coe evidence supports. Four mechanisms operationalise it.

Mid-cycle conversational check-ins as a first-class workflow. Koji is designed for short, focused AI-moderated check-ins through the term — typically 5–10 minutes for the student — that produce structured qualitative evidence specific enough to act on. The product flow assumes the instructor will see results in time to change something for the same cohort.

Conversational depth surfaces the behaviours to consult about. Penny and Coe's mechanism depends on having specific behavioural evidence to discuss. A Likert mean ("3.8 on clarity") is not consultation-grade input; "students described examples in lectures as too abstract and wanted more worked problems before the exercise sheet" is. Koji's thematic analysis surfaces the behavioural strand of the open responses precisely so the consultation has something specific to work with.

Structured consultation prompts. Koji generates an interpretation summary that scaffolds the conversation: what students raised, where the evidence is strong vs thin, what specific behaviours could change. This is designed to function as the "augmentation" layer Cohen and Penny–Coe identified as the active ingredient — whether the actual consultation is delivered by a human educational developer, a peer, or in some institutional contexts as a structured self-review.

Separation of formative and summative streams. Koji's data model distinguishes mid-cycle formative collection from end-of-term summative collection. Mid-cycle results are owned by the instructor and the QA partner they choose to share with; they are not pooled into a year-end ranking. This separation is important both for the evidence — students respond differently when they know the data is formative — and for the institutional politics that determine whether the formative loop actually runs.

For teams that also run user, customer, or product research outside the course-evaluation context, Koji's core research platform at koji.so applies the same conversational-interview engine and thematic-analysis pipeline.

The central design choice: course evaluation is at its empirically strongest when it is a cycle — specific evidence, time to interpret, support to act, repeat — and Koji is built to make that cycle the default, not the exception.

Related resources

References

  • Penny, A. R., & Coe, R. (2004). Effectiveness of consultation on student ratings feedback: A meta-analysis. Review of Educational Research, 74(2), 215–253. https://doi.org/10.3102/00346543074002215
  • Cohen, P. A. (1980). Effectiveness of student-rating feedback for improving college instruction: A meta-analysis of findings. Research in Higher Education, 13(4), 321–341. https://doi.org/10.1007/BF00976252
  • Marsh, H. W., & Roche, L. A. (1997). Making students' evaluations of teaching effectiveness effective: The critical issues of validity, bias, and utility. American Psychologist, 52(11), 1187–1197. https://doi.org/10.1037/0003-066X.52.11.1187
  • Uttl, B., White, C. A., & Gonzalez, D. W. (2017). Meta-analysis of faculty's teaching effectiveness: Student evaluation of teaching ratings and student learning are not related. Studies in Educational Evaluation, 54, 22–42. https://doi.org/10.1016/j.stueduc.2016.08.007
  • Wisniewski, B., Zierer, K., & Hattie, J. (2020). The power of feedback revisited: A meta-analysis of educational feedback research. Frontiers in Psychology, 10, 3087. https://doi.org/10.3389/fpsyg.2019.03087