Why Was the Course Rated Well, Not Just Whether? Mediation Analysis and the Mechanism Behind an Evaluation Score
A high rating tells you students were satisfied; it doesn't tell you why. Mediation analysis tests the pathway — did clarity raise satisfaction by increasing engagement? — but drawing that causal chain from a single end-of-term survey is far harder than the classic recipe suggests.
Koji Education Team
Product
The short answer
Mediation analysis asks not whether something affected an evaluation outcome but how — through what intervening variable the effect travelled. Did clearer teaching raise overall satisfaction because it increased student engagement? Engagement is the hypothesised mediator, and the "indirect effect" through it is what mediation estimates. This is a genuinely useful question for course evaluation, because knowing the mechanism tells you what to change. But the classic four-step recipe most people learned is underpowered and partly obsolete, and — more importantly — a mediation claim drawn from a single cross-sectional survey is causally fragile no matter how the statistics are done.
BLUF for busy readers: Mediation decomposes an effect into a direct path and an indirect path through a mediator (X → M → Y). Modern practice estimates the indirect effect directly and bootstraps its confidence interval, rather than using the low-power Baron-Kenny causal steps. The statistics are the easy part; the hard part is that mediators are almost never randomised, so "clarity works through engagement" is a causal story your correlational survey data can support only weakly.
What the research says
The foundational reference is Baron and Kenny (1986) in the Journal of Personality and Social Psychology — one of the most-cited papers in all of social science. They distinguished a mediator (a variable through which an effect is transmitted; it explains how or why) from a moderator (a variable that changes the strength or direction of an effect; it explains for whom or when), and proposed a four-step "causal steps" procedure for establishing mediation: show X predicts Y, X predicts M, M predicts Y controlling for X, and that the X–Y effect shrinks when M is added.
That recipe dominated for two decades, but methodologists have since shown it is the weakest of the available approaches. MacKinnon, Lockwood, Hoffman, West, and Sheets (2002) in Psychological Methods compared 14 methods for testing intervening-variable effects and found that the causal-steps strategy — especially the requirement that the direct effect become non-significant — had the lowest statistical power, reaching acceptable power only with samples above roughly 500. Hayes (2009) in Communication Monographs drove the point home in "Beyond Baron and Kenny," arguing that researchers should estimate and test the indirect effect itself (the product of the X→M and M→Y paths) using bootstrapped confidence intervals, which make no assumption that the indirect effect is normally distributed — the basis of the widely used PROCESS approach. A crucial corollary: you do not need a significant total X→Y effect to have real mediation, overturning Baron and Kenny's first step.
But the deepest caution is causal, not statistical. Bullock, Green, and Ha (2010) in JPSP — "Yes, but what's the mechanism? (Don't expect an easy answer)" — show that mediation analyses on non-experimental data are "likely to be biased," and that even randomising X does not rescue you, because the mediator is still not randomised. If some unmeasured variable causes both the mediator and the outcome, the estimated indirect effect is confounded. Their verdict: studying mechanism credibly is far harder than the popular recipe implies, and usually requires designs that manipulate the mediator, not just measure it.
Why it matters for course evaluation in practice
Course evaluation is full of implicit mediation claims, usually stated far too confidently:
- Diagnosing why a teaching change worked. Suppose a redesigned course scores higher. Was it because the redesign improved perceived clarity, which raised satisfaction? Mediation formalises that chain and estimates how much of the total improvement ran through clarity versus other paths — turning "students liked it" into "students liked it because X," which is what tells you what to keep.
- Separating warmth from substance. The literature on the warmth halo and on teacher clarity predicting learning are both, at heart, mediation questions: does immediacy raise ratings through genuine perceived learning, or directly through affect independent of learning? A mediation model makes those competing pathways explicit and estimable.
- Prioritising interventions. If engagement mediates most of the clarity→satisfaction effect, then interventions that raise engagement are leverage points. Mediation is, done well, a map of where to push.
- Avoiding the moderator/mediator mix-up. Institutions routinely confuse "the effect was bigger for first-years" (moderation) with "the effect worked through first-years' engagement" (mediation). Baron and Kenny's core contribution was insisting these are different questions with different analyses; conflating them produces incoherent action plans.
Limitations and honest caveats
This is a method where the statistics can look rigorous while the inference underneath is weak, so the caveats are the point:
- Cross-sectional mediation is causally weak. Measuring X, M, and Y on the same end-of-term survey cannot establish that X preceded M preceded Y. The temporal ordering is assumed, not observed, and reversing M and Y often fits the data just as well. A satisfied student may report more engagement, not the reverse.
- The mediator is almost never randomised. As Bullock and colleagues stress, even a randomised teaching intervention leaves the mediator (engagement, perceived clarity) self-selected. Any unmeasured common cause of mediator and outcome biases the indirect effect. This is not a nuisance to footnote; it is the central threat.
- Common-method bias inflates indirect effects. When X, M, and Y all come from one student rating instrument, shared method variance can create the very correlations mediation feeds on. Single-source mediation should be read with heavy scepticism.
- The old recipe is underpowered. Following Baron and Kenny's requirement that the direct effect vanish will miss real mediation in typical class-sized samples. Use bootstrapped indirect-effect tests instead — but remember better statistics do not fix the causal problems above.
- Mediation is not decomposition of blame. A statistically significant indirect effect is a consistency check on a hypothesised mechanism, not proof of it. Competing models with different mediators frequently fit equally well.
The honest framing: mediation analysis is valuable for generating and sharpening mechanism hypotheses, and modern indirect-effect testing is the right way to do the statistics. But a mediation result from a one-shot course-evaluation survey is a suggestive account of mechanism, not a demonstrated causal chain — and it should be reported that way.
How Koji incorporates this
Koji is built around the conviction that the number is only the beginning and the mechanism is the prize — which is exactly the terrain mediation tries, imperfectly, to map:
- Probing the mechanism the number hides. Where a mediation model can only infer an intervening variable from covariance, Koji's AI-moderated conversational interviews ask students directly why they rated a course as they did, following an adaptive
open_endedchain that surfaces the pathway — "the examples made it click, so I actually kept up with the reading" — that a purely quantitative mediation would have to guess at. Automatic thematic analysis then aggregates these mechanisms across a cohort. This is qualitative triangulation for a mediation hypothesis, not a replacement for the statistics. - Structured variables for a proper model. Koji's
scaleandsingle_choicequestions capture candidate mediators (engagement, perceived clarity) and outcomes (satisfaction, perceived learning) as clean, separable variables, so an analyst can specify and bootstrap an indirect-effect model rather than reverse-engineering one from a PDF export. - Reducing single-source bias. Because Koji can collect from multiple cohorts and pair self-report with LMS trace data, it offers a route out of the common-method trap that inflates cross-sectional indirect effects — measuring the mediator and the outcome from different sources where possible.
- Supporting temporal ordering. Koji's mid-cycle and end-of-cycle collection means a hypothesised mediator can be measured earlier than the outcome, giving a mediation claim the temporal precedence a single survey cannot — a partial, honest improvement on cross-sectional mediation, not a cure for the unmeasured-confounding problem Bullock and colleagues identify.
Koji does not compute a mediation model behind a dashboard tile and present the indirect effect as established causation — doing so would commit exactly the overreach the methodological literature warns against. Instead it supplies clean, separable, appropriately-timed variables and direct qualitative evidence of mechanism, so a competent analyst can build a defensible model and report its causal limits honestly. Koji's core research platform at koji.so applies the same mechanism-probing interview engine to product and customer research, where "why did they churn" is a mediation question every bit as fragile as "why did they rate the course highly."
References
- Baron, R. M., & Kenny, D. A. (1986). The moderator–mediator variable distinction in social psychological research: Conceptual, strategic, and statistical considerations. Journal of Personality and Social Psychology, 51(6), 1173–1182. https://doi.org/10.1037/0022-3514.51.6.1173
- MacKinnon, D. P., Lockwood, C. M., Hoffman, J. M., West, S. G., & Sheets, V. (2002). A comparison of methods to test mediation and other intervening variable effects. Psychological Methods, 7(1), 83–104. https://doi.org/10.1037/1082-989X.7.1.83
- Hayes, A. F. (2009). Beyond Baron and Kenny: Statistical mediation analysis in the new millennium. Communication Monographs, 76(4), 408–420. https://doi.org/10.1080/03637750903310360
- Bullock, J. G., Green, D. P., & Ha, S. E. (2010). Yes, but what's the mechanism? (Don't expect an easy answer). Journal of Personality and Social Psychology, 98(4), 550–558. https://doi.org/10.1037/a0018933
Frequently asked questions
What is the difference between a mediator and a moderator? A mediator is a variable through which an effect travels — it explains how or why X affects Y (X → M → Y). A moderator changes the strength or direction of the X–Y effect — it explains for whom or under what conditions. Confusing the two, as Baron and Kenny warned, produces incoherent conclusions and action plans.
Why shouldn't I use the classic Baron-and-Kenny causal steps anymore? Because MacKinnon and colleagues (2002) showed the causal-steps approach — especially requiring the direct effect to become non-significant — has the lowest power of the available methods, working only with very large samples. Modern practice estimates the indirect effect directly and bootstraps its confidence interval, and does not require a significant total effect to detect mediation.
Can I trust a mediation result from a single end-of-term survey? Only weakly. Measuring X, M, and Y at the same time cannot establish that X preceded M preceded Y, and the mediator is not randomised, so unmeasured common causes bias the indirect effect. Single-source measurement also invites common-method bias. Treat cross-sectional mediation as a hypothesis about mechanism, not a demonstrated causal chain.
Does randomising the intervention fix mediation's causal problems? No. Bullock, Green, and Ha (2010) show that even when X is randomised, the mediator is still self-selected, so an unmeasured cause of both mediator and outcome still biases the estimate. Credible mechanism claims usually require manipulating the mediator itself, not just measuring it.
How does Koji help with mechanism questions? Koji probes the mechanism directly through AI-moderated interviews that ask students why they rated a course as they did, captures candidate mediators and outcomes as clean separable variables, supports measuring a mediator earlier than the outcome, and pairs self-report with LMS data to reduce single-source bias. It supplies the ingredients for a defensible model rather than asserting causation behind a dashboard.
Related resources
- Teacher Clarity Predicts Learning Better Than Charisma: What Course Evaluations Should Measure
- The Warmth Halo: How Instructor Immediacy Shapes Course Evaluations More Than Learning Does
- Why a Single Student Survey Can't Stand Alone: Common-Method Bias in Course Evaluation
- Contribution Analysis: Making Credible Causal Claims from Course Evaluation
- Do the Clicks Confirm the Comments? Triangulating Course Evaluations with LMS Learning-Analytics Data
- Did the Teaching Change Cause the Score to Move? Difference-in-Differences for Course Evaluation
Related articles
Why a Single Student Survey Can't Stand Alone: Common-Method Bias in Course Evaluation
Common-method bias (Podsakoff et al., 2003) explains why correlations inside a single end-of-term student survey are inflated by the shared method itself - and why triangulating teaching evidence matters. A research-grounded guide for quality assurance.
Did the Teaching Change Cause the Score to Move? Difference-in-Differences for Course Evaluation
Difference-in-differences lets you estimate whether a course redesign actually moved evaluation scores by comparing a treated course against a similar untouched one over time. Here is how the design works, when its parallel-trends assumption holds, and how to use it honestly in quality assurance.
The Warmth Halo: How Instructor Immediacy Shapes Course Evaluations More Than Learning Does
What the meta-analytic evidence on teacher nonverbal immediacy and warmth tells us about course evaluations — why warmth strongly predicts how much students like a course and think they learned, but only weakly predicts what they actually learn.
Teacher Clarity Predicts Learning Better Than Charisma: What Course Evaluations Should Measure
A meta-analysis of 144 effects and 73,000+ students shows teacher clarity explains roughly 13% of the variance in student learning. Here is what that means for the items you put on a course evaluation.