New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
analysis-reporting11 min read

Does the Effect Depend on Who or What? Moderation and Interaction Effects in Course-Evaluation Analysis

A bias or a teaching effect on evaluations often holds only for some students, courses or conditions. Moderation analysis tests those "it depends" claims properly, and it is far harder than it looks.

Koji Education Team

Product

In brief: A moderator is a variable that changes the strength or direction of the relationship between two others — for example, if grading leniency inflates ratings more in large classes than small ones, class size moderates the leniency–rating link. Moderation is tested with an interaction term in a regression, not by eyeballing subgroup averages. Baron and Kenny (1986) drew the sharp line between moderation and mediation, and McClelland and Judd (1993) showed that interactions are genuinely hard to detect in observational data — field tests can have under 20% of the efficiency of an ideal experiment. So most "it works differently for group X" claims about course evaluation are under-powered and over-stated.

Deans and QA officers constantly make conditional claims: "the redesign helped, but only in the seminars," "gender bias shows up in quantitative disciplines," "difficulty lowers ratings for weaker students." These are all moderation hypotheses, and they are among the easiest claims to assert and the hardest to establish. This article explains how to test them and why the bar is higher than intuition suggests.

What the research says

The conceptual foundation is Baron and Kenny (1986), "The moderator–mediator variable distinction in social psychological research" (Journal of Personality and Social Psychology, 51(6), 1173–1182), one of the most cited methods papers in the social sciences. They define a moderator as a variable that affects the direction or strength of the relation between a predictor and an outcome — statistically, an interaction between the predictor and the moderator. A mediator, by contrast, is a variable that transmits the effect (predictor → mediator → outcome). Confusing the two is a common and consequential error: moderation answers "for whom or under what conditions does this hold?"; mediation, covered in our mediation analysis article, answers "why or through what?"

The mechanics of testing moderation are laid out in Aiken and West (1991), Multiple Regression: Testing and Interpreting Interactions. You regress the outcome on the predictor, the moderator, and their product term; a significant product term indicates moderation. Their practical guidance is essential: centre (or standardise) continuous predictors before forming the product to reduce nonessential multicollinearity and make main effects interpretable; probe a detected interaction with simple slopes (the predictor–outcome relationship computed at low, medium, and high values of the moderator); and always plot the interaction rather than trusting the coefficient alone.

The sobering result is McClelland and Judd (1993), "Statistical difficulties of detecting interactions and moderator effects" (Psychological Bulletin, 114(2), 376–390). They show that observational field studies are dramatically underpowered for interactions relative to designed experiments — often with less than 20% of the efficiency — because in naturally-occurring data the joint extremes of predictor and moderator (which carry most of the information about an interaction) are rare. The consequence: real moderators frequently go undetected, and the interactions that do reach significance in small samples are prone to the sign and magnitude exaggeration that afflicts any low-power estimate. Detecting moderation reliably usually needs large samples, wide variation, or deliberately extreme sampling.

Why it matters for course evaluation in practice

Conditional claims are the currency of evaluation debate, and getting the analysis wrong has real costs:

  1. "The intervention worked, but only for some." Whether a teaching redesign's effect on ratings depends on prior student interest or discipline is a predictor-by-moderator interaction. Reading two subgroup means and declaring a difference is not a moderation test — and, as the "difference between significant and not significant is not itself significant" principle warns, a significant effect in one subgroup and a non-significant one in another does not establish that the effect differs between them. Only the interaction term does.

  2. Bias that is conditional. Claims that gender or accent bias is stronger in some disciplines are moderation hypotheses. Testing them properly, with an interaction term and adequate power, separates a real conditional bias from noise — and protects against both complacency and over-claiming.

  3. Confound versus moderator. A variable such as class size can act as a confound (a common cause to be adjusted away) or as a moderator (changing how another factor relates to ratings). These are different roles requiring different models; treating a moderator as a mere covariate, or vice versa, misstates what is going on. Our note on propensity-score adjustment for confounds covers the confounding case; moderation is the distinct question of effect heterogeneity.

The practical discipline is: state the moderation hypothesis explicitly, fit a proper interaction model on ordinal-appropriate methods (see ordinal regression for evaluation data), centre the predictors, judge the interaction term with honest uncertainty, and probe and plot simple slopes before telling any story. And expect most honest tests on modest data to be inconclusive rather than confirmatory.

Limitations and honest caveats

  • Chronic low power. Following McClelland and Judd (1993), observational moderation tests are usually underpowered. A null interaction is weak evidence of no moderation, and a significant one on a small sample is likely exaggerated. Report power or design-analysis expectations alongside the estimate.
  • Product terms are unreliable. The reliability of an interaction term is roughly the product of the reliabilities of its components, so interactions built from noisy course-evaluation items are measured even less reliably than the main effects — compounding the detection problem.
  • Scaling and nonlinearity create phantom interactions. On a bounded, ceiling-prone 1–5 scale, an unmodelled nonlinearity or a monotone transformation can manufacture or erase an interaction. An apparent moderation may be an artefact of the response scale, not a substantive conditional effect.
  • Multiple interaction tests inflate false positives. Fishing across many candidate moderators without correction guarantees spurious "significant" interactions. Pre-specify the moderators you will test.
  • Correlational moderation is not causal. Finding that an effect is statistically stronger in one group does not explain why, and the moderator itself may be standing in for something else. Moderation describes heterogeneity; it does not by itself identify a mechanism.

Moderation analysis is indispensable for honest "it depends" claims, but it demands more data, cleaner measurement, and more restraint than a subgroup comparison.

How Koji incorporates this

Koji for Education is designed to help institutions form and interrogate conditional claims responsibly, while being candid that small samples limit what any analysis can conclude.

  • Structured context capture for principled subgroups. Koji records the course, cohort, and respondent metadata that candidate moderators (discipline, level, class-size band, modality) are built from, so a moderation question can be tested against pre-specified, meaningful groupings rather than post-hoc data-dredged ones.
  • Hypothesis generation from open text. Koji's AI-moderated conversational interviews and automatic thematic analysis surface for whom an experience differed — "students new to the topic found the pace hard" — generating moderation hypotheses grounded in what students actually said, which can then be tested properly rather than assumed.
  • Mechanism plus condition. Because Koji captures both the rating and the reasoning behind it, it helps distinguish a moderation ("the effect is stronger for X") from a mediation ("the effect runs through Y"), the very distinction Baron and Kenny drew — reducing the common conflation of the two.
  • Estimation-first, uncertainty-honest reporting. Consistent with our guidance on small-sample inference, Koji foregrounds effect sizes and uncertainty, discouraging the leap from a suggestive subgroup gap to a confident conditional claim.
  • Triangulation across cohorts. Pooling comparable cohorts over time is often the only way to accumulate the joint variation McClelland and Judd show is needed to detect a real interaction; Koji is built to aggregate across terms and cohorts to support that.

Koji does not claim to make under-powered moderation tests conclusive — no platform can repeal the statistics. What it is designed to do is help you ask conditional questions with the right groupings, generate them from real qualitative signal, and read the answers with appropriate humility. The same analysis-first engine serves product and customer research on Koji's core platform at koji.so, where "does this hold for every segment?" is an equally common — and equally over-claimed — question.

Frequently asked questions

What is the difference between a moderator and a mediator? A moderator changes the strength or direction of the relationship between a predictor and an outcome (an interaction). A mediator transmits the effect from predictor to outcome (a mechanism). Baron and Kenny (1986) formalised the distinction; moderation asks "for whom or when?", mediation asks "why or through what?".

Can I test moderation by comparing subgroup averages? No. Comparing group means, or noting that an effect is significant in one group and not another, does not test moderation. You need an interaction term in a regression, and a significant difference in significance is not itself evidence that the effect differs between groups.

Why is moderation so hard to detect in course-evaluation data? Because observational data rarely contains the joint extremes of predictor and moderator that carry information about an interaction. McClelland and Judd (1993) show field tests can have under 20% of the efficiency of an ideal experiment, so real moderators often go undetected without large samples or wide variation.

Should I centre my variables before testing an interaction? Yes, for continuous predictors. Centring or standardising before forming the product term reduces nonessential multicollinearity and makes the main-effect coefficients interpretable, as Aiken and West (1991) recommend. Then probe the interaction with simple slopes and plot it.

Could an interaction I found just be an artefact of the rating scale? Yes. On a bounded, ceiling-prone scale, unmodelled nonlinearity or a transformation can create or remove an apparent interaction. Check robustness to scaling and use ordinal-appropriate models before interpreting a moderation substantively.

How many moderators should I test? As few as you can pre-specify from theory. Testing many candidate moderators without correction inflates false positives, so decide in advance which conditional effects you will examine rather than dredging the data.

Related resources

References

  • Baron, R. M., & Kenny, D. A. (1986). The moderator–mediator variable distinction in social psychological research: Conceptual, strategic, and statistical considerations. Journal of Personality and Social Psychology, 51(6), 1173–1182. https://doi.org/10.1037/0022-3514.51.6.1173
  • Aiken, L. S., & West, S. G. (1991). Multiple Regression: Testing and Interpreting Interactions. Newbury Park, CA: Sage.
  • McClelland, G. H., & Judd, C. M. (1993). Statistical difficulties of detecting interactions and moderator effects. Psychological Bulletin, 114(2), 376–390. https://doi.org/10.1037/0033-2909.114.2.376
  • Gelman, A., & Carlin, J. (2014). Beyond power calculations: Assessing Type S (sign) and Type M (magnitude) errors. Perspectives on Psychological Science, 9(6), 641–651. https://doi.org/10.1177/1745691614551642