Why Is the Effect Bigger in Some Sections? Meta-Regression for Course Evaluation
Meta-regression explains between-section heterogeneity in pooled course-evaluation results by modelling section-level moderators — powerful for generating explanations, weak for confirming them.
Koji Education Team
Product
In brief
Pooling your own course-evaluation results across sections, terms, or cohorts almost always leaves heterogeneity — the effect you care about is bigger in some strata than others. Meta-regression is the tool that asks why: it regresses the section-level effect sizes on section-level moderators (class size, modality, discipline, instructor experience) to see what explains the spread. It is a genuine step beyond a random-effects meta-analysis, which only quantifies heterogeneity; but because it operates on aggregated, observational study-level data, it is prone to false positives and aggregation bias, and Higgins and Thompson (2004) showed you can reliably examine only about one moderator per ten sections.
What the research says
A random-effects meta-analysis of your own sections gives you three things: a pooled estimate, a between-section variance (τ²), and a heterogeneity statistic (I²) that says how much the true effect varies from section to section. What it does not tell you is what accounts for that variation. Meta-regression is the extension that answers that question.
Thompson and Higgins (2002), in the paper that set the modern standard, define meta-regression as a weighted regression in which the unit of analysis is the study (here, the section or cohort), the outcome is that study's effect size, and the predictors are study-level characteristics. Crucially, the weighting must reflect both the within-study sampling variance and the residual between-study heterogeneity — a random-effects meta-regression — otherwise small sections dominate and the residual variance is understated. They warn explicitly that meta-regression is an observational analysis at the aggregate level: an association between a moderator and the effect size across studies does not license an individual-level causal claim, an error they call aggregation bias (the study-level cousin of the ecological fallacy).
Higgins and Thompson (2004) then quantified how dangerous the method is when used casually. Using simulation, they showed that the type I error rate of meta-regression — the chance of "discovering" a moderator that has no real effect — is badly inflated when heterogeneity is present and when several covariates are screened. Their headline practical rule of thumb is that roughly ten studies are needed for each covariate examined, and they proposed a permutation test that recomputes the moderator's significance against many random re-labellings of the studies, which restores honest error rates. The heterogeneity statistics themselves — I² and τ² — come from Higgins and Thompson (2002), and the wider framework of fixed-effect versus random-effects synthesis, effect-size choice, and forest and bubble plots is laid out in Borenstein, Hedges, Higgins and Rothstein (2009).
The practical implementation for applied users is well documented. Harbord and Higgins (2008) describe the standard estimators (method-of-moments, REML) and the residual-heterogeneity and permutation options in widely used software, and emphasise the "bubble plot" — effect size against moderator, with point size proportional to precision — as the honest visual summary. Across this literature the consistent message is the same: meta-regression is powerful for generating explanations of heterogeneity and weak for confirming them.
Why it matters for course evaluation in practice
A quality-assurance office that runs the same instrument across a faculty is sitting on exactly the data meta-regression was built for. Suppose a new active-learning redesign was rolled out across eleven parallel sections and the pooled effect on the "the course helped me learn" item is a modest but heterogeneous improvement. The interesting managerial question is never the average; it is whether the gain is concentrated in small seminars and absent in large lectures, or whether it depends on whether the section was compulsory or elective. Meta-regression turns those hunches into a modelled, weighted answer, and its bubble plot communicates the finding to a committee far better than a table of eleven means.
It also disciplines the most common reporting mistake in a dean's office: reading a subgroup difference off a cross-tab and treating it as real. Meta-regression forces the between-section noise into the standard error, so a "big-classes-did-worse" pattern that is really just three noisy small sections gets the wide confidence interval it deserves. In that sense it is the multi-section sibling of empirical-Bayes shrinkage: both refuse to over-trust a single stratum. And where a multilevel model fits the raw student responses directly, meta-regression is the pragmatic alternative when all you can retrieve from historical terms is the summary per section — a mean, an n, and an SD — which is frequently all a legacy evaluation system exports.
Limitations and honest caveats
Three cautions dominate, and a credible report states all three.
First, power and false positives. With eleven sections you can honestly test one moderator, not five. Screening class size, modality, discipline, timetable slot and instructor rank on a dozen sections and reporting the one that reached p < 0.05 is precisely the behaviour Higgins and Thompson (2004) modelled as a near-guaranteed false discovery. Pre-specify the single moderator that theory nominates, or use their permutation test and report it as exploratory. This is the same multiple-comparisons trap in a synthesis wrapper.
Second, aggregation bias. A relationship between average class size and average rating across sections need not hold within sections or across students. Meta-regression cannot see individual students; it can only see section summaries. Any causal-sounding sentence ("larger classes cause lower ratings") over-reaches the design, which is observational and confounded — bigger classes differ from smaller ones in many unmodelled ways.
Third, it does not fix confounding. Moderators are not randomly assigned to sections. If experienced instructors teach the small electives, an "experience effect" and a "class-size effect" are hopelessly entangled, and meta-regression will happily attribute the whole gap to whichever one you happened to enter. The honest framing is hypothesis-generating: meta-regression tells you where to look, not what caused it, and its output should feed a specification-curve or multiverse check rather than a personnel decision.
How Koji incorporates this
Koji is built to produce the clean, section-level inputs meta-regression needs and to keep users from over-reading its outputs. Every study in Koji carries structured metadata — cohort, modality, class size, discipline, term — so the moderators for a meta-regression are captured as fields rather than reconstructed by hand later. When results are pooled across sections, Koji's reporting is designed to surface heterogeneity first (how much the effect varies) before offering any explanation, so a user is nudged toward "is there spread to explain?" before "what explains it?".
Where a moderator analysis is run, Koji is designed to weight by section precision and to flag the studies-per-covariate ratio, warning when a user is trying to explain the variation among eight sections with four moderators — the exact over-fitting Higgins and Thompson caution against. Because Koji's AI-moderated conversational interviews also capture why a cohort responded as it did, a quantitatively-flagged moderator (say, large lectures rating the redesign lower) can be triangulated against the open-text themes from those same large sections, rather than left as an unexplained coefficient. The aim throughout is framed as designed to support honest heterogeneity analysis, not to certify causal moderators — the coefficient is a lead, and the interview transcript is where the lead gets tested. Teams that also run product or customer research on Koji's core platform at koji.so use the same section-level synthesis logic to compare segments across studies, so the discipline transfers directly.
Frequently asked questions
How is meta-regression different from a random-effects meta-analysis?
A random-effects meta-analysis pools your sections into one estimate and reports how much the true effect varies between them (τ² and I²). Meta-regression goes one step further and models that variation as a function of section-level characteristics, so it answers why the effect is larger in some sections than others. You almost always run the meta-analysis first and reach for meta-regression only when it reveals substantial heterogeneity worth explaining.
How many sections do I need before I can trust a moderator?
Higgins and Thompson's (2004) rule of thumb is roughly ten studies per covariate. With eleven sections you can credibly examine one pre-specified moderator; with twenty you might examine two. Screening many moderators on a handful of sections produces false positives at a very high rate, so treat any moderator found that way as exploratory and, ideally, confirm it with their permutation test.
What is aggregation bias and why does it matter here?
Meta-regression works on section averages, not individual students. A relationship that holds across section averages (bigger average class size, lower average rating) may not hold within sections or across students — a study-level version of the ecological fallacy. It means you can never read an individual-level causal claim off a meta-regression coefficient; the association is between whole sections.
Can meta-regression prove that class size lowers ratings?
No. Sections are not randomly assigned their class size, modality or instructor, so any moderator is confounded with everything else that differs between sections. Meta-regression is observational and hypothesis-generating: it tells you where an association sits, not what caused it. Treat a strong coefficient as a lead to investigate, not as evidence for a policy or personnel decision.
What is a bubble plot?
A bubble plot graphs each section's effect size against the moderator value, sizing each point by its precision (larger points for sections with more responses), and overlays the fitted meta-regression line. It is the honest visual summary because it shows at a glance whether the fitted slope is driven by a couple of small, noisy sections or by a consistent trend across well-measured ones.
Do I need student-level data to run it?
No — that is a large part of its appeal. Meta-regression needs only each section's summary: an effect size (or mean and SD) and a sample size, plus the moderator values. When a legacy evaluation system exports only per-section summaries and not the raw responses, meta-regression is often the only defensible synthesis available, whereas a multilevel model would require the individual rows.
References
- Thompson, S. G., & Higgins, J. P. T. (2002). How should meta-regression analyses be undertaken and interpreted? Statistics in Medicine, 21(11), 1559–1573. https://doi.org/10.1002/sim.1187
- Higgins, J. P. T., & Thompson, S. G. (2002). Quantifying heterogeneity in a meta-analysis. Statistics in Medicine, 21(11), 1539–1558. https://doi.org/10.1002/sim.1186
- Higgins, J. P. T., & Thompson, S. G. (2004). Controlling the risk of spurious findings from meta-regression. Statistics in Medicine, 23(11), 1663–1682. https://doi.org/10.1002/sim.1752
- Harbord, R. M., & Higgins, J. P. T. (2008). Meta-regression in Stata. The Stata Journal, 8(4), 493–519. https://doi.org/10.1177/1536867X0800800403
- Borenstein, M., Hedges, L. V., Higgins, J. P. T., & Rothstein, H. R. (2009). Introduction to Meta-Analysis. Chichester: Wiley. https://doi.org/10.1002/9780470743386
Related resources
- Pooling Your Own Sections Without Faking Precision: Random-Effects Meta-Analysis for Course Evaluation
- Stop Comparing Raw Averages: Empirical-Bayes Shrinkage for Course-Evaluation Scores
- The Department Average Is Not the Student: The Ecological Fallacy in Course Evaluation
- Students Are Nested in Courses: Why Multilevel Models Beat Raw Averages
- Does the Effect Depend on Who or What? Moderation and Interaction Effects
- One Instructor, Two Hundred Defensible Scores: Specification-Curve Analysis
Related articles
Pooling Your Own Sections Without Faking Precision: Random-Effects Meta-Analysis for Course Evaluation
When you combine ratings across many small sections or terms, averaging the averages pretends they all measure one fixed truth. Random-effects meta-analysis treats each section as a noisy estimate of a genuinely varying effect, weights them properly, and reports how much real spread remains — with a prediction interval, not just a mean.
Stop Comparing Raw Averages: Empirical-Bayes Shrinkage for Course-Evaluation Scores
Why raw course-evaluation means mislead — especially for small classes — and how empirical-Bayes shrinkage (partial pooling), the method Kane and Staiger applied to teacher value-added, produces fairer, more accurate instructor estimates.
Students Are Nested in Courses: Why Multilevel Models Beat Raw Averages for Course Evaluation
Course-evaluation data has a nested structure — students within sections within instructors — and a flat average ignores it. This guide explains how multilevel (hierarchical linear) models partition variance, why ignoring clustering understates uncertainty, and what it means for fair reporting.
One Instructor, Two Hundred Defensible Scores: Specification-Curve Analysis for Course Evaluation
Every course-evaluation "score" is the product of dozens of defensible analytic choices — which items count, how to weight, mean or median, whether to adjust for class size. Specification-curve and multiverse analysis compute the result across all of them, so a personnel decision rests on the pattern, not on one analyst's arbitrary path.