New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
analysis-reporting10 min read

Pooling Your Own Sections Without Faking Precision: Random-Effects Meta-Analysis for Course Evaluation

When you combine ratings across many small sections or terms, averaging the averages pretends they all measure one fixed truth. Random-effects meta-analysis treats each section as a noisy estimate of a genuinely varying effect, weights them properly, and reports how much real spread remains — with a prediction interval, not just a mean.

Koji Education Team

Product

In brief

Every programme eventually faces the same question: taking all our sections (or all the terms this course has run), what is the overall picture — and how sure are we? The tempting move is to average the section means, or pool all students into one big mean. Both quietly assume every section is estimating the same fixed number, so any differences are pure noise. Random-effects meta-analysis — the workhorse of evidence synthesis in medicine — makes the more honest assumption: each section estimates a truth that itself varies from section to section, and the data contain two kinds of variation, within-section sampling error and between-section real heterogeneity. It weights each section by its precision, estimates how much genuine spread exists (τ² and I²), and — critically — reports a prediction interval telling you the range a new section is likely to fall in. That last number is what a fixed-effect average or a raw grand mean can never give you, and it is exactly what a dean deciding whether a result generalises needs.

What the research says

The random-effects model was formalised for applied synthesis by DerSimonian and Laird (1986, Controlled Clinical Trials), whose method for "combining evidence from a series of experiments" — incorporating between-study heterogeneity into the overall estimate — became the standard approach, with tens of thousands of citations. The model gives each unit a weight of 1/(within-unit variance + between-unit variance τ²): precise, large sections count more, but never so much that a single big section swamps the rest, because the added τ² term stops precise units from dominating the way they do in a fixed-effect analysis.

Quantifying heterogeneity is the second pillar. Higgins and Thompson (2002, Statistics in Medicine) introduced , "the proportion of total variation in study estimates that is due to heterogeneity rather than chance," alongside the H statistic. I² near 0% means the sections really are interchangeable estimates of one number; I² of 60–80% means most of the spread is real — the course genuinely behaves differently across sections, and reporting a single pooled mean would hide the most important finding.

The third pillar is the honest expression of uncertainty. IntHout, Ioannidis, Rovers and Goeman (2016, BMJ Open) made the influential "plea for routinely presenting prediction intervals in meta-analysis." Their argument: the confidence interval around the mean effect shrinks as you add sections, but it describes only the average — it says nothing about the range of true effects across sections. The prediction interval, which incorporates τ², answers the question people actually care about ("what should we expect a new section to do?") and is often startlingly wider than the confidence interval, puncturing false precision. Borenstein, Hedges, Higgins and Rothstein's (2009) Introduction to Meta-Analysis is the standard textbook tying these pieces together.

Why it matters for course evaluation in practice

Course-evaluation data are naturally meta-analytic: many small sections, repeated offerings, parallel cohorts — each a small, noisy study of "how good is this course?" Treating them correctly changes the conclusions.

  1. It stops big sections from silently dominating. A raw pooled mean across all students weights a 300-person lecture 15× a 20-person seminar. Random-effects weighting tempers that, so a programme-level figure is not just the big service course wearing a trenchcoat.
  2. It separates real spread from noise. I² tells you whether "sections vary from 3.9 to 4.6" reflects genuine differences worth investigating or is what you would expect from small-sample luck. That single diagnostic reframes many "trend" conversations.
  3. It reports what generalises. The prediction interval is the deliverable for accreditation and planning: not "our mean is 4.3 (95% CI 4.2–4.4)" but "a new section is likely to land between 3.8 and 4.7" — an honest statement about consistency of delivery.
  4. It gives a principled forest plot. Laying sections out with their confidence intervals and the pooled diamond makes outliers, precision, and heterogeneity legible at a glance to non-statistical committees.

Limitations and honest caveats

Meta-analysing your own data inherits meta-analysis's well-known hazards, and a critical reader will raise them.

  • Heterogeneity is estimated poorly with few units. With a handful of sections, τ² and I² are themselves very imprecise; a confident-looking I² of 70% from five sections should be read as a rough signal, not a fact. Report the uncertainty in the heterogeneity estimate, and prefer methods (e.g., Hartung–Knapp adjustment) that behave better with few units.
  • Garbage in, garbage out. Pooling does not fix biased inputs. If several sections share the same confound — a grading-leniency culture, a timetable disadvantage — meta-analysis will average the bias, not remove it. The unit-level validity work still has to happen first.
  • It describes, it does not causally explain. A meta-regression relating section effects to moderators (class size, modality) is observational; it generates hypotheses about why sections differ, not proof.
  • Independence assumptions can be violated. Sections taught by the same instructor, or successive terms of one course, are correlated; treating them as independent studies overstates precision. Multilevel or three-level meta-analytic models handle this and should be used when the nesting is strong — which is also why this overlaps, but does not replace, hierarchical modelling of nested evaluation data.

How Koji incorporates this

Koji is designed to hold evaluation data in exactly the structure meta-analysis needs and to report it without manufacturing precision. Because Koji organises responses by section, cohort, and term as first-class units, the platform can treat each as its own small study — carrying not just a mean but the standard error and response count that proper precision-weighting requires, rather than collapsing everything into one undifferentiated average. Koji's bias-aware reporting is built to surface heterogeneity as a finding — flagging when sections of the "same" course genuinely diverge — instead of papering over it with a single headline number, and its quality scoring guards against letting low-quality, low-response sections distort the pooled picture. When real between-section spread appears, Koji's automatic thematic analysis of each section's open-text can help explain why the outliers differ, turning a heterogeneity statistic into an actionable diagnosis. Koji is designed to mitigate the two classic errors — big-section dominance and false precision — that naive pooling produces; it does not remove shared confounds or turn observed heterogeneity into causal explanation. Programmes running multi-study or multi-cohort research beyond course evaluation can apply the same synthesis-friendly engine on Koji's core platform at koji.so.

A worked example

A course has run eight times, with section means from 3.9 to 4.6 and response counts from 18 to 240. Leadership reports the pooled mean, 4.35, and moves on. A random-effects meta-analysis tells a richer story: I² is 68%, so most of the spread is real, not sampling noise; the confidence interval around the mean is a tight 4.25–4.45, but the prediction interval is 3.85–4.75. The honest conclusion is not "the course scores 4.35" but "the course averages about 4.3, yet delivery is inconsistent enough that the next section could plausibly land anywhere from the high-3s to the high-4s" — a quality-assurance flag that the grand mean actively concealed.

Frequently asked questions

How is this different from empirical-Bayes shrinkage of instructor scores?

They are close cousins. Empirical-Bayes shrinkage pulls each unit's estimate toward the overall mean to produce better individual scores. Random-effects meta-analysis focuses on the overall estimate, the amount of heterogeneity, and the prediction interval for a new unit. Both rest on the same variance-components logic and can be used together.

Isn't this the same as the Uttl or Cohen meta-analyses you cite elsewhere?

No. Those are external meta-analyses synthesising many published studies about SET validity. This article is about synthesising your own sections or terms as if each were a small study — an internal quality-assurance technique, not a literature review.

Why not just pool all students into one big mean?

Because that weights sections purely by size, letting a large lecture dominate, and it assumes every section estimates one fixed truth. If sections genuinely differ, a single pooled mean hides the heterogeneity that is often the most important finding.

What is the difference between the confidence interval and the prediction interval?

The confidence interval describes uncertainty about the average effect and narrows as you add sections. The prediction interval describes the range a new section is likely to fall in and incorporates real between-section spread, so it stays wide when delivery is genuinely inconsistent.

How many sections do I need?

More is better, because heterogeneity is hard to estimate from few units. With only three to five sections, treat I² and the prediction interval as rough guidance and consider the Hartung–Knapp adjustment, which is more reliable with small numbers of studies.

What if the same instructor teaches several of the sections?

Then the sections are not independent, and standard meta-analysis overstates precision. Use a three-level or multilevel meta-analytic model that nests sections within instructors, mirroring the hierarchical structure of the data.

References

Related Resources

Related articles

analysis-reporting

How Many Responses Do You Need for a Reliable Course Evaluation?

Reliability of a class-average rating rises steeply with the number of students who respond. Generalizability theory — from Marsh''s syntheses to Dzakadzie (2026) — shows raters matter more than items, and a class of 12 cannot be evaluated like a class of 120.

analysis-reporting

Generalizability Theory and the Reliability of Student Ratings: Why One Class Is Not Enough

Reliability is not one number. Generalizability theory (Gillmore, Kane & Naccarato 1978; Marsh 1984) decomposes the variance in student ratings into student, teacher, course and occasion components — and shows that a single class can be reliable for the course yet a poor estimate of the teacher. What that means for fair evaluation.

analysis-reporting

The 4.2 vs 4.4 Trap: Why Small Differences in Course-Evaluation Means Are Usually Noise

Faculty and administrators routinely read meaning into tiny gaps between course-evaluation means. Boysen (2015) and Boysen et al. (2014) show this happens even when confidence intervals say the difference is nothing — and that warnings barely help. How to report uncertainty honestly.

analysis-reporting

Stop Comparing Raw Averages: Empirical-Bayes Shrinkage for Course-Evaluation Scores

Why raw course-evaluation means mislead — especially for small classes — and how empirical-Bayes shrinkage (partial pooling), the method Kane and Staiger applied to teacher value-added, produces fairer, more accurate instructor estimates.