Build a Synthetic Comparison Department: Synthetic Control Methods for a Teaching Change
When one programme redesigns its teaching, you rarely have a clean control group. Synthetic control methods build a weighted "synthetic" comparator from other programmes so you can estimate whether the change actually moved evaluation scores.
Koji Education Team
Product
In brief
When a single department, programme, or module redesigns its teaching and you want to know whether that change caused evaluation scores to move, you almost never have a randomised control group. The synthetic control method (SCM) solves this by constructing a synthetic comparison unit — a weighted average of other, untouched programmes chosen so that their combined pre-change evaluation trajectory closely tracks the treated programme. The post-change gap between the real unit and its synthetic twin is the estimated effect. SCM is the most credible quasi-experimental tool when you have one treated unit, many candidate comparators, and several terms of comparable pre-change data — a situation that describes most curriculum reforms in higher education.
What the research says
The synthetic control method was introduced by Alberto Abadie and Javier Gardeazabal (2003) to estimate the economic cost of terrorism in the Basque Country. They built a synthetic Basque Country from a weighted combination of other Spanish regions that matched its pre-conflict economic path, and found per-capita GDP fell roughly 10 percentage points relative to that synthetic comparator after terrorism intensified.
Abadie, Diamond, and Hainmueller (2010) formalised the estimator in the Journal of the American Statistical Association and applied it to California Proposition 99, a 1988 tobacco-control programme. Their synthetic California — a convex combination of other US states — tracked California cigarette sales almost exactly before 1988 and then diverged, implying annual per-capita sales about 26 packs lower by 2000 than California would have seen without the programme. Two design features make the estimator trustworthy: the donor weights are non-negative and sum to one, so the synthetic unit is an interpolation of real units rather than an extrapolation; and inference is done with placebo (permutation) tests — the method is re-run pretending each untreated donor was treated, and the real gap is judged unusual only if it lies in the tail of the placebo gaps.
Abadie (2021), reviewing the method in the Journal of Economic Literature, sets out the conditions under which SCM is defensible: a good pre-treatment fit over a reasonably long window, a donor pool of genuinely comparable untreated units, the treated unit lying inside the "convex hull" of donor characteristics, no anticipation of the intervention, and no interference (spillovers) between treated and donor units. Where difference-in-differences assumes parallel trends between a single treated and single control group, SCM builds the counterfactual trend from data and reports transparently which donors carry weight.
Why it matters for course evaluation in practice
Curriculum and pedagogy changes in universities are almost always aggregate, one-off, and un-randomised: a chemistry department flips to active learning; a business school introduces a new assessment regime; a faculty adopts a new mid-cycle feedback policy. You cannot randomise students across the reformed and unreformed versions, and you rarely have a single obviously-parallel comparator. That is exactly the setting SCM was built for.
Concretely, treat the reformed programme as the treated unit; assemble a donor pool of other programmes at your institution (or peer institutions) that did not change; use several terms of pre-change evaluation data — overall satisfaction, teacher-clarity items, workload — plus predictors such as class size, discipline, and student composition; and let the algorithm find the weighted blend of donors whose pre-change evaluation trajectory best matches the reformed programme. If the real programme and its synthetic twin move together for years and then separate after the reform, you have a far more credible causal story than a naive before-and-after mean, and a far more honest one than comparing the reformed programme to a single hand-picked "similar" department. The placebo test tells you whether the post-change gap is larger than what you would see by chance across programmes that changed nothing.
Limitations and honest caveats
SCM is powerful but easy to misuse, and a critical reader will press on several points. First, it needs a long, stable pre-period. A convincing fit over two or three terms can be coincidence; ten-plus terms are safer, and course-evaluation series are often short. Second, interpolation bias: if the reformed programme is extreme (the largest cohort, the only clinical programme), no convex combination of donors can match it, and the synthetic unit is unreliable — Abadie (2021) is explicit that the treated unit must sit inside the donor convex hull. Third, donor-pool contamination: if reforms spill over (shared instructors, a faculty-wide policy that quietly touched "control" programmes), the no-interference assumption fails and the effect is biased. Fourth, instrument drift: synthetic control assumes the outcome is measured the same way over time; if the evaluation form, scale, or administration mode changed mid-series, the "effect" may be a measurement artefact. Fifth, inference is permutation-based, not large-sample — with a small donor pool, placebo tests have limited power, and a non-significant result may simply reflect too few donors. Finally, SCM estimates the effect on the rating, which — as the wider evaluation literature stresses — is not the same as the effect on learning. Report SCM as one triangulated strand of evidence, never as proof on its own.
How Koji incorporates this
Koji does not silently run a synthetic control behind a dashboard button, and we would distrust any vendor who claimed to — SCM demands analyst judgement about the donor pool and pre-period. What Koji does is make the analysis feasible and defensible, which is usually the binding constraint:
- Comparable longitudinal series. Koji stores per-cohort, per-course results as a structured time series using stable question types (scale, single_choice, yes_no), so the multi-term pre-change trajectory SCM needs actually exists and is comparable across programmes — rather than being scattered across incompatible annual PDFs.
- Instrument stability. Because Koji versions its questionnaires, you can hold the evaluation instrument constant across the pre/post window (or flag exactly when it changed), directly addressing the measurement-drift threat above.
- Rich predictors and covariates. Alongside ratings, Koji captures the structured context (cohort size, delivery mode, discipline) that SCM uses to build a good pre-treatment match.
- Beyond the number. Koji's AI-moderated conversational interviews and automatic thematic analysis of open text let you check why a gap opened — did students describe the specific reform, or something else happening that term? — which is how you probe the no-interference and confounding assumptions rather than assuming them away.
- Export for estimation. Koji is designed to export clean, analysis-ready data so your institutional-research team can run the estimator in R (the
Synthortidysynthpackages) and report the placebo distribution transparently.
The same AI-moderated interview engine powers Koji's core research platform at koji.so for product and customer research, where teams face the identical problem of estimating the effect of a single launch with no clean control.
Frequently asked questions
How is synthetic control different from difference-in-differences? Difference-in-differences compares one treated group to one (or a simple average of) control group(s) and assumes their trends would have stayed parallel. Synthetic control constructs the comparator as a data-driven weighted blend of many donors chosen to match the treated unit's pre-change trajectory, and it reports which donors carry weight — a more transparent and often more credible counterfactual when no single comparator is obviously parallel.
How many pre-change terms do I need? There is no hard rule, but a fit built on only two or three terms is fragile. Aim for enough pre-period observations that a close match is unlikely to be coincidence — Abadie (2021) stresses that a good, extended pre-treatment fit is the main evidence that the synthetic unit is a valid counterfactual.
How do I get a p-value if there is only one treated programme? Through placebo tests: re-run the method treating each untreated donor as if it had been reformed, collect the distribution of post-period gaps, and see whether your real programme's gap is extreme relative to that distribution. This is a permutation-style inference, not a conventional large-sample test, and its power depends on having enough donors.
Can I use synthetic control if a faculty-wide policy affected every programme? No. Synthetic control requires untouched donor units. If the change reached the comparators — even indirectly, through shared staff or a blanket policy — the no-interference assumption fails and the estimated effect is biased. You need donors that genuinely did not experience the change.
Does a positive synthetic-control result prove the teaching change improved learning? No. It estimates the effect on evaluation scores, which the wider literature shows are only weakly related to learning gains. Treat it as one strand of causal evidence to be triangulated with direct learning outcomes, peer observation, and qualitative feedback.
Related resources
- Difference-in-Differences for Course Evaluation
- Interrupted Time Series for Course-Evaluation Trends
- Regression Discontinuity for Course-Evaluation Cutoff Effects
- Propensity Score Matching for Course-Evaluation Confounds
- Contribution Analysis: Credible Causal Claims from Course Evaluation
- Specification-Curve Analysis for Course Evaluation
References
- Abadie, A., & Gardeazabal, J. (2003). The Economic Costs of Conflict: A Case Study of the Basque Country. American Economic Review, 93(1), 113–132. https://doi.org/10.1257/000282803321455188
- Abadie, A., Diamond, A., & Hainmueller, J. (2010). Synthetic Control Methods for Comparative Case Studies: Estimating the Effect of California's Tobacco Control Program. Journal of the American Statistical Association, 105(490), 493–505. https://doi.org/10.1198/jasa.2009.ap08746
- Abadie, A. (2021). Using Synthetic Controls: Feasibility, Data Requirements, and Methodological Aspects. Journal of Economic Literature, 59(2), 391–425. https://doi.org/10.1257/jel.20191450
Related articles
Did the Teaching Change Actually Work? Interrupted Time Series for Course-Evaluation Trends
Comparing this year''s evaluation mean to last year''s cannot tell you whether a curriculum redesign worked. Interrupted time series with segmented regression can — here is how to apply it, and where it breaks.
Propensity Score Matching for Course-Evaluation Confounds
How propensity score matching (PSM) adjusts student-evaluation-of-teaching comparisons for confounds like class size, course level, discipline, and grading — and where it honestly falls short.
Did the Teaching Change Cause the Score to Move? Difference-in-Differences for Course Evaluation
Difference-in-differences lets you estimate whether a course redesign actually moved evaluation scores by comparing a treated course against a similar untouched one over time. Here is how the design works, when its parallel-trends assumption holds, and how to use it honestly in quality assurance.
One Instructor, Two Hundred Defensible Scores: Specification-Curve Analysis for Course Evaluation
Every course-evaluation "score" is the product of dozens of defensible analytic choices — which items count, how to weight, mean or median, whether to adjust for class size. Specification-curve and multiverse analysis compute the result across all of them, so a personnel decision rests on the pattern, not on one analyst's arbitrary path.