Did the Teaching Change Cause the Score to Move? Difference-in-Differences for Course Evaluation
Difference-in-differences lets you estimate whether a course redesign actually moved evaluation scores by comparing a treated course against a similar untouched one over time. Here is how the design works, when its parallel-trends assumption holds, and how to use it honestly in quality assurance.
Koji Education Team
Product
Difference-in-differences (DiD) estimates the causal effect of a teaching or curriculum change by comparing how evaluation scores moved in the course you changed against how they moved in a similar course you left alone over the same period. The estimate is a simple double subtraction — the treated group's before-to-after change minus the comparison group's before-to-after change — and it works because the comparison group nets out the university-wide trends and one-off shocks that a naive before/after comparison would mistake for your intervention's effect. It is only as good as its central assumption: that, absent the change, both courses would have drifted in parallel.
Answer box (BLUF): If you want to claim that a redesign, a new textbook, or an active-learning switch caused your course-evaluation scores to rise, a before/after comparison of one course is not enough — a semester-wide dip, a cohort change, or regression to the mean can produce the same pattern. Difference-in-differences fixes this by subtracting the change in a comparable comparison course from the change in your treated course. The remaining difference is your best estimate of the intervention's effect, valid only if the two courses were on parallel trends before you intervened. Check that assumption with several pre-intervention periods; never assume it.
What the research says
Difference-in-differences is one of the most widely used quasi-experimental designs in the social and health sciences, precisely for situations — like curriculum reform — where a randomised controlled trial is impractical or unethical. The canonical modern illustration is Card and Krueger's (1994) study of a New Jersey minimum-wage increase: rather than simply looking at New Jersey employment before and after, they subtracted the contemporaneous change in neighbouring Pennsylvania, which had no wage change, to strip out region-wide economic trends. The logic transfers directly to a faculty of education asking whether a redesigned statistics module lifted its ratings.
The clearest methodological treatment for practitioners is Wing, Simon, and Bello-Gomez's (2018) Annual Review of Public Health paper, "Designing Difference in Difference Studies: Best Practices for Public Health Policy Research." They frame DiD not as a formula to run once, but as a design to be actively constructed: choosing a defensible comparison group, testing the parallel-trends assumption with pre-treatment data, and stress-testing the result with sensitivity analyses and placebo checks. Their core estimand — the "difference of differences" — is (Treated_after − Treated_before) − (Comparison_after − Comparison_before). In a regression form, you model the outcome on a treatment-group indicator, a post-period indicator, and their interaction; the coefficient on the interaction term is the DiD estimate, and standard errors are typically clustered at the unit level.
A crucial refinement comes from Goodman-Bacon (2021), "Difference-in-differences with variation in treatment timing" (Journal of Econometrics). When different courses adopt an intervention at different times — a staggered rollout, common in phased curriculum reform — the conventional two-way fixed-effects estimator can produce a misleading weighted average, because already-treated units end up acting as comparison units for later-treated ones. This is not a reason to abandon DiD; it is a reason to be careful about how you handle rollouts, and to prefer clean two-group, two-period comparisons or the newer robust estimators when timing varies.
Why it matters for course evaluation in practice
Quality-assurance teams constantly face causal questions dressed up as descriptive ones. "Our scores went up after we introduced peer instruction" is a causal claim in disguise. So is "the new assessment tanked satisfaction." Acting on either without a comparison group risks two expensive errors: crediting an intervention that did nothing, and scrapping a change that actually helped but happened to launch in a rough semester.
Three everyday confounds make the single-course before/after comparison unreliable, and DiD addresses all three:
- Institution-wide time trends. A university-wide LMS migration, an exam-timetable change, or a national mood shift moves everyone's scores. Subtracting a comparison course removes any shock that hit both.
- Regression to the mean. Interventions are often aimed at courses that scored badly last year — exactly the courses statistically likely to rebound anyway. This is a documented trap (see our companion article on regression to the mean); a comparison group of similarly low-but-untreated courses helps separate genuine improvement from mean reversion.
- Fixed differences between courses. A required quantitative module and an elective seminar have permanently different baseline ratings. DiD does not require them to have the same level — only the same trend — so it tolerates these fixed gaps.
The practical payoff is credibility. When a programme director tells an accreditation panel that a reform improved the student experience, "scores rose 0.4 points more than in our matched comparison cohort, on parallel pre-trends" is an argument that survives scrutiny. "Scores rose 0.4 points" is not.
Limitations and honest caveats
DiD is a strong design, but it is assumption-driven, and a critical reader will probe exactly where it can fail:
- Parallel trends is untestable for the period that matters. You can show the two courses moved together before the intervention, but you can never directly observe whether they would have continued to move together afterwards. Pre-trend evidence is necessary, not sufficient. Reporting several pre-periods, not one baseline, is the minimum honest standard.
- Composition changes break it. If the reform also changed who takes the course — attracting stronger students, or driving weak ones to drop — then the treated cohort is no longer the same population, and the "effect" partly reflects selection. Track enrolment and prior-attainment composition alongside the scores.
- Spillovers contaminate the comparison. If the redesigned module shares instructors, materials, or students with the comparison module, the treatment can leak, biasing the estimate toward zero.
- Ceiling effects and skew. Course-evaluation distributions pile up near the top of the scale. A course already at 4.7/5 cannot rise as much as one at 3.5, mechanically violating the equal-movement logic. Consider ordinal models rather than differencing raw means.
- Small numbers of units. DiD's inference assumptions strain when you have one treated course and one comparison; clustered standard errors behave badly with few clusters. Treat single-pair DiD as suggestive, and strengthen it with multiple comparison courses.
- Staggered timing. As Goodman-Bacon (2021) shows, phased rollouts analysed with naive two-way fixed effects can give the wrong sign. Match the estimator to the rollout.
- Satisfaction is not learning. DiD on evaluation scores tells you the intervention changed ratings. Whether it changed learning is a separate question that needs a learning outcome, not a satisfaction proxy.
None of these are fatal; each is a reason to design the study deliberately rather than run a spreadsheet subtraction and call it causal.
How Koji incorporates this
Koji is an evaluation platform, not a statistics package — but a good DiD depends far more on how the data was collected than on the final subtraction, and that is where the platform is designed to help.
- Aligned instruments across treated and comparison courses. DiD requires the same outcome, measured the same way, in both groups and both periods. Koji lets you deploy an identical structured instrument — the same
scale,single_choice, andopen_endeditems — across the treated course and a matched comparison course each cycle, so the two arms are genuinely comparable rather than stitched together from mismatched forms. - Multiple pre-period observations for a real parallel-trends check. Because Koji supports mid-cycle and formative collection, not just a single end-of-term snapshot, you can accumulate several pre-intervention data points and actually test whether the two courses were trending in parallel before you changed anything — the evidence Wing et al. (2018) insist on.
- Mechanism evidence the numbers cannot give. A DiD estimate says the score moved; it does not say why. Koji's AI-moderated conversational interviews probe beyond the Likert number, surfacing whether students attribute a change to the specific thing you altered. That qualitative mechanism evidence is what turns a plausible correlation into a defensible causal story.
- Triangulation against non-survey outcomes. To avoid mistaking a satisfaction bump for a learning gain, Koji's reporting is built to sit alongside learning-analytics and outcome data, so a DiD on ratings can be cross-checked against a DiD on something that isn't a rating.
- A documented definition of the "treatment." DiD needs a precise, dated statement of what changed. Koji's closing-the-loop action tracking records exactly which change was introduced and when — the audit trail that lets a panel see your treatment was a real, bounded intervention rather than a vague "we tried some things."
Koji is designed to support credible causal inference, not to manufacture it: the platform frames these estimates as "designed to mitigate confounding," never as proof. The same AI-moderated interview engine powers Koji's core research platform at koji.so for product and customer research, where before/after comparisons against a control cohort raise identical design questions.
Frequently asked questions
What exactly is the difference-in-differences estimate? It is a double subtraction: the change in the outcome for the group that received the intervention, minus the change in the outcome for a comparable group that did not, over the same time window. That second subtraction removes any trend or shock that affected both groups, leaving your best estimate of the intervention's own effect.
What is the parallel-trends assumption, and how do I check it? It is the assumption that, in the absence of the intervention, the treated and comparison courses would have moved by the same amount. You cannot verify it for the post-period, but you build confidence by showing the two courses tracked each other across several pre-intervention periods. One baseline point is not enough; you need a trend.
Can I use difference-in-differences with just one course before and after? No — a single course with no comparison group is exactly the before/after design DiD is meant to replace, because it cannot separate your intervention from semester-wide trends or regression to the mean. You need at least one comparison course that did not receive the change.
How is DiD different from an interrupted time series? An interrupted time series models a single unit's trajectory before and after a change and looks for a break in level or slope. DiD instead compares two units and relies on a control group rather than on modelling the trend. They answer the same causal question with different assumptions, and are often strongest used together.
Does a difference-in-differences result prove the change caused the score to move? It provides credible causal evidence if its assumptions hold — parallel pre-trends, no differential composition change, no spillovers. It is stronger than a raw before/after comparison, but it is a quasi-experiment, not a randomised trial, so it should be reported with its caveats, not as proof.
Should I run DiD on satisfaction scores or on learning outcomes? Ideally both, kept separate. A DiD on evaluation scores tells you the intervention changed how students rated the course; a DiD on a learning outcome tells you whether it changed what they learned. Conflating the two is how institutions end up optimising for satisfaction at the expense of learning.
Related resources
- Interrupted Time Series for Course-Evaluation Trends
- Propensity Score Matching for Course-Evaluation Confounds
- Contribution Analysis: Making Credible Causal Claims
- Regression to the Mean in Course Evaluations
- Emergency Remote Teaching as a Natural Experiment in Confounding
- Simpson's Paradox in Course-Evaluation Data
References
- Wing, C., Simon, K., & Bello-Gomez, R. A. (2018). Designing difference in difference studies: Best practices for public health policy research. Annual Review of Public Health, 39, 453–469. https://doi.org/10.1146/annurev-publhealth-040617-013507
- Card, D., & Krueger, A. B. (1994). Minimum wages and employment: A case study of the fast-food industry in New Jersey and Pennsylvania. American Economic Review, 84(4), 772–793. https://www.jstor.org/stable/2118030
- Goodman-Bacon, A. (2021). Difference-in-differences with variation in treatment timing. Journal of Econometrics, 225(2), 254–277. https://doi.org/10.1016/j.jeconom.2021.03.014
- Dimick, J. B., & Ryan, A. M. (2014). Methods for evaluating changes in health care policy: The difference-in-differences approach. JAMA, 312(22), 2401–2402. https://doi.org/10.1001/jama.2014.16153
Related articles
Why a Low-Scoring Course Usually 'Improves' Next Year Even If Nothing Changed: Regression to the Mean in Course Evaluations
Year-over-year movements in course-evaluation scores are mostly statistical noise regressing toward an instructor's true average. Here is why regression to the mean fools quality-assurance committees, and how to read evaluation trends responsibly.
Did the Teaching Change Actually Work? Interrupted Time Series for Course-Evaluation Trends
Comparing this year''s evaluation mean to last year''s cannot tell you whether a curriculum redesign worked. Interrupted time series with segmented regression can — here is how to apply it, and where it breaks.
Propensity Score Matching for Course-Evaluation Confounds
How propensity score matching (PSM) adjusts student-evaluation-of-teaching comparisons for confounds like class size, course level, discipline, and grading — and where it honestly falls short.
Did COVID Lower Course Evaluations? Emergency Remote Teaching as a Natural Experiment in Confounding
What student-evaluation data from the 2020 emergency shift to remote teaching reveals about how much SET scores reflect factors outside an instructor's control — and why pandemic-era ratings need a giant asterisk in any personnel decision.