New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
analysis-reporting10 min read

Did the Teaching Change Actually Work? Interrupted Time Series for Course-Evaluation Trends

Comparing this year''s evaluation mean to last year''s cannot tell you whether a curriculum redesign worked. Interrupted time series with segmented regression can — here is how to apply it, and where it breaks.

Koji Education Team

Product

Answer in brief. If a department redesigns a course and its evaluation score rises the next term, a simple before-versus-after comparison cannot tell you whether the redesign caused the rise, whether an existing upward trend continued, or whether the score simply regressed toward the mean. Interrupted time series (ITS) analysis with segmented regression — the strongest quasi-experimental design for a single-group intervention measured repeatedly over time (Wagner et al. 2002; Bernal, Cummins & Gasparrini 2017) — separates a genuine level or slope change at the point of the intervention from the pre-existing trend. For course evaluations it is the honest alternative to "the number went up, so it worked," provided you have enough pre-intervention terms and you handle autocorrelation, seasonality and regression to the mean.

What the research says

The core problem ITS solves is old and well understood in survey and program evaluation: a pre–post comparison of two time points confounds the intervention with the trend the series was already on. If evaluations were drifting upward for unrelated reasons — grade inflation, cohort change, an easier assessment — a post-intervention rise looks like success even when the change did nothing.

The canonical methods reference is Wagner, Soumerai, Zhang and Ross-Degnan (2002), "Segmented regression analysis of interrupted time series studies in medication use research," Journal of Clinical Pharmacy and Therapeutics, 27(4), 299–309. They describe ITS as "the strongest quasi-experimental design" for evaluating the longitudinal effects of an intervention introduced at a known point, and lay out segmented regression as the estimation tool. The model fits a regression line to the outcome across time with terms that let the line change at the intervention:

  • a baseline level (the intercept),
  • a baseline trend (the pre-intervention slope over time),
  • a level change immediately after the intervention (a step up or down), and
  • a trend change (a change in slope) after the intervention.

The intervention "works" only if the level change and/or the post-intervention slope change are statistically and substantively distinguishable from the continuation of the baseline trend — not merely if the post period looks higher than the pre period.

The most-cited applied tutorial is Bernal, Cummins and Gasparrini (2017), "Interrupted time series regression for the evaluation of public health interventions: a tutorial," International Journal of Epidemiology, 46(1), 348–355. It walks through the segmented-regression model and, crucially for a careful analyst, catalogues the design''s real threats: over-dispersion, autocorrelation (successive observations are correlated, which deflates standard errors if ignored), seasonality, and time-varying confounders that happen to coincide with the intervention. It also stresses the practical requirement that you need a sufficient number of data points before and after the interruption — a handful of terms will not support credible slope estimation.

A useful corrective from the same literature is the methodological note by Turner and colleagues (2020) in the International Journal of Epidemiology on a common parameterisation error in segmented regression — mis-specifying the time and intervention terms so that the "level change" is estimated at the wrong point. It is a reminder that ITS is easy to run and easy to get subtly wrong.

Why it matters for course evaluation in practice

Quality-assurance work is full of interventions applied at a known moment: a curriculum redesign, a new assessment structure, a teaching-and-learning grant, a switch to blended delivery, a change of coordinator. The standard evidence offered to an accreditation panel or a programme committee is a before-and-after mean. ITS turns that anecdote into a defensible causal-ish claim.

  • It distinguishes a real step from a continuing drift. If evaluations were already climbing, ITS attributes only the extra jump at the intervention to the intervention.
  • It captures slope as well as level. Some interventions do not produce an instant jump but bend the trajectory — a mentoring scheme whose effect accumulates. A two-point comparison is blind to this; segmented regression models it directly.
  • It disciplines the "it improved" narrative. For programme review under ESG/ENQA or national frameworks, showing you can separate intervention effects from trend is exactly the methodological maturity panels reward. See our companion note on turning evaluation data into ESG/ENQA accreditation evidence.
  • It guards against being fooled by regression to the mean. A course flagged because it scored unusually low will usually "recover" next year even if nothing changed; ITS with an adequate baseline makes that artefact visible rather than claimable as a win. See regression to the mean in course evaluations.

Limitations and honest caveats

ITS is stronger than pre–post, but it is still quasi-experimental, and a sceptical reader should hold it to account.

  • It needs many time points. Segmented regression on three terms before and two after is not ITS in any meaningful sense; you typically want on the order of eight or more pre-intervention observations to estimate a baseline trend with any stability. Many course-level series are simply too short, which is why ITS is often better applied at programme or department level, where more repeated measures exist.
  • Autocorrelation and seasonality are not optional. Course evaluations have obvious periodicity (autumn versus spring cohorts, exam-timing effects). Ignoring seasonal structure or serial correlation produces standard errors that are too small and false "significant" effects. Bernal et al. are explicit about adjusting for these.
  • Co-interventions confound. If the redesign coincided with a new lecturer, a room change and a grading-policy shift, ITS cannot disentangle them — it estimates the effect of everything that happened at that time point.
  • The outcome is still a student rating. ITS cleans up the temporal inference; it does nothing to fix the underlying validity problems of student evaluations of teaching (bias, the weak link to learning). A well-estimated change in a biased measure is still a change in a biased measure.
  • Parameterisation is error-prone. As Turner et al. (2020) show, mis-coding the intervention indicator silently biases the level-change estimate. ITS should be run by someone who can read the model, not pasted from a template.

How Koji incorporates this

Koji for Education is designed so that the inputs to a credible ITS analysis actually exist and are comparable across time — which is where most institutions fail before any statistics begin.

  • Consistent, versioned instruments over time. Because Koji administers the same structured question set (scale, single_choice, open_ended and others) across cycles, the series you feed into segmented regression is measuring the same construct each term rather than a moving target. Comparability is the precondition ITS assumes.
  • Frequent, mid-cycle collection increases the number of data points. Koji supports formative, mid-semester collection as well as end-of-term, which lengthens the time series and improves the feasibility of estimating a baseline trend — directly addressing the "too few observations" limitation.
  • Structured export for analysis. Koji exports tidy, time-stamped, cohort-labelled data so an institutional-research analyst can fit a segmented-regression model (including seasonal and autocorrelation terms) in R or Stata without hand-reconciling spreadsheets.
  • Triangulation to offset the biased-measure caveat. Koji''s automatic thematic analysis of open-text responses lets you check whether a modelled level change in the numeric score is corroborated by a shift in what students actually say — a qualitative cross-check on the quantitative interruption.
  • Closing-the-loop action tracking records when an intervention was introduced, giving you the defensible interruption point that ITS requires rather than a reconstructed guess.

Koji is designed to support rigorous longitudinal analysis, not to perform causal inference on your behalf — the modelling, and the judgement about confounds, remain the analyst''s. Koji''s core research platform at koji.so applies the same structured, longitudinally comparable collection to product and customer research, where interrupted time series is equally the right tool for judging whether a launch or a change actually moved a metric.

Related resources

A minimal specification, in words

The simplest useful model regresses the evaluation outcome on four terms: time (a counter of consecutive terms, capturing the baseline trend), an intervention indicator (0 before the change, 1 from the change onward, capturing an immediate level shift), a post-intervention time counter (0 before, then 1, 2, 3… after, capturing a change in slope), and the intercept for the baseline level. The coefficient on the intervention indicator is your estimated step change; the coefficient on the post-intervention time counter is your estimated change in trajectory. To this you add the adjustments the tutorials insist on: a seasonal term (for example, a spring-versus-autumn indicator) so cohort periodicity is not mistaken for an effect, and a correction for autocorrelation — such as Newey–West standard errors or an autoregressive error term — so your confidence intervals are honest. Plotting the fitted segments against the observed points, with the counterfactual continuation of the baseline trend drawn through the post period, is the single most persuasive figure you can put in front of a quality committee: it shows the effect as the gap between what happened and what the trend alone would have predicted.

References

  • Wagner, A. K., Soumerai, S. B., Zhang, F., & Ross-Degnan, D. (2002). Segmented regression analysis of interrupted time series studies in medication use research. Journal of Clinical Pharmacy and Therapeutics, 27(4), 299–309. https://doi.org/10.1046/j.1365-2710.2002.00430.x
  • Bernal, J. L., Cummins, S., & Gasparrini, A. (2017). Interrupted time series regression for the evaluation of public health interventions: a tutorial. International Journal of Epidemiology, 46(1), 348–355. https://doi.org/10.1093/ije/dyw098
  • Turner, S. L., Karahalios, A., Forbes, A. B., Taljaard, M., Grimshaw, J. M., Cheng, A. C., Bero, L., & McKenzie, J. E. (2020). Design characteristics and statistical methods used in interrupted time series studies evaluating public health interventions: a review. Journal of Clinical Epidemiology, 122, 1–11. https://doi.org/10.1016/j.jclinepi.2020.02.006