New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Methodology9 min read

Did Your Teaching Change Actually Work? Quasi-Experimental Designs for Course Evaluation

A score that went from 3.8 to 4.2 after you redesigned a module is not evidence the redesign worked. Comparison cohorts, difference-in-differences, and interrupted time series tell you what a single before-and-after number never can.

Koji Education Team

Product ·

The short answer: When an evaluation score rises after a teaching change, the rise is not proof the change caused it. Scores drift, cohorts differ, and low scores rebound on their own. To make a defensible causal claim you need a counterfactual — what would the score have been without the change. Three quasi-experimental designs supply one: a comparison cohort, difference-in-differences, and (comparative) interrupted time series. None require a randomized trial, and all are within reach of an ordinary quality office.

The before-and-after trap

Here is the most common evaluation story in higher education. A module scores 3.8. The programme director redesigns the assessment, adds formative checkpoints, reruns the module, and the score comes back 4.2. The redesign goes into the annual report as a success and gets rolled out to three more modules.

The problem is that a single before-and-after comparison cannot distinguish a real effect from at least four rival explanations:

  • Regression to the mean. A low score is partly bad luck — a hard cohort, a bad term, sampling noise. Left alone, extreme scores tend to move back toward the average on their next measurement, with no intervention at all. We unpack this in Did Your Teaching Improve, or Is It Just Regression to the Mean?
  • Cohort differences. This year's students are not last year's. Prior attainment, motivation, and even the weather of a term shift the baseline.
  • Secular trends. If satisfaction is creeping up across the whole faculty — a new VLE, a post-pandemic recovery — your module would have risen anyway.
  • Measurement noise. With 25 respondents, a 0.4-point move can sit comfortably inside the margin of error. See Is a 0.3-Point Difference Real?

A before-and-after design controls for none of these. To claim the change caused the improvement, you need an estimate of the counterfactual: the score the redesigned module would have shown if you had changed nothing.

Design 1: A comparison cohort (the minimum viable counterfactual)

The simplest upgrade is to find a group that did not get the change but is otherwise similar — a parallel module, another campus running the old design, the same module taught by a colleague who kept the old format. If the redesigned module rose from 3.8 to 4.2 and the comparison module sat flat at 3.9 across the same period, the case strengthens. If the comparison module also rose to 4.2, your redesign explains nothing — something faculty-wide did.

The threat to this design is selection: the comparison group may differ from the treated group in ways that also affect the outcome. You manage it by choosing the most similar comparator you can and by checking that the two groups looked alike before the change.

Design 2: Difference-in-differences

Difference-in-differences (DiD) formalises the comparison-cohort logic. You measure both groups before and after, then take the difference of the two differences. The treated group's change minus the comparison group's change is your estimated effect — and the subtraction cancels out anything that moved both groups equally (a faculty-wide trend, a calendar effect, a platform change).

DiD rests on one central, testable assumption: parallel trends. The two groups must have been moving in parallel before the intervention. If the redesigned module was already climbing faster than the comparator, DiD will mistake that pre-existing trajectory for an effect. The honest move is to plot several pre-intervention periods and show the lines were parallel before you intervened.

Design 3: (Comparative) interrupted time series

If you have a run of scores over many terms, an interrupted time series (ITS) uses the module's own history as the counterfactual. You model the pre-intervention trend, project it forward, and test whether the post-intervention scores depart from that projection — in level (an immediate jump) or slope (a changed trajectory). Segmented regression is the standard estimator (BMJ-indexed methods overview, Bernal et al., 2017).

ITS is more demanding than DiD: you typically want at least four pre-intervention time points to estimate the baseline trend reliably, which is not always feasible for an annual module (MDRC, Comparative ITS vs DiD). Its strongest form is the comparative ITS (CITS), which adds a comparison series and controls for both baseline level and trend differences. Within-study comparisons against randomized benchmarks in education found CITS bias of roughly 0.03 standard deviations across four studies — evidence that a well-conducted CITS can approximate experimental estimates (Journal of Research on Educational Effectiveness, 2022).

But isn't this overkill for a course evaluation?

A fair challenge. Three honest answers:

  • Match the rigour to the stakes. If you are privately tinkering with your own seminar, a before-and-after glance is fine — you are not making a generalizable causal claim. The moment a result is used to roll out a change across a programme, justify a resource decision, or appear in an accreditation self-assessment, the claim becomes causal and deserves a counterfactual. Quasi-experimental designs are for consequential claims.
  • These designs have real assumptions that can fail. Parallel trends can be violated; ITS can be confounded by a co-occurring change (you redesigned the assessment and changed rooms and the exam timetable moved). Naming the threat is part of the method, not a footnote. Non-experimental evaluations of school programmes can still carry meaningful bias when the comparison is poorly chosen (Wong et al., quantifying bias in non-experimental evaluations).
  • They do not need a statistician for every module. A comparison cohort is arithmetic. DiD is two subtractions. The discipline is in the design — deciding the comparator and the timing before you run the change — not in exotic computation.

What this means for how you collect evaluations

Quasi-experimental analysis is only possible if your data collection is built for it. Three practical implications:

  • Keep a stable core of items over time. ITS and DiD need comparable measurements across terms. If you rewrite the questionnaire every year, you destroy your own counterfactual. Pair a stable longitudinal core with rotating probes.
  • Capture enough context to build a credible comparator. Cohort size, prior attainment, modality, and discipline are what let you argue two groups are comparable. Bare star ratings cannot support a DiD.
  • Measure mid-cycle, not just at the end. Formative, mid-cycle collection gives you the extra time points ITS needs and lets you intervene within a term rather than only between years.
  • Beware multiple comparisons. If you test twenty modules for "significant" change, one will look significant by chance. Plan your comparisons in advance — see the multiple-comparisons trap.

Where Koji fits

A causal claim is only as good as the measurements feeding it, and this is where the static end-of-term survey quietly undermines its own analysis. Koji for Education is designed to produce the kind of data these designs depend on: a stable, standardized core of structured questions (the six question types — open_ended, scale, single_choice, multiple_choice, ranking, yes_no) that stays comparable term over term, plus formative mid-cycle collection that generates the additional time points an interrupted time series needs.

Because Koji runs AI-moderated conversational interviews, every wave is moderated to the same standard — removing the human-moderator drift that would otherwise masquerade as a "change" in your time series. Its programme- and institution-level reporting is built to compare cohorts and modules rather than just average one. And its automatic thematic analysis answers the question a DiD coefficient cannot: not just whether the score moved, but what students say changed — the mechanism behind the number. Koji does not turn observational data into a randomized trial; it gives you cleaner, comparable, time-stamped evidence so your counterfactual is defensible rather than wishful. (The main Koji platform applies the same engine to product and customer research, where the before-and-after trap is just as common.)

A score that rose after you changed something is the beginning of an investigation, not the end of one. The institutions whose evaluation data actually drives improvement are the ones that designed the counterfactual before they ran the change.

Build evaluation data that can support a real causal claim — explore Koji for Education.