New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods10 min read

Who Dropped Out of Your Longitudinal Evaluation — and Did It Break the Comparison? Differential Attrition and Panel Mortality

In any pre/post or multi-wave evaluation, the students who stop responding are rarely a random subset — and when they leave one condition faster than another, differential attrition can manufacture an effect that was never there.

Koji Education Team

Product

The short answer

In a longitudinal evaluation — mid-semester versus end-of-semester, pre versus post an intervention, or a follow-up months later — the danger is not that some students drop out, but that they drop out unevenly. When the students who disappear from one wave, or from one condition, differ systematically from those who stay, your before/after comparison is quietly comparing two different populations. This is differential attrition (also called panel mortality), and it is one of the most underappreciated threats to internal validity in course evaluation. It can create an apparent improvement — or destroy a real one — without anyone touching the teaching.

BLUF for busy readers: Overall dropout weakens a panel study; differential dropout can reverse its conclusion. If disengaged students leave your end-of-term survey faster than engaged ones, end-of-term scores rise on their own. Always report attrition rates by wave and by condition, test whether leavers differ from stayers, and treat a large difference in attrition rates between groups as a red flag — the What Works Clearinghouse treats even a few percentage points of differential attrition as potentially disqualifying.

What the research says

The most vivid modern demonstration comes from Zhou and Fishbach (2016) in the Journal of Personality and Social Psychology. Running standard experimental paradigms on online samples, they observed attrition rates of 30% to 50%, and — crucially — those rates differed across experimental conditions. Because participants abandoned the harder or less pleasant condition at a higher rate, the survivors in each condition were no longer comparable, even though assignment had originally been random. The unattended attrition introduced confounds that led the authors to "surprising (yet false)" conclusions — for instance, that recalling a few happy events is more effortful than recalling many. The lesson is stark: randomisation protects you only if everyone you randomised is still in the data. Selective drop-out silently un-randomises an experiment.

The methodological groundwork is older. Goodman and Blum (1996), in the Journal of Management, laid out a still-used procedure for diagnosing attrition's damage. They advise checking four things: (1) whether drop-out is non-random, using logistic regression to predict who leaves from their wave-1 characteristics; (2) whether means differ between leavers and stayers on the study's key variables; (3) whether attrition has restricted or inflated variances; and (4) whether it has changed the relationships among variables. In their own employed-adult panel, attrition did produce non-random sampling and shifted some means and variances — but, reassuringly, did not distort the correlations among variables. That nuance matters: attrition can bias some estimates (a mean satisfaction score) while leaving others (a correlation between clarity and satisfaction) intact. You have to check, not assume.

Finally, the evaluation world has codified how much attrition is too much. The What Works Clearinghouse Standards Handbook evaluates studies on both overall attrition and differential attrition, because both contribute to potential bias. Under its cautious threshold, an overall attrition rate near zero still tolerates only about 5.7 percentage points of differential attrition before a study is rated "high attrition"; under the more optimistic threshold the boundary is about 10 percentage points. The exact numbers trade off against overall attrition on a published boundary, but the headline is that a difference of a handful of percentage points between your two groups is enough to put the whole comparison in question.

Why it matters for course evaluation in practice

Course evaluation is riddled with longitudinal designs, and every one of them is exposed:

  • Mid-cycle to end-of-cycle tracking. If you collect formative feedback at week 6 and summative feedback at week 12, the students most likely to vanish by week 12 are the disengaged, the failing, and the withdrawn. Because those are disproportionately the harsher raters, end-of-term averages can rise from week 6 to week 12 even if nothing improved. A naïve reading celebrates a mid-course "fix" that is pure survivorship.
  • Pre/post intervention studies. You pilot a new teaching approach, survey before and after, and compare. If the intervention frustrates weaker students into dropping the post-survey (or the course), your post sample is enriched with the students who were always going to rate it highly. The intervention looks effective; the attrition did the work.
  • Comparing two sections or two modalities. If an online section loses respondents faster than the face-to-face section, a difference in end-of-term scores partly reflects who remained, not what was taught. This is the course-evaluation version of Zhou and Fishbach's condition-dependent attrition.
  • Alumni and longer-horizon follow-up. Evidence on the longitudinal stability of student ratings depends on who is still reachable years later — and reachability correlates with graduation, achievement, and goodwill toward the institution.

The critical distinction from ordinary non-response is that panel attrition is about losing the same individuals across waves. Cross-sectional non-response bias asks "are my respondents representative of the class?" Differential attrition asks "did the composition of my sample change between the two measurements I am comparing?" — and it is that change, correlated with condition, that fabricates or masks effects.

Limitations and honest caveats

A rigorous reader will push back, and should:

  1. Attrition does not always bias the estimate you care about. As Goodman and Blum showed, drop-out can shift means while leaving relationships intact. If your question is about a correlation rather than a level, attrition may be survivable. The point is to test, following their four checks, not to panic or to wave it away.
  2. You can only test on observed variables. Comparing leavers and stayers on wave-1 characteristics detects attrition related to measured traits. If drop-out depends on something unmeasured — a private disappointment, an unrecorded grade shock — no balance test will reveal it. This is the same missing-not-at-random problem discussed in MAR vs MNAR, and it caps how reassuring any diagnostic can be.
  3. The WWC thresholds are calibrated for RCTs, not surveys. The 5.7 / 10 percentage-point boundaries come from a randomised-trial framework with specific bias models. They are a useful sanity check for evaluation panels, but do not treat them as an exact pass/fail law for an observational course-evaluation study.
  4. Fixes are partial. Inverse-probability weighting, multiple imputation, and selection models can mitigate attrition bias, but each rests on assumptions about why people left that you usually cannot fully verify. The honest posture is that the best defence against attrition is preventing it, and the second best is transparent reporting of how much occurred and who left.

How Koji incorporates this

Koji treats attrition as a first-class fact about a dataset, not an embarrassment to hide:

  • Wave- and condition-level response accounting. Because Koji captures each response with structured metadata and timestamps, an analyst can compute attrition rates by wave and by condition directly — the exact quantities the What Works Clearinghouse cares about — rather than reconstructing them from disconnected survey exports. That makes it possible to flag a dangerous difference in drop-out between two sections before anyone compares their scores.
  • Leaver-versus-stayer diagnostics. With wave-1 characteristics stored alongside later responses, running the Goodman and Blum checks — is drop-out predictable, do means differ, have variances or relationships shifted — becomes a query rather than a bespoke project. Koji's role is to keep the data in a shape where those tests are cheap to run.
  • Reducing attrition at the source. The most reliable mitigation is fewer drop-outs, and Koji is designed to lower the burden that drives them: adaptive AI-moderated interviews that stay short and relevant rather than marching every student through a fixed battery, which directly attacks the survey fatigue that thins out later waves. Structured question types (scale, single_choice, open_ended) collect the same construct with less respondent effort.
  • Reaching the students who would otherwise vanish. Conversational, mobile-first collection is aimed at the disengaged tail that abandons long end-of-term forms — precisely the group whose selective loss biases summative scores upward. Koji is careful not to overclaim here: it cannot eliminate attrition, and no platform can recover a student who has genuinely left the course. What it can do is shrink avoidable drop-out and make the remaining attrition measurable and reportable.

Framed honestly, Koji does not "solve" panel mortality — nothing does. It lowers avoidable attrition and, just as importantly, refuses to let the attrition go unmeasured, so that a comparison rests on evidence about who stayed rather than on an unstated hope that everyone did. Koji's core research platform at koji.so applies the same adaptive-interview engine to longitudinal product and customer studies, where selective drop-out between waves is an identical and equally underestimated threat.

References

Frequently asked questions

What is the difference between attrition and non-response? Non-response is failing to reach part of a single sample at one point in time. Attrition (panel mortality) is losing the same individuals across repeated waves. Differential attrition is when that loss happens at different rates in different conditions or groups — which is what can fabricate or hide an effect in a before/after comparison.

Why is differential attrition worse than overall attrition? Overall attrition mainly shrinks your sample and widens confidence intervals. Differential attrition changes who remains in each group unevenly, so a difference in outcomes between groups partly reflects composition rather than treatment. Randomisation or matched design protects you only if drop-out is similar across conditions.

How much differential attrition is too much? There is no single number, but the What Works Clearinghouse treats even a few percentage points as potentially disqualifying: near-zero overall attrition tolerates roughly 5.7 points of differential attrition under a cautious threshold and about 10 points under an optimistic one. Treat any sizeable gap in drop-out rates between groups as a reason to investigate before trusting the comparison.

Can I just weight or impute my way out of attrition? Partly. Inverse-probability weighting, multiple imputation, and selection models can mitigate the bias, but each assumes something about why people left that you usually cannot fully verify — especially if drop-out depends on unmeasured factors (missing not at random). These are damage control, not a cure; preventing attrition and reporting it transparently matter more.

Does Koji fix attrition? No platform can recover a student who has left. Koji reduces avoidable drop-out with short, adaptive, mobile-first interviews that fight survey fatigue, and it records attrition by wave and condition so you can run leaver-versus-stayer diagnostics and report the loss honestly. The mitigation is real but partial, and Koji frames it that way.

Related resources