New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
analysis-reporting9 min read

Where Exactly Do Students Abandon Your Evaluation? Discrete-Time Survival Analysis of Breakoff

A completion rate tells you how many students quit; it cannot tell you where or why. Discrete-time survival analysis models the hazard of breakoff question by question, turning a single number into an actionable map of where your evaluation loses people.

Koji Education Team

Product

In brief

Every course evaluation has a completion rate, but that single number hides the information you actually need: at which question do students give up, and what about that question makes them quit? Discrete-time survival analysis answers this by treating the questionnaire as a sequence of moments and estimating the hazard — the probability of breaking off at each question, given the respondent reached it. Peytchev (2009) established breakoff as a distinct response behaviour with its own predictors, and Mittereder and West (2022) showed a dynamic survival model can predict breakoff well enough to intervene in real time. The method converts a flat completion statistic into a per-question map of where your instrument bleeds respondents.

This matters because breakoff is not random. It clusters at grids, at sensitive items, and at points of fatigue — so where students quit is a design diagnosis. It complements the descriptive treatment of questionnaire length and breakoff by giving the modelling method behind the phenomenon, and it connects to nonresponse and selection bias because who breaks off is rarely a representative slice of who started.

What the research says

Breakoff is a distinct behaviour. Peytchev (2009), Survey Breakoff (Public Opinion Quarterly, 73(1), 74-97), drew the key conceptual line: unit nonresponse is refusing to start, while breakoff is abandoning after starting. Because breakoff happens only after respondents have seen questions, it is predicted by features inside the survey — question type, burden, sensitivity, layout — that unit nonresponse cannot be. Peytchev found breakoff rates in web surveys far exceed those in interviewer-administered ones, and that design characteristics encountered mid-survey drive them. Practically, this reframes breakoff from a respondent failing to a design signal.

The discrete-time survival framing. The natural model treats "time" as the cumulative count of questions the respondent has seen, which is discrete. The hazard is the probability that a person breaks off at a given question, conditional on not having broken off earlier. Singer and Willett (2003), Applied Longitudinal Data Analysis, provide the canonical treatment of discrete-time survival (event-history) analysis: you restructure the data into person-period format — one row per respondent per question reached — and fit a logistic (or complementary log-log) regression predicting the breakoff event, with the question position entered flexibly. This yields a hazard function across the questionnaire: a curve showing the risk of losing a respondent at each step.

Prediction good enough to act on. Mittereder and West (2022), A Dynamic Survival Modeling Approach to the Prediction of Web Survey Breakoff (Journal of Survey Statistics and Methodology, 10(4), 945-978), extended this to a dynamic model that updates a respondent's breakoff risk as they progress, using paradata (response times, backing up, item-level behaviour). The model predicts breakoff accurately enough that a live survey could trigger a tailored intervention — a progress reassurance, a shortened path — for respondents at high risk before they quit. Related work by Chen, Cernat, Shlomo and Eckman (2022) shows question topic and filter-question format shift breakoff hazards, confirming that specific, fixable design choices move the curve.

Together these studies establish three facts a quality office can use: breakoff is patterned and predictable; the pattern lives at the question level; and the driver is design, which means it is changeable.

Why it matters for course evaluation in practice

Course evaluations are a near-ideal case for this method, and a poorly served one in current practice:

  1. A completion rate is a lagging, uninformative summary. Knowing 38% abandoned the survey tells you that you have a problem, not where it is. A hazard function tells you the breakoff spikes at the open-text question in position 14, or at the demographic grid — turning "improve completion" into a specific edit.

  2. Breakoff and nonresponse bias are linked. If the students who break off differ systematically from those who finish — for example, dissatisfied students quitting at the "would you recommend" item — then the completed responses are a biased sample of opinion. This is the same threat analysed under selection bias and early-vs-late wave nonresponse. Locating where breakoff happens is the first step to reasoning about whose voice is being lost.

  3. It targets design fixes precisely. Because the hazard is estimated per question with covariates, you can test whether long grids, mandatory items, sensitive climate questions, or mobile display raise the hazard, and by how much. That is a defensible, evidence-based basis for redesign rather than guesswork — the survey-methods analogue of using control charts to separate signal from noise.

  4. It rewards moving sensitive and burdensome items thoughtfully. If the hazard model shows a spike at a sensitive item, you learn that placing it late costs you the responses of everyone who quits there. The model quantifies the trade-off between asking hard questions and retaining respondents.

Limitations and honest caveats

  • Breakoff position is confounded with content and fatigue. A spike at question 14 could be the question's topic, its format, or simple exhaustion by that point. Position and content are entangled; disentangling them requires experimental variation (randomising item order), not the observational hazard model alone. Without randomisation, a hazard spike is a flag, not a cause.

  • Paradata quality drives the dynamic model. Mittereder and West's real-time prediction depends on clean response-time and navigation paradata. If your platform does not capture per-question timing reliably, the dynamic model degrades to a static one, which still maps the hazard but cannot intervene live.

  • Rare events and small classes. In a class of 25, breakoff events are few and the per-question hazard is estimated with wide uncertainty. Pooling across many courses is usually necessary to estimate a stable hazard function, and pooling assumes the breakoff process is similar across courses — an assumption worth checking, not presuming.

  • Interventions can backfire. A progress bar or reassurance prompt aimed at high-risk respondents can itself annoy, and aggressive completion pressure risks converting a would-be breakoff into a careless straightliner — trading missing data for low-quality data. The goal is genuine burden reduction, not coercion to finish.

  • Modelling breakoff does not fix the underlying instrument. The method diagnoses; the cure is design. A beautiful hazard analysis that leads to no shortening or reordering of the questionnaire has changed nothing for students.

How Koji incorporates this

Koji for Education is architected around the insight that abandonment is a design signal, not a respondent failure:

  • Question-level paradata by default. Because Koji collects evaluations as a structured, moderated conversation, it naturally records where in the flow a respondent stops and how long each step took. That per-question, person-period record is precisely the input a discrete-time survival model needs — no reconstruction from server logs required.

  • Adaptive, burden-aware flow. The AI-moderated interview does not force every respondent down one long fixed path. It can probe where a student has something to say and move on where they do not, which is designed to lower the hazard at exactly the fatigue and grid points where fixed forms lose people. This operationalises the "reduce burden, not coerce completion" lesson from the breakoff literature.

  • Bias-aware reporting of who was lost. Koji is designed to report not just a completion rate but a view of where drop-off concentrates and whether those who stopped differ from those who finished, so a quality committee can judge whether the completed responses are a representative sample — the selection-bias question — rather than assuming they are.

  • Sensitive items handled with care. Because the breakoff hazard spikes at sensitive questions, Koji's conversational framing and anonymity design aim to reduce the cost of asking climate or wellbeing items, mitigating (not eliminating) the drop-off they otherwise trigger. Koji's core research platform at koji.so applies the same moderated-flow engine to product and customer surveys, where breakoff at long grids is an equally expensive, equally diagnosable problem.

Related resources

References

Frequently asked questions

What is the difference between breakoff and nonresponse?

Unit nonresponse means a student never starts the evaluation. Breakoff means they start, answer some questions, then abandon it. The distinction matters because breakoff is driven by things inside the survey that the respondent has already seen — question burden, sensitivity, layout — whereas nonresponse is driven by factors before the first question. Peytchev (2009) established breakoff as a separate behaviour with its own predictors.

What does the hazard function actually show?

It shows, for each question in your evaluation, the probability that a respondent who reached that question breaks off there. Plotted across the questionnaire, it is a curve that reveals the exact positions where you lose people — a grid, a sensitive item, a long open-text box — rather than a single completion rate that hides all of that.

How is the data structured for a discrete-time survival model?

You reshape it into person-period format: one row per respondent per question they reached, with an indicator of whether they broke off at that question. Then you fit a logistic or complementary log-log regression predicting breakoff, with question position entered flexibly and any design covariates added. Singer and Willett (2003) give the standard recipe.

Can this predict breakoff in real time?

Yes, within limits. Mittereder and West (2022) built a dynamic survival model that updates each respondent's breakoff risk as they progress, using response times and navigation paradata, accurately enough to trigger an intervention before they quit. Real-time prediction depends on capturing clean per-question paradata; without it, you can still map the hazard after the fact but cannot intervene live.

Does modelling breakoff work for a small class?

Not on its own. In a class of 25 there are too few breakoff events to estimate a stable per-question hazard, so uncertainty is wide. You usually need to pool across many courses, which assumes the breakoff process is similar across them — an assumption to check rather than presume.

Could reducing breakoff hurt data quality?

It can, if you pursue completion by coercion. Aggressive progress pressure can convert a would-be breakoff into a careless straightliner, trading missing data for low-quality data. The productive fix is genuine burden reduction — shorter, better-ordered, less punishing questions — not forcing reluctant respondents to the end.

Related articles

analysis-reporting

Students Are Nested in Courses: Why Multilevel Models Beat Raw Averages for Course Evaluation

Course-evaluation data has a nested structure — students within sections within instructors — and a flat average ignores it. This guide explains how multilevel (hierarchical linear) models partition variance, why ignoring clustering understates uncertainty, and what it means for fair reporting.

analysis-reporting

Did the Teaching Change Cause the Score to Move? Difference-in-Differences for Course Evaluation

Difference-in-differences lets you estimate whether a course redesign actually moved evaluation scores by comparing a treated course against a similar untouched one over time. Here is how the design works, when its parallel-trends assumption holds, and how to use it honestly in quality assurance.

research-methods

How Long Should a Course Evaluation Be? Questionnaire Length, Breakoff, and Answer Quality

What the survey-methodology evidence says about questionnaire length: longer instruments depress participation and degrade answers to later questions, but ruthless shortening is not automatically the answer. A research-grounded guide for designing course evaluations.

research-methods

Latent Transition Analysis: Modelling How Student Segments Move Between Evaluation Waves

Latent transition analysis (LTA) extends latent profile analysis over time, estimating how students move between response segments from a mid-semester to an end-of-semester evaluation. This guide explains the method, the evidence, its limitations, and how Koji uses multi-wave data to support it.