New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Evaluation bias8 min read

The Students Who Withdrew Never Got to Evaluate: Survivorship Bias in Course Evaluation

Non-response bias is the familiar worry. The deeper problem is survivorship: the students who struggled most—who dropped, withdrew, or quietly disengaged—are structurally absent from your end-of-term ratings. Here is what that does to your data, and how to hear the missing voices.

Koji Education Team

Product ·

Bottom line: Most course-evaluation debates focus on whether the students who answer differ from those who stay silent. But there is a more fundamental distortion: the students who withdrew from the course, switched programmes, or stopped attending are never even in the room when the survey goes out. Your "average rating" is computed over the survivors. For courses with high attrition—and for the struggling students any quality system most needs to hear—this survivorship bias systematically flatters the data. The fix is not a better Likert scale; it is collecting feedback earlier, from everyone, in a format that reaches the disengaged before they vanish.

Two different missing-data problems

Quality officers are rightly trained to watch response rates. If only 35% of enrolled students complete an evaluation, the worry is non-response bias: the 35% who answered may not represent the 65% who did not. That is real, and the evidence is sobering. Online evaluation response rates that begin near 60% commonly fall into the 30–40% range, and research on non-response bias in student evaluations finds that respondents and non-respondents differ systematically—enough that correcting for it can change how instructors rank against one another.

But non-response bias assumes a fixed population—the enrolled cohort—from which some choose not to respond. Survivorship bias is different and more insidious: the population itself has already been filtered before the survey exists. Students who withdrew, deferred, or transferred out are not "non-responders." They are gone. They never receive the end-of-term invitation, or if they do, the system excludes their record. As one analysis of the practice put it bluntly, student surveys of professors do not account for withdrawals—and withdrawals, while a small share of all grades in a typical semester, are concentrated in exactly the courses and the students a quality process should scrutinise most.

Why the survivors give you a rosier picture

This is the statistical heart of the matter. Attrition is not random. Students leave courses disproportionately when they are struggling, when the workload feels unmanageable, when the teaching did not work for them, or when the subject was mismatched to their preparation. These are not neutral departures—they are, in many cases, the strongest negative signals a course could generate. By the time the evaluation runs, those signals have self-selected out of the dataset.

The result is a dataset weighted toward students who coped, kept up, and stayed. The 4.2/5 you report is a 4.2 among those who made it to week twelve. The student who withdrew in week four because the pacing was punishing, or because the foundational concepts were never scaffolded, contributes nothing. Worse, the same forces that drive withdrawal—difficulty, poor alignment, feeling lost—are the forces a formative evaluation is supposed to catch. Survivorship bias guarantees that the louder a course's problems, the quieter they appear in the final numbers.

There is a compounding effect with response rates. Evidence on student evaluations shows that as response rates rise, average scores tend to fall for highly-rated teachers and rise for poorly-rated ones, while low response rates are associated with lower variance—a flatter, less discriminating distribution. In other words, thin, late, survivor-only data does not just shift the mean; it compresses the very differences administrators then over-interpret.

What this means for high-stakes decisions

If survivor-only ratings feed into promotion, renewal, or programme-review decisions, the institution is making consequential judgements on a biased sample and treating it as a census. A module with a 25% withdrawal rate and a glowing 4.4 from survivors is not obviously "better" than a module with full retention and a 3.9. The first may have shed its dissatisfied students; the second kept and challenged them. Averaging hides this completely. (For why even the surviving numbers are shakier than they look, see our pieces on small-class statistics and why averaging Likert scores misleads.)

"But doesn't this just mean we should chase higher response rates?"

This is the strongest objection, and it is half right. Higher response rates do reduce non-response bias, and mandatory or incentivised completion can help—though coercing response rates carries its own data-quality costs. But no amount of chasing responses at the end of term fixes survivorship. The withdrawn student cannot be surveyed in week twelve about a course they left in week four. A 95% response rate among survivors is still a 0% response rate among the departed.

A second objection: "Withdrawals are only a few percent of grades—why obsess over them?" Because the bias is not about volume; it is about direction. A small, systematically negative group removed from the denominator pulls every aggregate in the same flattering direction, and the effect concentrates precisely in the at-risk modules and at-risk student populations that equity-minded quality assurance is meant to protect. Rare and consequential is not the same as negligible.

The honest position is that no instrument fully recovers the voice of someone who has left. But you can dramatically shrink the window in which they disappear—by asking while they are still there.

How to estimate how much you are missing

You cannot interview the departed, but you can quantify the hole they leave. Three diagnostics make survivorship visible rather than invisible:

  1. Pair evaluation scores with withdrawal and completion rates per module. A high rating sitting on top of a high withdrawal rate is a warning, not a triumph. The two numbers belong on the same dashboard.
  2. Compare the demographics and prior attainment of responders against the full enrolled cohort. If your respondents skew toward higher-attaining, more-engaged students—as the non-response literature repeatedly finds—your sample is doubly filtered: survivors first, then volunteers.
  3. Run a short, separate exit touchpoint for students who withdraw or defer. Even a two-question conversational check-in at the point of withdrawal captures the reason while it is fresh—turning a silent departure into recoverable signal. This is the single most direct counter to survivorship bias available to a quality team.

None of these recovers a perfect picture. All of them convert an unknown, comfortable distortion into a measured, uncomfortable one—which is exactly the trade an evidence-led process should want.

How Koji narrows the survivorship gap

Koji does not claim to eliminate survivorship bias; no tool can interview a student who has already gone. What it changes is when and how you listen, so that fewer voices are lost before they are heard:

  • Formative, mid-cycle collection. Koji is built for feedback during the course, not only the post-mortem. A mid-semester conversational check-in reaches students before the withdrawal decision crystallises—capturing the struggling student in week five rather than mourning their absence in week twelve.
  • AI-moderated conversational interviews that probe beyond a rating. When a student signals difficulty, the interview follows up—"what specifically wasn't working, and when did you start to feel that?"—surfacing the early-warning detail a static Likert item cannot.
  • Automatic thematic analysis of open-text responses, so that even a handful of at-risk students produce structured, themed signal rather than a few anecdotes lost in an average. (A theme is not a sentiment score; it is actionable.)
  • Standardised, bias-aware AI moderation that asks every student the same well-formed questions, reducing the human-moderator inconsistency that distorts who feels comfortable speaking up.
  • Programme- and institution-level reporting that can pair feedback with retention and withdrawal data, so survivorship is made visible rather than silently baked into a single number.

The same conversational interview engine underpins koji.so for general user and customer research, where survivorship bias—hearing only from the customers who stayed—is just as corrosive. The cure is identical: ask earlier, ask everyone, and let the conversation find the people about to leave.

The takeaway for evidence-led evaluation

Survivorship bias is the quiet companion of non-response bias, and it resists the usual remedies. You cannot fix it at the end of term, and you cannot average your way out of it. Treat your end-of-course ratings as what they are—a measurement of the students who survived the course—and triangulate them with other evidence and, above all, with feedback collected while every student is still enrolled. The students who withdrew were trying to tell you something. The least an evaluation system can do is ask before they are gone.

Curious how mid-cycle conversational evaluation works in practice? See Koji for Education.