New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Methodology9 min read

Missing Not at Random: Why No Amount of Weighting Rescues a Biased Course Evaluation

Everyone worries about the response rate. The deeper problem is the reason students go silent — and if that reason is tied to the opinion you never captured, no weighting or imputation can put it back.

Koji Education Team

Product ·

Answer up front: A low response rate is not, by itself, the fatal flaw in a course evaluation. What matters is why students stay silent. Statistician Donald Rubin's foundational 1976 taxonomy divides missing data into three mechanisms — Missing Completely At Random (MCAR), Missing At Random (MAR), and Missing Not At Random (MNAR) — and the mechanism decides whether any repair is possible. Weighting and multiple imputation, the corrections most vendors quietly rely on, only recover an unbiased picture under MCAR or MAR. Course-evaluation non-response is almost always MNAR: students go quiet for reasons bound up with the very opinion you failed to record — they disengaged, they are heading for a fail, they withdrew before the survey opened. Under MNAR there is no statistical formula guaranteed to fix the bias. The lever is the design of collection, not the arithmetic you apply afterwards.

The comfortable story about response rates

The usual framing goes like this: response rates are falling, a low rate means a small and possibly unrepresentative sample, so we should either chase the rate up or statistically reweight the respondents we have to look like the class as a whole. Both instincts are reasonable, and we have written about the evidence on what actually raises response rates and about what post-stratification weighting can and cannot fix. But both instincts share a hidden assumption: that the students who answered and the students who did not differ only in ways you can see and therefore adjust for.

That assumption is exactly what Rubin's taxonomy forces you to interrogate.

Rubin's three mechanisms, in plain language

Rubin's 1976 paper Inference and Missing Data gave the field the vocabulary it still uses (Biometrika, 1976). Strip out the notation and it says this:

  • MCAR — Missing Completely At Random. The probability that a student skips the survey has nothing to do with anything — not their grade, not their satisfaction, not any characteristic. Missingness is pure coin-flip. Your respondents are a genuine random subsample. This is the friendliest case and the least believable one.
  • MAR — Missing At Random. Whether a student responds depends only on things you observed. Perhaps final-year students respond less than first-years, or large lectures less than seminars. As long as the missingness is fully explained by variables you hold, you can condition on them — through weighting or imputation — and recover an unbiased estimate. MAR is the assumption under which the standard fixes are valid.
  • MNAR — Missing Not At Random. Whether a student responds depends on the unobserved value itself — the very rating they did not give. The disengaged student who would have rated the course 2 is precisely the student least likely to bother. Here the missingness carries information you do not have, and no amount of reweighting on observed variables can conjure it back.

The methodological literature is blunt about the consequence: maximum likelihood and multiple imputation are valid under MCAR and MAR, but under MNAR "the missing data problem cannot be handled" by multiple imputation alone; you are reduced to selection or pattern-mixture models that require untestable assumptions about the missing values (van Buuren, Flexible Imputation of Missing Data).

Why course evaluation is structurally MNAR

Here is the uncomfortable part. Course evaluation is close to a textbook case of MNAR, because the mechanisms that drive silence are correlated with the opinion being measured:

  • Disengagement. A student who stopped attending in week 6 is both the least satisfied and the least likely to open an end-of-term email.
  • Attrition. The students who withdrew never got to evaluate — and their exit is often the single loudest verdict on the course. Their absence is not random noise; it is survivorship bias.
  • Grade expectation. Anticipated grade is one of the most consistently studied correlates of both rating and response. A widely cited finding is that students expecting higher grades are more likely to complete evaluations, which can tilt the surviving sample upward.

None of these are observed-only mechanisms you can cleanly weight away. They are tied to the latent satisfaction you are trying to estimate.

There is a subtle tell that the mechanism really is MNAR rather than MAR: the direction of non-response bias in the SET literature is not even consistent. A 2016 study of online student evaluations in Marketing Education Review found that raising the response rate was associated with lower average scores for high-rated teachers and higher average scores for low-rated teachers — non-response was compressing the spread, not simply inflating or deflating the mean (Wright & Meade / Marketing Education Review, 2016). When the bias flips sign depending on the instructor, you are not looking at a stable, correctable offset. You are looking at missingness entangled with the outcome — the signature of MNAR.

What this means for the fixes you have been sold

Rubin's taxonomy lets us be precise about each proposed remedy:

  1. "We reached the magic response-rate threshold." A high response rate reduces sampling variance and shrinks the room for non-response bias, but it does not eliminate it. Even an 80% rate can be badly biased if the missing 20% are systematically the disaffected. There is no threshold that converts MNAR into MAR.
  2. "We weight the respondents to match the class demographics." Valid only under MAR — that is, only if the demographics you weighted on fully explain the missingness. They rarely do, because the strongest driver (unobserved satisfaction) is not in your covariate set.
  3. "We impute the missing responses." Multiple imputation is a genuine advance, but it assumes MAR unless you build an explicit MNAR model — and an MNAR model requires you to assume how the non-responders would have answered, which is the very thing you cannot observe. You can run a sensitivity analysis; you cannot run a fix.

This is not a counsel of despair. It is a redirection.

Design beats arithmetic

If the disease is MNAR, the cure is upstream — collect feedback in a way that does not systematically lose the students most likely to be dissatisfied:

  • Collect while they are still there. Mid-cycle, formative collection reaches students before disengagement turns into silence and before the ones who will withdraw have gone. In-the-moment experience sampling does the same.
  • Follow up the non-responders directly, and treat their late responses as the most informative you will get — a targeted double-sample is one of the few honest ways to probe the missing mass.
  • Remove avoidable barriers. An inaccessible survey manufactures its own non-response, and that non-response is not random either.
  • Make responding low-cost and worth it, so the marginal, less-engaged student still participates rather than self-selecting out.

But doesn't every survey have this problem?

Yes — and that is the honest counterargument. Non-response is a universal survey affliction, not a course-evaluation peculiarity, and it would be unfair to hold SET to a standard no instrument meets. Two things follow. First, the point is not that course evaluations are worthless; it is that their non-response should be treated as plausibly outcome-related by default, so uncertainty is reported rather than hidden behind a tidy mean. The Total Survey Error framework puts non-response alongside measurement and coverage error precisely so no single number pretends to be the whole truth. Second, "everyone has it" is not "nothing can be done" — the design moves above genuinely shift the mechanism closer to MAR by shrinking the pool of informative missingness. The goal is not the impossible elimination of MNAR; it is reducing how much of your evidence is quietly self-selected, and being candid about the rest.

Where Koji fits

Koji for Education is built around the design response, not the arithmetic patch. Because it runs AI-moderated conversational interviews that can be deployed at any point in the term — not just as a retrospective end-of-course form — it supports formative, mid-cycle collection that captures students before disengagement hardens into non-response. The conversational format lowers the effort of a meaningful reply, which helps recruit the marginal respondent who would otherwise drop out of the sample. And rather than reporting a single average as if it were the settled truth, Koji surfaces where feedback is thin and where a result rests on a small, possibly self-selected group — so a programme director reads the number and its fragility. Koji does not claim to eliminate non-response bias; no tool can. It reduces the amount of informative missingness you build in, and it is honest about what remains. (The same AI interview engine underpins the general-purpose research platform at koji.so for teams doing wider user and customer research.)

Course evaluation's real problem was never that too few students answered. It is that the ones who stayed silent were trying to tell you something — and no weight you apply after the fact can hear it.

Closing the loop

If you want feedback that does not quietly discard your most disengaged students, see how Koji for Education collects evidence across the term. Design your way out of MNAR — because you cannot compute your way out.