New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods9 min read

Is a Low Response Rate Biasing Your Course Evaluations? MCAR, MAR, and Missing-Not-at-Random

A low response rate does not automatically bias a course-evaluation score — the missing-data mechanism (MCAR, MAR, or MNAR) does. How to diagnose nonresponse bias instead of chasing a rate threshold.

Koji Education Team

Product

Answer first. A low response rate does not automatically bias a course-evaluation score. Bias depends on the mechanism behind the missing responses. If the students who skip the survey differ from those who complete it on the very thing you are measuring — their satisfaction with teaching — the missingness is "not at random" (MNAR) and the mean is biased, sometimes badly. If they differ only on characteristics you have recorded (grade, discipline, cohort), the missingness is "at random" (MAR) and is correctable. The response rate alone tells you almost nothing; the response pattern tells you everything.

What the research says

The intuition that "low response rate = untrustworthy result" is so widespread that it has become institutional folklore. The survey-methodology literature has spent two decades dismantling it.

The foundational distinction comes from Rubin's (1976) taxonomy of missing-data mechanisms, later formalised by Little and Rubin in Statistical Analysis with Missing Data. Missingness is:

  • MCAR (missing completely at random): the probability of skipping the evaluation is unrelated to anything — a student was ill on survey day for reasons having nothing to do with the course. The respondents are a genuine random subsample, and the mean is unbiased, just less precise.
  • MAR (missing at random): the probability of skipping depends on observed variables. Final-year students respond less than first-years, say — but once you condition on year of study, respondents and non-respondents have the same underlying opinion. This is correctable with weighting or imputation.
  • MNAR (missing not at random): the probability of skipping depends on the unobserved outcome itself. The students who are indifferent, disengaged, or quietly furious are exactly the ones who never open the survey. No amount of adjustment on observed covariates fully fixes this, because the thing driving both the opinion and the non-response is unmeasured.

The crucial empirical result is that response rate and nonresponse bias are only weakly linked. In his influential Public Opinion Quarterly paper, Groves (2006) assembled cases showing that surveys with low response rates can have negligible bias and surveys with high response rates can carry substantial bias. The Groves and Peytcheva (2008) meta-analysis of 59 methodological studies made the point quantitatively: across 959 estimates, the correlation between an item's nonresponse rate and its nonresponse bias was weak and highly variable. Bias arises when the cause of non-response is correlated with the survey variable — not when the rate is merely high.

For course evaluations specifically, Adams and Umbach (2012) provide the most cited direct evidence. Studying roughly 22,000 undergraduates who were sent about 135,000 online evaluations at a single research university, they fitted multilevel models predicting who responds. Response was systematically related to salience (students responded more to courses in their major), survey fatigue (each additional simultaneous evaluation request lowered the odds of completing any of them), and academic environment. Non-response, in other words, was demonstrably not random — it was patterned, and several of the patterns (engagement, major-relevance) plausibly correlate with the evaluation ratings themselves. That is the signature of an MNAR risk, not a reassuring one.

Why it matters for course evaluation in practice

Three practical consequences follow.

1. Chasing response-rate thresholds is the wrong target. Many quality-assurance offices impose a minimum response rate (often 50%) below which results are "not reported." This is defensible as a precision safeguard, but as a bias safeguard it is close to superstition. A 70% response rate composed disproportionately of the most enthusiastic students is more biased than a 35% rate that happens to be demographically balanced. The mechanism, not the rate, determines validity.

2. The direction of MNAR bias in evaluations is not obvious. It is tempting to assume the disgruntled abstain, inflating the mean. But the opposite "grievance-driven response" pattern is also documented — students with a strong negative experience are motivated to voice it, while the broadly satisfied middle stays silent. Because the direction is context-dependent, you cannot simply "mentally discount" a high score. You need evidence about who is missing.

3. Small classes amplify everything. In a seminar of 12 with 5 respondents, a single MNAR-driven omission moves the mean visibly. Nonresponse bias and small-sample instability compound; this is where over-interpretation of a single number does the most damage.

The productive response is to diagnose the mechanism rather than assume it. Compare respondents and non-respondents on everything you do observe — grade distribution, attendance, major status, year, demographics. If they look alike on observables, MAR is plausible and the score can be trusted (or weighted). If they differ sharply on observables, treat the divergence on the unobserved opinion as a live possibility and report the score with an explicit caveat. Sensitivity analysis — asking "how different would the non-respondents have to be to overturn this conclusion?" — is more honest than a pass/fail response-rate gate.

Limitations and honest caveats

A PhD reader will raise several objections, and they are correct to.

  • MNAR is, by definition, untestable from the data alone. You cannot prove missingness is MNAR versus MAR using only the observed responses, because the distinguishing information is the very thing you failed to collect. Analysts can only reason about plausibility, use auxiliary data, and run sensitivity analyses. Anyone claiming to have "corrected for MNAR" without external information or explicit modelling assumptions is overclaiming.
  • The Adams and Umbach (2012) evidence is single-institution and now over a decade old. Response behaviour is shaped by local policy, reminder culture, and platform design; generalisation to a European quality-assurance context should be cautious. Their models predict response, not bias directly — the inference from "non-response is patterned" to "the mean is biased" requires the additional (untestable) step that the patterns correlate with opinion.
  • Weighting and selection models trade one assumption for another. Post-stratification weighting fixes MAR but can increase variance and does nothing for genuine MNAR. Heckman-type selection models can address MNAR but rely on strong distributional and exclusion-restriction assumptions that are hard to satisfy for evaluations. There is no assumption-free fix.
  • Response-rate floors still have a legitimate role — for precision and for protecting individual instructors from volatile small-sample scores. The argument here is against treating the rate as a bias certificate, not against the rate as a reporting safeguard.

How Koji incorporates this

Koji is designed to attack the nonresponse problem at its two sources — reducing patterned nonresponse and diagnosing what remains — rather than papering over it with a rate threshold.

  • Fatigue-aware collection. Because Adams and Umbach identify survey fatigue as a primary driver of non-response, Koji favours short, AI-moderated conversational evaluations over long grids, and supports staggered mid-cycle collection so a student is not hit with a wall of simultaneous requests at term end. The conversational format is engineered to convert reluctant, low-salience respondents who would abandon a 40-item matrix.
  • Respondent-vs-non-respondent diagnostics. Koji's reporting can compare the observable profile of students who completed an evaluation against the enrolled cohort (grade band, year, demographic composition where available), surfacing the MAR-vs-MNAR question directly instead of hiding it behind a single response-rate figure.
  • Bias-aware reporting language. Rather than a binary "reported / suppressed," Koji is built to attach uncertainty and representativeness context to a score, so a programme director reads "4.1, but respondents skew toward higher-attending students" rather than a bare number.
  • Probing the silent middle. The AI-moderated interview is designed to elicit substance from students who would otherwise leave a Likert item blank or give a non-committal midpoint — directly targeting the disengaged group whose absence drives MNAR risk. Koji frames this as designed to mitigate nonresponse bias, not eliminate it: no collection method can recover the opinion of a student who never participates.

For teams running non-teaching research, Koji's core platform at koji.so applies the same fatigue-aware, AI-moderated interview engine to customer and product studies, where self-selection into feedback is an identical threat to validity.

Related resources

Frequently asked questions

Does a low response rate automatically mean my course evaluation is biased?

No. Bias depends on the mechanism of non-response, not the rate. If non-respondents differ from respondents only on characteristics you have recorded (missing at random), the score is correctable and often trustworthy. Bias arises only when the reason for skipping is correlated with the opinion being measured (missing not at random). Groves (2006) and the Groves & Peytcheva (2008) meta-analysis show the rate–bias link is weak.

What is the difference between MAR and MNAR for course evaluations?

Missing at random (MAR) means the chance of skipping depends on observed variables — year of study, discipline, grade — so conditioning on those variables removes the bias. Missing not at random (MNAR) means the chance of skipping depends on the unobserved rating itself, for example when disengaged students systematically ignore the survey. MNAR cannot be fully corrected from the collected data alone.

Can I test whether my missingness is MNAR?

Not directly. MNAR is untestable from the observed data because the distinguishing information was never collected. You can only assess plausibility by comparing respondents and non-respondents on observed variables, using auxiliary institutional data, and running sensitivity analyses that ask how extreme the non-respondents would have to be to change your conclusion.

Should we still enforce a minimum response-rate threshold?

A threshold is reasonable as a precision and individual-fairness safeguard — small, volatile samples produce unstable scores. But it is a poor bias safeguard, because a high rate can still be unrepresentative and a low rate can be balanced. Treat the threshold as protecting against noise, not as a certificate that the score is unbiased.

Which non-response finding is most relevant to online evaluations specifically?

Adams and Umbach (2012), studying ~135,000 online evaluations, found that response is systematically driven by course salience (students respond more in their major) and survey fatigue (each extra simultaneous request lowers completion). Because those drivers plausibly correlate with satisfaction, the study is direct evidence that online-evaluation non-response is patterned rather than random.

References

  • Rubin, D. B. (1976). Inference and missing data. Biometrika, 63(3), 581–592. https://doi.org/10.1093/biomet/63.3.581
  • Little, R. J. A., & Rubin, D. B. (2019). Statistical Analysis with Missing Data (3rd ed.). Wiley.
  • Groves, R. M. (2006). Nonresponse rates and nonresponse bias in household surveys. Public Opinion Quarterly, 70(5), 646–675. https://doi.org/10.1093/poq/nfl033
  • Groves, R. M., & Peytcheva, E. (2008). The impact of nonresponse rates on nonresponse bias: A meta-analysis. Public Opinion Quarterly, 72(2), 167–189. https://doi.org/10.1093/poq/nfn011
  • Adams, M. J. D., & Umbach, P. D. (2012). Nonresponse and online student evaluations of teaching: Understanding the influence of salience, fatigue, and academic environments. Research in Higher Education, 53(5), 576–591. https://doi.org/10.1007/s11162-011-9240-5