Is Your Missing Course-Evaluation Data Random? MAR, MNAR, and What to Do About It
A low response rate is a missing-data problem. Rubin's MCAR/MAR/MNAR framework explains when a course-evaluation mean is biased and when multiple imputation or maximum likelihood can help.
Koji Education Team
Product
In brief
A low course-evaluation response rate is, statistically, a missing-data problem — and what you can do about it depends entirely on why the data are missing. Donald Rubin's framework (1976) distinguishes three mechanisms: missing completely at random (MCAR), missing at random (MAR), and missing not at random (MNAR). Under MCAR and MAR, principled methods such as multiple imputation (MI) and full-information maximum likelihood (FIML) can recover unbiased estimates; under MNAR — where the students who skip the survey differ on the very thing you are measuring even after accounting for what you know about them — no statistical fix is guaranteed, and only better collection or explicit sensitivity analysis helps. For course evaluation, the uncomfortable truth is that non-response is often plausibly MNAR, which is why response-rate improvement and transparent caveats matter more than any imputation trick.
What the research says
The conceptual foundation is Donald Rubin's 1976 Biometrika paper, "Inference and missing data," which introduced the now-standard taxonomy of missingness mechanisms. The definitions are precise and worth stating carefully:
- MCAR (missing completely at random): the probability that a value is missing is unrelated to any data, observed or unobserved. A student fails to respond for reasons entirely independent of their opinion of the course or anything you measured. Respondents are then a random subsample, and the simple mean of respondents is unbiased.
- MAR (missing at random): the probability of missingness depends only on observed variables. For example, response depends on year of study, discipline, or prior GPA — all of which you have on file — but, conditional on those, not on the unobserved rating itself. Under MAR, methods that use the observed variables can correct the bias.
- MNAR (missing not at random): missingness depends on the unobserved value itself even after conditioning on observed data. The classic course-evaluation case: dissatisfied students disproportionately do not bother to respond (or, in some settings, only the dissatisfied bother). Here the respondent mean is biased and no method recovers the truth without additional assumptions.
The authoritative modern treatment is Schafer and Graham (2002), "Missing data: our view of the state of the art," in Psychological Methods — the most-cited practical synthesis. They clear up a persistent misunderstanding (MAR does not mean the missingness is "random" in the colloquial sense) and establish multiple imputation and maximum likelihood as the state-of-the-art, both vastly preferable to the older defaults of listwise deletion, mean substitution, or single imputation, which bias estimates and understate uncertainty. John Graham's 2009 Annual Review of Psychology paper, "Missing data analysis: making it work in the real world," extends this with the single most useful practical lever: auxiliary variables. Including variables that predict either the outcome or the missingness (e.g., GPA, attendance, prior-course ratings) in the imputation/estimation model both reduces bias and recovers statistical power — and crucially, can make an otherwise-MNAR situation behave more like MAR.
Multiple imputation works in three steps: (1) generate several (e.g., 20–100) complete datasets, each filling missing values with draws that reflect the predictive uncertainty; (2) analyse each completed dataset normally; (3) pool the results using Rubin's rules, which combine the within- and between-imputation variance so that standard errors honestly reflect the fact that the missing values were estimated, not known. This last point is what single imputation gets wrong: it pretends the filled-in values are real and reports falsely precise results.
Why it matters for course evaluation in practice
Course evaluations rarely achieve high response rates, especially online; Nulty (2008), a widely cited higher-education reference, documents the gap between paper and online response rates and the response levels needed for defensible inference. But the response rate is the wrong thing to fixate on. The decisive question is the mechanism:
- A course with 35% response that is MCAR can yield a trustworthy mean. A course with 70% response that is MNAR (the missing 30% are systematically the most dissatisfied) can yield a biased mean. Rate alone does not tell you which you have.
- The most defensible institutional posture treats non-response as plausibly MAR at best and tests sensitivity to MNAR. Practically that means: (1) use the administrative variables you already hold (year, programme, GPA band, attendance) as auxiliary variables so that comparisons and any model-based adjustment lean on observed structure; (2) check whether early and late respondents differ (the wave-analysis logic is a poor-man's MNAR probe — late responders are a proxy for non-responders); and (3) report a range under different assumptions rather than a single point estimate when the stakes are high.
- Weighting and imputation are complements, not competitors. Post-stratification weighting (covered in Can You Weight Your Way Out of a Low Response Rate?) corrects for who responded on observed margins; multiple imputation and FIML additionally borrow strength from auxiliary variables and propagate uncertainty into standard errors. Both rest on the same MAR assumption and both fail silently under MNAR.
The operational upshot for a quality-assurance office: stop reporting a bare mean from a 25% sample as if it were the class's verdict. Either raise response toward the level where the mechanism matters less, or accompany the estimate with an honest statement of the assumed mechanism and a sensitivity check.
Limitations and honest caveats
A PhD reader will immediately raise the central one: the missingness mechanism is fundamentally untestable from the observed data alone. You cannot prove MAR versus MNAR, because doing so would require the very values that are missing. Everything downstream is an assumption you should state, not a fact you can verify. Second, multiple imputation and FIML are only as good as the model: omit a strong auxiliary variable and you bias the result; impose the wrong distributional form and you distort it. Third, imputation does not create information — it makes honest use of what you have and propagates uncertainty correctly; with a 15% response rate the confidence intervals will (rightly) be wide, and no method narrows them legitimately. Fourth, small classes make all of this fragile: with 8 respondents and 20 non-respondents, imputation models are unstable and the better move is qualitative caution (see small mean differences and confidence intervals). Fifth, sensitivity analysis for MNAR (e.g., pattern-mixture models, selection models, delta-adjustment) requires explicit, contestable assumptions about how much non-respondents differ — useful for bracketing, not for producing a definitive number. The honest framing is that statistical machinery disciplines and bounds the uncertainty introduced by non-response; it does not abolish it.
How Koji incorporates this
Koji's primary contribution to the missing-data problem is upstream: the best cure for non-response bias is fewer non-respondents, and the most defensible analysis is one where the mechanism is closer to MAR.
- Higher, more representative response through better experience. Koji's AI-moderated conversational interview is faster and more engaging than a long Likert grid, which lifts completion — and, more importantly, lifts it across the spectrum of opinion rather than only among the highly motivated, pushing the realised mechanism away from the MNAR worst case. This is the same logic behind the response-rate evidence in What Actually Raises Course-Evaluation Response Rates?.
- Auxiliary-variable-aware reporting. Because Koji collects structured metadata alongside each interview (cohort, programme, mode of study), its reporting can stratify and contextualise rather than present a single undifferentiated mean — exactly the observed-variable structure that MAR-based correction depends on. Koji surfaces respondent composition next to the result so a committee can judge representativeness instead of assuming it.
- Early-vs-late and composition diagnostics. Koji can compare early and late responders and flag when the responding sample diverges from the enrolled population on known margins — an operational, non-response sensitivity check in the spirit of Graham's auxiliary-variable advice.
- Triangulation instead of a single contaminated source. Where one survey is likely MNAR, Koji supports combining course-level evaluation with mid-cycle formative collection and other evidence (see Triangulation) so conclusions do not rest on a single, possibly-biased instrument.
Koji is careful not to overclaim: it does not "fix" MNAR non-response, and no platform can. It is designed to reduce the missingness, expose the composition of who responded, and support the honest, assumption-explicit reporting the literature demands. The same engine underpins Koji's core research platform at koji.so, where non-response bias is an equally central concern for customer and product research.
Related resources
- Can You Weight Your Way Out of a Low Response Rate? Post-Stratification and Nonresponse Adjustment
- Are Your Respondents Representative? Early-vs-Late Wave Analysis
- Online vs. Paper Course Evaluations: Do Lower Response Rates Mean Worse Data?
- How Many Responses Do You Need for a Reliable Course Evaluation?
- What Actually Raises Course-Evaluation Response Rates? The Experimental Evidence
- Selection Bias in Course Evaluations: What Goos and Salomons Found
References
- Rubin, D. B. (1976). Inference and missing data. Biometrika, 63(3), 581–592. https://doi.org/10.1093/biomet/63.3.581
- Schafer, J. L., & Graham, J. W. (2002). Missing data: our view of the state of the art. Psychological Methods, 7(2), 147–177. https://doi.org/10.1037/1082-989X.7.2.147
- Graham, J. W. (2009). Missing data analysis: making it work in the real world. Annual Review of Psychology, 60, 549–576. https://doi.org/10.1146/annurev.psych.58.110405.085530
- Nulty, D. D. (2008). The adequacy of response rates to online and paper surveys: what can be done? Assessment & Evaluation in Higher Education, 33(3), 301–314. https://doi.org/10.1080/02602930701293231
Related articles
How Many Responses Do You Need for a Reliable Course Evaluation?
Reliability of a class-average rating rises steeply with the number of students who respond. Generalizability theory — from Marsh''s syntheses to Dzakadzie (2026) — shows raters matter more than items, and a class of 12 cannot be evaluated like a class of 120.
Are Your Respondents Representative? Using Early-vs-Late Wave Analysis to Estimate Nonresponse Bias
A low response rate is not automatically biased — what matters is whether respondents differ from non-respondents. Armstrong and Overton (1977) gave us a cheap diagnostic: compare early and late responders. Here is how to use wave analysis on course-evaluation data and where it breaks down.
What Actually Raises Course-Evaluation Response Rates? The Experimental Evidence
Online course evaluations chronically under-perform paper. We review the experimental evidence — Dommeyer''s grade-incentive trials and Nulty''s adequacy thresholds — on what genuinely lifts response rates, what it costs in data quality, and how to hit a defensible rate without coercion.
Selection Bias in Course Evaluations: What Goos and Salomons Found
A research-grounded reading of Goos & Salomons (2017) on selection bias in online course evaluations, with practical implications for response-rate policy and reporting at European universities.