Can You Weight Your Way Out of a Low Response Rate? Post-Stratification and Nonresponse Adjustment in Course Evaluations
A high response rate does not guarantee unbiased course evaluations, and a low one is not automatically wrong. Here is what survey methodology says about post-stratification weighting, when it corrects nonresponse bias, and when it just adds noise.
Koji Education Team
Product
Bottom line: Response rate is a weak predictor of nonresponse bias; what matters is whether the students who skip the evaluation differ on the rating itself. Post-stratification and nonresponse weighting can partly correct that bias — but only when you hold auxiliary data (such as final grade) that is related to both the decision to respond and the rating given. Weighting on variables that fail that test adds variance without removing bias. Weighting is a useful repair, not a substitute for representative collection.
The question universities keep getting wrong
Almost every quality office treats the response rate as the headline quality metric: 70% is "trustworthy," 30% is "unusable." Survey methodology says this framing is mistaken in both directions. The classic statement is Groves (2006), whose review of household-survey experiments found no consistent relationship between nonresponse rate and nonresponse bias — high-response surveys can be badly biased and low-response surveys can be nearly unbiased. The reason is structural: bias is not a function of how many people are missing but of who is missing and how their answers would have differed.
Formally, nonresponse bias in a mean is approximately the product of the nonresponse rate and the covariance between response propensity and the survey outcome. If the students who respond rate the course the same as those who do not, even a 20% response rate is unbiased for the mean. If response propensity is correlated with the rating — say, students who got poor grades or disengaged early are both less likely to respond and more critical — then even a 70% rate leaves a biased estimate. This is why Adams and Umbach (2012), studying nonresponse in online SET, conclude that obtaining a representative sample matters more than a high raw rate, and document that nonresponse is systematically driven by salience, survey fatigue, and the academic environment rather than being random.
What post-stratification weighting does
Post-stratification and nonresponse weighting attempt to repair a non-representative sample after the fact. The logic: if you have an auxiliary variable measured for both respondents and non-respondents, and respondents are skewed on that variable, you reweight respondents so their distribution matches the known population. For course evaluations the population is the enrolled class, and the registry often holds exactly the kind of auxiliaries you need — final grade, year of study, gender, programme, attendance, prior GPA. If high-grade students are over-represented among respondents, post-stratifying to the true grade distribution pulls the weighted mean back toward what the whole class would have said.
The pivotal caveat comes from Little and Vartivarian (2005). They show that weighting only reduces bias when the auxiliary variable is related to both the response propensity and the outcome of interest. An auxiliary correlated with response but not with the rating buys you nothing on bias and inflates the variance of the estimate — you have made the number noisier for no gain. An auxiliary correlated with the rating but not with response also does little. The sweet spot is a variable that predicts both, and for SET the most defensible candidate is usually final grade: it is plausibly related to whether a student bothers to respond and to how generously they rate the course (the grading-leniency literature makes the second link explicit). Weighting on grade is therefore far more likely to help than weighting on, say, an ID-number stratum.
A simple worked example shows both the mechanism and its ceiling. Suppose a 100-student class splits evenly into 50 high-grade and 50 low-grade students, and high-grade students rate the course 4.5 on average while low-grade students rate it 3.5. If 40 of the 50 high-grade students respond but only 10 of the low-grade students do, the raw respondent mean is dragged toward the satisfied high-grade group: (40 × 4.5 + 10 × 3.5) ÷ 50 = 4.3. Post-stratifying back to the true 50/50 grade split recovers the population mean of 4.0 — a 0.3-point correction that no amount of chasing a higher overall response rate would have delivered. But the correction is only as good as the assumption that, within each grade band, the handful of low-grade respondents represent the many who stayed silent. If the 10 who answered were the angriest of the 50, post-stratification over-corrects. The repair is real but conditional, never automatic.
Why it matters for course evaluation in practice
Three practical implications follow for a QA process:
- Stop reading the response rate as the verdict. A 35% response rate is not automatically invalid. The right diagnostic is not the rate but evidence about whether respondents differ from non-respondents on grade, engagement, or demographics — the same logic behind early-vs-late wave analysis. A high rate with a skewed respondent profile can be worse than a moderate rate that is representative.
- Use the registry you already have. Because the enrolled population is fully known, course evaluation is one of the few survey settings where post-stratification is genuinely feasible without an external benchmark. Weighting to the class's true grade and demographic distribution is a concrete, auditable correction — far stronger than the usual hand-waving about "low response."
- Weight deliberately, not reflexively. Following Little and Vartivarian, choose auxiliaries that predict both response and rating, accept the variance cost, and report it. Indiscriminate weighting on whatever variables are handy can degrade the estimate. The decision should be documented as part of the methodology, especially where scores feed personnel or accreditation evidence.
Limitations and honest caveats
Weighting is a partial repair, and a sophisticated reader will press on its limits:
- It only corrects bias on observed auxiliaries. If non-respondents differ on something you did not measure — motivation, workload that term, an unrecorded grievance — no amount of weighting on grade and gender will recover it. This is the fundamental, unfalsifiable weakness of all nonresponse adjustment: you are assuming that, within weighting cells, respondents and non-respondents are comparable (missing at random). That assumption is untestable from the respondent data alone.
- Variance-bias trade-off is real. Little and Vartivarian's central result is that weighting can increase the variance of estimates. For a small class, aggressive post-stratification with many cells can produce unstable, over-influenced weights. There is a point where the cure is worse than the disease.
- Small classes break the method. Post-stratification needs enough respondents per cell to estimate cell means. A 15-student seminar cannot be meaningfully post-stratified into grade-by-gender cells. For small cohorts, weighting is often infeasible and the honest move is to widen the reliability and uncertainty caveats instead.
- It does not fix the construct. Weighting addresses who answered, not whether the question measures teaching quality. Bias from grading leniency, halo, or selection into the course is a separate problem that representativeness cannot touch.
How Koji incorporates this
Koji treats representativeness, not raw response rate, as the quantity to manage.
- Representativeness diagnostics over vanity metrics. Koji reports who has responded against the known enrolled population and supports early-vs-late wave comparison, so a programme can see whether late or reluctant responders differ from eager ones — the practical, data-grounded test of nonresponse bias that Groves and Adams and Umbach point to, rather than a single rate.
- Auxiliary-aware reporting where data is available. Where an institution supplies registry auxiliaries (grade band, year, programme) for the enrolled class, Koji can present results against that population distribution and flag when respondents are skewed on a variable plausibly tied to the rating — operationalising the Little and Vartivarian condition that an auxiliary must predict both response and outcome before it is worth weighting on. Koji frames any adjustment as a documented, reversible step, never a silent correction.
- Lifting representativeness at the source. The most reliable defence against nonresponse bias is fewer non-respondents of the systematically different kind. Koji's conversational, mobile-friendly, AI-moderated format is designed to raise engagement among students who abandon traditional grid surveys, and its mid-cycle and formative collection spreads response across the term so the sample is less dominated by the most (or least) satisfied end-of-term voices.
- Honest uncertainty for small cohorts. Where a class is too small to post-stratify, Koji widens uncertainty and small-sample flags rather than implying a weighted point estimate is more precise than it is.
This is the same survey-methodology backbone Koji applies commercially: its core research platform at koji.so uses representativeness checks and auxiliary-aware reporting for customer and market research, where teams over-trust a high completion rate exactly as universities over-trust a high response rate.
Related Resources
- Are Your Respondents Representative? Using Early-vs-Late Wave Analysis to Estimate Nonresponse Bias
- Online vs. Paper Course Evaluations: Do Lower Response Rates Mean Worse Data?
- What Actually Raises Course-Evaluation Response Rates? The Experimental Evidence
- How Many Responses Do You Need for a Reliable Course Evaluation?
- Selection Bias in Course Evaluations: What Goos and Salomons Found
- Why "I Learned a Lot" Can't Be Compared Across Courses: Reference Bias
References
- Groves, R. M. (2006). Nonresponse rates and nonresponse bias in household surveys. Public Opinion Quarterly, 70(5), 646-675. https://doi.org/10.1093/poq/nfl033
- Little, R. J. A., & Vartivarian, S. (2005). Does weighting for nonresponse increase the variance of survey means? Survey Methodology, 31(2), 161-168. https://www150.statcan.gc.ca/n1/pub/12-001-x/2005002/article/9046-eng.pdf
- Adams, M. J. D., & Umbach, P. D. (2012). Nonresponse and online student evaluations of teaching: Understanding the influence of salience, fatigue, and academic environments. Research in Higher Education, 53(5), 576-591. https://doi.org/10.1007/s11162-011-9240-5
Related articles
How Many Responses Do You Need for a Reliable Course Evaluation?
Reliability of a class-average rating rises steeply with the number of students who respond. Generalizability theory — from Marsh''s syntheses to Dzakadzie (2026) — shows raters matter more than items, and a class of 12 cannot be evaluated like a class of 120.
Are Your Respondents Representative? Using Early-vs-Late Wave Analysis to Estimate Nonresponse Bias
A low response rate is not automatically biased — what matters is whether respondents differ from non-respondents. Armstrong and Overton (1977) gave us a cheap diagnostic: compare early and late responders. Here is how to use wave analysis on course-evaluation data and where it breaks down.
What Actually Raises Course-Evaluation Response Rates? The Experimental Evidence
Online course evaluations chronically under-perform paper. We review the experimental evidence — Dommeyer''s grade-incentive trials and Nulty''s adequacy thresholds — on what genuinely lifts response rates, what it costs in data quality, and how to hit a defensible rate without coercion.
Online vs. Paper Course Evaluations: Do Lower Response Rates Mean Worse Data?
Online course evaluations consistently draw lower response rates than in-class paper forms, but the research shows the resulting scores are largely equivalent. Here is what Dommeyer, Nulty, and Stowell actually found, and what an adequate response rate really requires.