New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Methodology9 min read

Can You Weight Your Way Out of a Bad Response Rate? What Post-Stratification Can and Cannot Fix

Weighting can correct the part of course-evaluation non-response your data explains — and only that part. Here is the honest line between post-stratification as a real statistical tool and as a machine for laundering a low response rate.

Koji Education Team

Product · July 29, 2026

Bottom line up front: Weighting can correct the part of non-response that your data can explain — and only that part. If disengaged or dissatisfied students skipped the survey because they were disengaged or dissatisfied, no amount of post-stratification will bring their missing verdict back. This is the single most important thing to understand before you reweight a course-evaluation dataset: it is a genuine statistical tool, not a machine for laundering a low response rate into a trustworthy mean.

Every quality office knows the anxiety. The response rate for a module comes in at 34%, someone asks whether the results are "valid", and a well-meaning analyst proposes weighting the data "to make it representative". Sometimes that is exactly the right move. Often it quietly makes a biased number look authoritative. The difference between the two depends on a distinction most dashboards never surface.

The response rate is not the bias

Start by dismantling the assumption that a low response rate automatically means a biased result. It does not. The standard reference is Groves and Peytcheva's meta-analysis of 59 methodological studies (The Impact of Nonresponse Rates on Nonresponse Bias, Public Opinion Quarterly, 2008, 72(2):167–189, doi:10.1093/poq/nfn011), which found that the response rate is, in general, a poor predictor of nonresponse bias.

The reason is in the arithmetic of bias itself. Nonresponse bias in a mean is, to a close approximation, the product of two things: the nonresponse rate, and the covariance between the propensity to respond and the value being measured. If the students who skip the survey feel roughly the same as those who answer, bias is near zero even at a 30% response rate. If the students who skip are systematically different — the ones who quietly disengaged in week four — then even an 80% response rate can carry real bias. The rate alone cannot tell you which world you are in.

That is the correct frame for weighting. Weighting does not "fix a low response rate". It attempts to reduce the covariance term by making your respondents resemble the full cohort on variables you can observe.

What post-stratification actually does

Post-stratification is the workhorse. You take known population totals — how many enrolled students fall in each year of study, grade band, programme, or attendance bracket — and you reweight respondents so their distribution matches the enrolled cohort. If final-year students are 40% of the class but 60% of respondents, each final-year response is down-weighted and each first-year response is up-weighted until the weighted composition matches the register.

A quick intuition: suppose a large lecture is half first-years, half finalists, but three-quarters of your respondents are finalists who happened to like the course. The raw mean over-represents the happy finalists. Post-stratification restores the 50/50 split, and the corrected mean drops toward what the whole class would have said — if year of study is what distinguishes responders from non-responders.

When you only have the marginal totals (the year split and the programme split separately, but not the joint cross-tabulation), raking — iterative proportional fitting — gets you there. When you can model an individual response probability from auxiliary variables, inverse-probability weighting does the same job. All three rest on the same assumption.

The assumption that decides everything

That assumption has a name. In Rubin's taxonomy (Little and Rubin, Statistical Analysis with Missing Data), data are Missing At Random (MAR) when the probability of responding depends only on things you can observe. Under MAR, weighting on those observed variables is unbiased. Data are Missing Not At Random (MNAR) when the probability of responding depends on the unobserved value itself — when students skip the survey because of the very opinion the survey would have recorded.

Course evaluation is the textbook habitat of MNAR. The most plausible non-response mechanism is precisely the opinion you are trying to measure: the disengaged student who stopped attending, the dissatisfied student who could not be bothered, the one who mentally checked out. Their silence is correlated with their verdict. Weighting by year, programme and grade band adjusts for demographic imbalance, but it recovers their missing rating only to the extent that those demographics correlate with the opinion. The residual — the part of the dissatisfaction not captured by any auxiliary variable — survives the weighting untouched.

This is why an honest analyst treats weighting as necessary but not sufficient. It buys back the MAR portion of the bias and leaves the MNAR portion exactly where it was.

What to do instead of pretending

Three disciplines separate defensible weighting from number-laundering.

First, collect auxiliary data worth weighting on. Weighting is only as good as the variables you can align to the enrolled frame. A student registry gives you programme, year and demographics; the LMS gives you engagement and attendance proxies; the assessment system gives you grade bands (used carefully, and only after grades are finalised). The more your auxiliary variables correlate with both response propensity and the rating, the more bias weighting removes.

Second, run a sensitivity analysis, not a point estimate. Because you can never confirm MAR from the data alone, the mature move is to ask: how wrong could this be? Assume non-respondents would have rated, say, 0.4 points lower on average, and recompute. Report the range. A course whose mean survives the pessimistic scenario is genuinely fine; one that flips from "above threshold" to "below" under a modest MNAR assumption was never as solid as its point estimate suggested.

Third, report the assumptions alongside the number. State which variables you weighted on, what the unweighted estimate was, and how much the weighting moved it. A weighted mean presented without its scaffolding is less trustworthy than an unweighted one presented honestly, because it hides a modelling choice inside a figure that looks like a fact.

"But isn't this just p-hacking with survey weights?"

The strongest objection deserves a straight answer. If an analyst can choose weighting variables until the mean lands where they would like, weighting becomes a tool for manufacturing a preferred result — a garden of forking paths with survey weights. The worry is legitimate and the guardrails are well established: pre-specify the weighting scheme before you see the outcome; weight to correct representativeness, never to move a specific instructor's score; always report the unweighted estimate next to the weighted one; and use weighting to widen your expressed uncertainty, not to project false precision. Weighting done in the open, on pre-declared variables, is the opposite of p-hacking. Weighting reverse-engineered to a target is exactly p-hacking. The procedure is identical; the governance is everything.

A second objection: weighting inflates variance, so a weighted estimate from a small, skewed sample can have such wide confidence intervals that it says almost nothing. That is a feature, not a bug — it is the data telling you that 34% of a class of 60, badly skewed, cannot support confident conclusions. Better an honest wide interval than a precise wrong one. (See our companion pieces on effect size and over-interpretation and why instructor averages are mostly noise.)

Where Koji fits

Weighting is a repair applied after the fact to a thin, ambiguous signal. The better lever is to reduce how much repair you need. Koji's AI-moderated conversational interviews improve representativeness at the source in two ways. Higher-quality, less tedious response experiences and formative, mid-cycle collection lift participation among students who would abandon a 30-item Likert grid — shrinking the nonresponse rate that feeds the bias term. And because Koji captures the reasoning behind a rating, not just the number, you gain qualitative signal about who is missing and why — interpretive context a bare weighted mean can never supply.

Koji does not claim to solve MNAR non-response; nothing can fully solve it. It gives you a richer instrument and honest partial data, so that when you do weight, you are adjusting a signal that carries its own explanation. The same conversational interview engine underpins koji.so for general user and customer research, where self-selection on opinion distorts feedback just as badly. Treat weighting as a scalpel, not a spray can. Use it, declare it, bound it — and never let it substitute for asking better questions of more of the cohort.