New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
analysis-reporting9 min read

Correcting for Who Chose to Respond: The Heckman Selection Model for Course-Evaluation Nonresponse

When students who respond to course evaluations differ from those who skip them on the very thing you are measuring, reweighting cannot help. The Heckman selection model models response and rating jointly to correct for selection on unobservables — under fragile assumptions this article makes explicit.

Koji Education Team

Product

In brief

When students who bother to complete a course evaluation differ systematically from those who skip it — and they differ on the very thing you are measuring — no amount of reweighting on observed characteristics will fix the bias, because the problem lives in the unobserved. The Heckman selection model (Heckman, 1979) is the classic parametric remedy: it models who responds and what they say jointly, then inserts a correction term (the inverse Mills ratio) into the rating equation. It can recover an unbiased estimate under strong assumptions, but decades of applied work show those assumptions are fragile, so treat Heckman output as a sensitivity check, not a magic fix.

What the research says

James Heckman's 1979 Econometrica paper, "Sample Selection Bias as a Specification Error," reframed a messy sampling problem as an ordinary omitted-variable problem (Heckman, 1979). His insight: if the probability of appearing in your sample depends on the outcome you are studying, then the observed cases carry a hidden regressor — the expected value of the error term conditional on being selected — and omitting it biases every coefficient in the outcome model.

Heckman's two-step estimator makes this operational. In the first step, a probit model predicts selection (here: did the student submit an evaluation?) from a set of predictors. From that probit you compute the inverse Mills ratio (IMR) for each respondent — a number that is large when a student was unlikely to respond yet did anyway. In the second step, you run the outcome regression (the rating model) on respondents only, adding the IMR as an extra covariate. If its coefficient is statistically distinguishable from zero, selection was biasing your estimates; the corrected coefficients net that bias out. A full-information maximum-likelihood (FIML) variant estimates both equations simultaneously and is generally more efficient.

The crucial ingredient is an exclusion restriction: at least one variable that affects whether a student responds but does not directly affect the rating they would give. Classic candidates in course evaluation are administrative nudges — the number of reminder emails sent, whether the instructor set aside in-class time to complete the survey, or the response deadline — factors plausibly linked to participation but not to how good the teaching was.

Two decades of methodological review temper the enthusiasm. Puhani (2000) surveyed the Monte Carlo evidence and concluded that the two-step estimator behaves badly under collinearity: when the selection and outcome equations share most of their predictors (a weak exclusion restriction), the IMR is nearly a linear function of the other covariates, standard errors explode, and estimates can be worse than uncorrected OLS. Bushway, Johnson, and Slocum (2007), auditing applied criminology, found the technique routinely misapplied — identification resting entirely on the model's non-linear functional form rather than a genuine exclusion restriction, hazard rates miscomputed, and standard errors that ignored the first-stage estimation. Their verdict, echoed widely, is that Heckman corrections are only as credible as the exclusion restriction behind them.

Why it matters for course evaluation in practice

Course-evaluation response rates commonly sit between 30% and 60%, and non-response is rarely random. The concern is not merely that responders are fewer but that they are different in ways that correlate with the rating: highly satisfied students and deeply aggrieved students are both over-motivated to respond, while the indifferent middle stays silent. This is a textbook missing-not-at-random (MNAR) pattern, and it is exactly the case where reweighting on observed variables — post-stratification, non-response weighting — cannot help, because those methods assume the missingness is explainable by things you measured.

The Heckman framework gives a quality-assurance office three concrete benefits. First, the IMR coefficient is a formal test for whether selection on unobservables is distorting a reported mean or a group comparison. Second, when the test fires, the model yields a corrected estimate you can report alongside the naive average. Third — and often most valuable — the exercise forces you to name your exclusion restriction, which disciplines the whole conversation: if you cannot point to a variable that drives participation but not the rating, you have learned that your data cannot separate low scores from low turnout, and you should report the ambiguity honestly rather than a falsely precise mean.

Limitations and honest caveats

The Heckman model is powerful and fragile in equal measure, and a PhD reader will press on four points.

  1. The exclusion restriction is rarely watertight. Reminder intensity or in-class time may themselves correlate with instructor conscientiousness, which correlates with teaching quality — puncturing the restriction. Without a defensible excluded variable, identification comes only from the bivariate-normal functional form, which is an assumption, not evidence.
  2. Bivariate normality of the two error terms is a strong, untestable assumption. When it fails, the correction can inject bias rather than remove it. Semiparametric alternatives exist but demand larger samples than a single course provides.
  3. Collinearity destroys precision. As Puhani (2000) showed, a weak exclusion restriction can leave you worse off than doing nothing. A near-significant IMR with huge standard errors is not a fix; it is a warning.
  4. Small samples. A 40-student class cannot support a stable two-equation model. Heckman corrections belong at the programme or institution level, pooling many courses, not on a single roster.

The honest framing is that Heckman is a sensitivity analysis, not a truth machine. If the naive mean and the corrected mean tell the same story, your conclusion is robust to selection on unobservables; if they diverge, you have quantified how much your reading depends on assumptions you cannot verify — which is itself a finding worth reporting. It is a natural companion to the assumption-free worst-case bounds of Manski's partial-identification approach.

How Koji incorporates this

Koji's design philosophy is to attack non-response before it forces you into a fragile statistical correction, and to give analysts the raw material a Heckman model needs when correction is unavoidable.

  • Raising participation to shrink the selection problem. Because Heckman corrections degrade as response rates fall, the first line of defence is turnout. Koji's AI-moderated conversational format is designed to feel less like a chore than a static grid, and mid-cycle and formative collection windows capture the indifferent middle before end-of-term fatigue sets in — directly targeting the students most likely to be missing-not-at-random.
  • Capturing paradata for the selection equation. Koji logs the participation signals a first-stage probit relies on — reminders delivered, invitation-to-completion timing, device, and whether collection was embedded in a session — so an institutional analyst can specify a credible selection model with a real candidate exclusion restriction rather than guessing.
  • Bias-aware reporting. Koji's reporting layer is built to flag low and skewed response rates rather than silently averaging whatever came in, so a 35%-response course mean is never presented with the same confidence as an 80%-response one.
  • Reaching the silent through richer signal. Where a Likert grid gives a non-responder no voice at all, Koji's conversational probes and open-text thematic analysis lower the effort of saying something substantive, pulling in students who would otherwise abstain. Koji's core research platform at koji.so applies the same AI-moderated interview engine to product and customer research, where self-selection into feedback is an identical hazard.

Koji is designed to mitigate selection bias by improving and instrumenting participation; it does not claim to eliminate it. Where selection on unobservables remains, the correct move is to report a corrected estimate or an honest bound, not to pretend the responders speak for everyone.

Frequently asked questions

What is the difference between the Heckman model and simply reweighting responses?

Non-response weighting and post-stratification correct for missingness that is explainable by observed characteristics (missing-at-random). The Heckman model targets the harder case where response depends on the unobserved rating itself (missing-not-at-random) — for example, when only the very satisfied and the very angry bother to respond.

What is the inverse Mills ratio?

It is a term computed from the first-stage selection (probit) model that captures, for each respondent, the expected value of the outcome error given that they chose to respond. Adding it to the rating regression absorbs the selection bias; a significant coefficient on it indicates selection was distorting your estimates.

What is an exclusion restriction and why does it matter so much?

It is a variable that influences whether a student responds but not the rating they would give (e.g., number of reminder emails). Without one, the Heckman correction is identified only by its functional-form assumptions, which makes the results unreliable — the single most common reason applied Heckman models fail.

Can I run a Heckman correction on one small class?

No. The two-equation model needs a substantial sample and variation in the selection predictors; a single roster of 30-50 students cannot support it. Apply it at the programme or institution level across many courses.

Is the Heckman model better than reporting the raw average?

Only when the exclusion restriction is credible and collinearity is low. Puhani (2000) showed that with a weak exclusion restriction the correction can be worse than the uncorrected mean. Treat it as a sensitivity check: report both numbers and note whether they agree.

Does a high response rate make selection correction unnecessary?

It reduces the risk substantially but does not guarantee its absence — even an 80%-response survey can be biased if the missing 20% are systematically different on the outcome. High response rates mainly narrow the range of plausible bias, which is why raising participation is the first-best strategy.

References

Related resources