When You Cannot Assume the Non-Responders Are Like the Responders: Manski Bounds for Course Evaluation
Weighting and imputation assume the silent majority resemble the responders. Partial identification refuses that assumption and reports the interval the true mean could occupy, showing exactly how much your conclusion rests on belief.
Koji Education Team
Product
In brief
When 40% of a class responds, the honest question is not "what is the mean rating?" but "what could the mean rating be, given that we know nothing about the 60% who stayed silent?" Partial identification answers exactly that: instead of assuming the non-responders resemble the responders (as weighting and imputation quietly do), it reports the interval of values the true mean could take under no such assumption. Manski's worst-case bounds are often wide — sometimes uninformatively so — but their width is the point: they show precisely how much your conclusion depends on assumptions the data cannot verify.
What the research says
The framework is Charles Manski's. In "Nonparametric Bounds on Treatment Effects" (Manski, 1990, American Economic Review 80(2):319-323) he showed that even when a parameter — a treatment effect, a population mean — cannot be pinned to a single value ("point-identified") from the available data, the data plus minimal, credible assumptions still confine it to a set. The classic worst-case bound on a mean when some responses are missing replaces every missing value once with the lowest possible score and once with the highest; the resulting interval is guaranteed to contain the truth without any assumption about why data are missing. These bounds are sharp — you cannot make them narrower without adding assumptions.
Manski's Partial Identification of Probability Distributions (2003, Springer) develops the general theory: identification is a spectrum, and the analyst's job is to trace how the identified set shrinks as assumptions are layered on. Manski and Pepper (2000, Econometrica 68(4):997-1012, "Monotone Instrumental Variables: With an Application to the Returns to Schooling") is the key applied refinement — they show that weak, defensible monotonicity assumptions (for example, "response propensity is weakly related to satisfaction in a known direction") can tighten worst-case bounds substantially while remaining far more credible than the point-identifying assumptions imputation requires. Tamer (2010, Annual Review of Economics 2:167-195, "Partial Identification in Econometrics") is the standard review, framing the approach as a deliberate trade: you give up the false comfort of a single number in exchange for conclusions that survive weaker, more honest assumptions.
The philosophical core is Manski's "Law of Decreasing Credibility": the credibility of an inference falls as the strength of the assumptions used to reach it rises. Point estimates sit at the low-credibility, high-precision end; worst-case bounds sit at the high-credibility, low-precision end; and the interesting work happens in between.
Why it matters for course evaluation in practice
Nonresponse is the defining problem of modern course evaluation, and the field's standard fixes all rest on an assumption the data cannot test. Complete-case reporting assumes responders and non-responders have the same mean (missing completely at random). Nonresponse weighting assumes missingness is ignorable given the weighting variables. Multiple imputation assumes a correctly specified model for the missing values. Each produces a confident single number — and each can be badly wrong if the silent majority differ systematically from the vocal minority, which nonresponse research repeatedly finds they do.
Partial identification offers a different report. With a 1-5 scale and a 40% response rate showing a responder mean of 4.2, the worst-case bound is stark: if every non-responder would have given a 1, the true mean is about 0.4x4.2 + 0.6x1 = 2.28; if every non-responder would have given a 5, it is about 0.4x4.2 + 0.6x5 = 4.68. The true mean lies in [2.28, 4.68] with no assumption about the non-responders at all. That interval is wide — deliberately — and it reframes the debate: a department cannot claim "4.2, solidly good" without an assumption, but it also cannot be accused of hiding a disaster if the upper bound is still acceptable. Adding a mild monotonicity assumption (say, silent students are on average no more satisfied than responders) collapses the interval on one side and often yields a genuinely useful, defensible range.
Used this way, partial identification is not defeatism; it is a discipline that makes the response-rate conversation precise. It tells you when your data are strong enough to support a claim regardless of the non-responders (the whole bound sits above your threshold) and when they are not (the bound straddles it), which is far more actionable than a point estimate whose validity is silently assumed. This complements our MCAR/MAR/MNAR framing, which explains why the assumptions matter; Manski bounds show what you can say when you refuse to make them.
Limitations and honest caveats
The obvious cost is width. Assumption-free worst-case bounds are frequently too wide to guide a decision — with a low response rate they can span most of the scale, and reporting "the true mean is somewhere between 2.3 and 4.7" will frustrate stakeholders who want a number. The method's value then lies in layering assumptions transparently and showing how each one narrows the interval, but each added assumption is itself a claim that must be defended; monotonicity is credible in some settings and not others.
Bounds also address only the identification problem, not sampling uncertainty. A complete treatment needs confidence intervals for the bounds themselves (the Imbens-Manski approach), and small classes make those intervals wide on top of the already-wide identified set. Partial identification is a framework for reasoning about assumptions, not a mechanical procedure that returns one defensible answer — two analysts who accept different monotonicity assumptions will report different intervals, and the honest presentation shows the whole ladder from assumption-free to strongly-assumed. Finally, bounds discipline aggregate claims; they do not rescue an evaluation whose response rate is so low that even the tightest defensible assumption leaves the decision-relevant threshold inside the interval. Sometimes the correct output is "we cannot tell — collect more responses."
How Koji incorporates this
Koji attacks the nonresponse problem at its source and reports what remains honestly. Because Koji's conversational, AI-moderated format is designed to raise completion — a dialogue that adapts to the student is less likely to be abandoned than a static grid — the response rate that drives the width of any Manski bound is higher to begin with, which mechanically tightens every interval before a single assumption is made. Its mid-cycle and always-on collection spreads response opportunities across the term rather than staking everything on one end-of-term email, further shrinking the silent fraction.
For what nonresponse remains, Koji's reporting is built to present ranges and assumptions rather than a lone point estimate dressed up as fact. Its bias-aware reporting can show a responder mean next to the assumption-free bound and the effect of a stated monotonicity assumption, so a quality committee sees explicitly how much a "good" score depends on beliefs about the students who did not answer. Koji's quality scoring and wave-level metadata provide exactly the covariates a monotonicity or weighting assumption would use, so the ladder from assumption-free to strongly-assumed can be drawn from real data. Koji does not pretend that a clever model recovers the opinions of students who never responded — it is designed to reduce nonresponse and to be candid about the identified range that is left. Teams running customer and product research on Koji's core platform at koji.so meet the same wall whenever a survey's response rate is low, and the same "bound it, don't assume it" discipline keeps their conclusions defensible.
Frequently asked questions
What does "partial identification" mean?
It means the data plus credible assumptions confine a quantity to a set of possible values rather than a single point. Instead of one estimate that requires strong assumptions, you report the interval consistent with weaker, more defensible ones.
What is a Manski worst-case bound for a mean?
Replace every missing response once with the lowest possible value and once with the highest, and compute the mean each time. The two results bracket the true mean with no assumption about why data are missing. The interval is sharp — it cannot be narrowed without adding assumptions.
Aren't these bounds uselessly wide?
Often the assumption-free bound is wide, and that honestly reflects how little a low response rate tells you. The method's power is in layering mild, defensible assumptions (such as monotonicity) that tighten the interval, while showing exactly how much each assumption buys.
How is this different from imputation or weighting?
Imputation and weighting assume the missing responses can be predicted from the observed ones (ignorable missingness) and return a single number. Partial identification refuses that assumption and instead reports the range the mean could occupy, making the dependence on assumptions visible rather than hidden.
When are bounds actually decisive?
When the entire identified set falls on one side of your decision threshold. If even the worst-case bound is above "acceptable", the conclusion holds regardless of the non-responders. If the bound straddles the threshold, the data cannot settle the question without an assumption.
Do Manski bounds account for small-sample noise?
The basic bounds address identification, not sampling error. For a full account you add confidence intervals for the bounds themselves (the Imbens-Manski method), which widen further for small classes — a reminder that a tiny, low-response class may simply be uninformative.
References
- Manski, C. F. (1990). Nonparametric bounds on treatment effects. American Economic Review, 80(2), 319-323. https://www.jstor.org/stable/2006592
- Manski, C. F. (2003). Partial Identification of Probability Distributions. Springer. https://doi.org/10.1007/b97478
- Manski, C. F., & Pepper, J. V. (2000). Monotone instrumental variables: With an application to the returns to schooling. Econometrica, 68(4), 997-1012. https://doi.org/10.1111/1468-0262.00144
- Tamer, E. (2010). Partial identification in econometrics. Annual Review of Economics, 2, 167-195. https://doi.org/10.1146/annurev.economics.050708.143401
Related resources
- Is a Low Response Rate Biasing Your Course Evaluations? MCAR, MAR, and Missing-Not-at-Random
- Deleting Incomplete Surveys Throws Away Good Data: Multiple Imputation for Missing Course-Evaluation Responses
- Can You Weight Your Way Out of a Low Response Rate? Post-Stratification and Nonresponse Adjustment in Course Evaluations
- Are Your Respondents Representative? Using Early-vs-Late Wave Analysis to Estimate Nonresponse Bias
- Is 4.2 Really Worse Than 4.4? Bayesian Estimation and Credible Intervals for Course Evaluations
- Can Instructors Buy Better Ratings With Easy Grades? Instrumental Variables and the Grade–Evaluation Endogeneity Problem
Related articles
Are Your Respondents Representative? Using Early-vs-Late Wave Analysis to Estimate Nonresponse Bias
A low response rate is not automatically biased — what matters is whether respondents differ from non-respondents. Armstrong and Overton (1977) gave us a cheap diagnostic: compare early and late responders. Here is how to use wave analysis on course-evaluation data and where it breaks down.
Is a Low Response Rate Biasing Your Course Evaluations? MCAR, MAR, and Missing-Not-at-Random
A low response rate does not automatically bias a course-evaluation score — the missing-data mechanism (MCAR, MAR, or MNAR) does. How to diagnose nonresponse bias instead of chasing a rate threshold.
Deleting Incomplete Surveys Throws Away Good Data: Multiple Imputation for Missing Course-Evaluation Responses
Listwise deletion of partially-complete course evaluations shrinks your sample and can bias results. How multiple imputation (Rubin 1987; Schafer & Graham 2002) handles item nonresponse honestly.
Is 4.2 Really Worse Than 4.4? Bayesian Estimation and Credible Intervals for Course Evaluations
Bayesian estimation answers the question institutions actually ask — how probable is it that this instructor is below standard? — with credible intervals and a region of practical equivalence instead of raw averages or p-values.