Shorter Surveys Without Losing Coverage: Planned Missingness for Course Evaluations
Planned missing data designs let you cover more questions while each student answers fewer. Here is how the three-form design works, what the evidence says, and where it fits in course evaluation.
Koji Education Team
Product
In brief
You can ask a full battery of evaluation questions and keep each student's survey short — by deliberately showing every student only a subset of items. This is planned missingness (also called a split-questionnaire or matrix-sampling design), and its statistical foundation is that the missing data are missing completely at random (MCAR) by design, so modern estimators (full-information maximum likelihood or multiple imputation) recover unbiased population estimates. Graham, Taylor, Olchowski and Cumsille (2006, Psychological Methods) describe the canonical three-form design, which lets researchers collect data on roughly 33% more questions than any single respondent answers, at the cost of statistical power that must be planned for. It is a genuine tool for the over-survey/fatigue problem — but it trades away per-item sample size, so it is not free, and it fits programme-level measurement better than high-stakes individual-instructor comparison.
What the research says
Graham and colleagues (2006) formalised planned missing data designs as a way to reduce respondent burden without abandoning breadth. Their flagship is the three-form design: the item pool is divided into a common block (X) that everyone answers and three variable blocks (A, B, C). Each respondent receives X plus two of the three variable blocks, so any given student sees three-quarters of the variable items and none sees all of them. Because the assignment of which block is omitted is under the researcher's control and randomised, the resulting missingness is MCAR — the single most benign missing-data mechanism. Under MCAR, FIML and multiple imputation produce unbiased parameter estimates; the price paid is in precision (larger standard errors for the relationships that rely on the partially observed pairs), not in bias. Graham et al. calculate that the design yields data on about 33% more questions than a respondent actually answers, and they also describe a two-method measurement design that pairs a cheap measure administered to everyone with an expensive gold-standard measure administered to a subsample.
Subsequent methodological work reinforced and extended this. Rhemtulla and Hancock, and Little and Rhemtulla, showed how planned missingness applies in educational and developmental research and how to choose block allocations to protect power for the estimates you care about. A commentary in Industrial and Organizational Psychology explicitly frames planned missingness as "an underused but practical approach to reducing survey and test length" — precisely the course-evaluation pain point. The consistent message across this literature: planned missingness is well understood, statistically principled, and under-adopted outside methodology circles.
Two connections matter for evaluators. First, planned missingness is the benign twin of the missing-data problem QA staff already worry about: ordinary non-response is often not random (MNAR), which biases results, whereas planned missingness is MCAR by construction and therefore analytically tractable. Second, it directly targets satisficing and straightlining — the careless responding that long surveys provoke — by keeping each form short enough that students stay engaged.
Why it matters for course evaluation in practice
Course evaluation lives in permanent tension between breadth and burden. Quality offices want items on teaching, assessment, feedback, workload, resources, inclusion, and learning gain; students, surveyed in every module every term, suffer survey fatigue, and long instruments depress both response rates and response quality. Planned missingness offers a way out of the trade-off for the questions that are analysed at aggregate level:
- More coverage per student-minute. A department can field a 24-item item bank while each student answers ~16, keeping completion time down and reducing drop-off and straightlining.
- Better data quality on what is answered. Shorter forms mean less fatigue, so the answers you do get are less contaminated by satisficing — a quality gain that can offset the sample-size cost.
- A principled alternative to simply cutting items. The usual response to "the survey is too long" is to delete questions, permanently losing that information. Planned missingness keeps the item in the bank and samples it, preserving programme-level coverage.
- Natural fit for aggregate reporting. Because estimates are pooled across students, the per-item sample loss is absorbed at the cohort level where course evaluation decisions are (properly) made.
Limitations and honest caveats
Planned missingness is powerful but far from a universal fix, and a rigorous reader should weigh several real constraints.
- It costs power, and power must be pre-planned. Every relationship that depends on two items in different variable blocks is estimated on a reduced sample. If you have not planned the block allocation around your key comparisons, precision for those comparisons can suffer. This is a design decision that must be made before data collection, not patched afterwards.
- It requires principled analysis. The benefits hold only if data are analysed with FIML or multiple imputation. Naive listwise deletion or pretending the blanks are "no opinion" throws the advantage away and can reintroduce bias.
- It is ill-suited to high-stakes individual scores. For a personnel decision about one instructor in a small class, deliberately reducing the per-item sample is the wrong direction. Planned missingness is a programme- and aggregate-level tool, not a tenure-file tool.
- Complexity and transparency. Split forms complicate survey logistics, item-level reporting, and communication to committees who expect "everyone answered everything." The administrative and explanatory overhead is real.
- Not a substitute for reducing over-surveying. If students are drowning in evaluations, the first-order fix may be to evaluate less often or sample students, not just to shorten each form. Planned missingness complements, but does not replace, sensible survey governance.
The honest framing: planned missingness elegantly dissolves the breadth-versus-length trade-off for aggregate measurement, provided you plan power in advance and analyse with modern missing-data methods — and provided you are not trying to score an individual instructor in a small class.
How Koji incorporates this
Koji for Education is designed to attack the survey-length problem from several directions, of which structured item sampling is one:
- Adaptive, conversational elicitation shortens the effective instrument. Rather than showing every student a fixed long battery, Koji's AI-moderated interview probes the areas that matter for each respondent, so breadth is achieved conversationally instead of by forcing all items on all students — a design cousin of planned missingness that keeps each student's experience short.
- Item-bank configuration with sampling. Institutions can maintain a wide bank of
scale,single_choice,multiple_choice,ranking,yes_noandopen_endedquestions and administer subsets, so cohort-level coverage is preserved without any single student answering everything. - Aggregate-first, uncertainty-honest reporting. Because Koji reports at cohort level with explicit sample size and uncertainty, it fits the aggregate logic where planned missingness pays off, and it flags when a per-item sample is too small to bear an individual comparison — directly respecting the "not for high-stakes individual scores" caveat.
- Fatigue and quality safeguards. Shorter, adaptive forms reduce the satisficing and straightlining that long surveys provoke, improving the quality of the responses that are collected.
- Principled handling of missing data. Koji is built to treat unasked items as by-design missing rather than as substantive "no opinion" responses, aligning with the requirement that planned missingness be analysed correctly rather than naively.
These are described as mechanisms designed to mitigate survey burden while protecting coverage; they do not repeal the statistical cost of collecting less data per student, and power still has to be planned. Koji's core research platform at koji.so applies the same adaptive-interview approach to product and customer research, where long questionnaires create the same fatigue and coverage dilemma.
Planning the block allocation
The single most important design decision in a planned-missing survey is which items share a form and which are split across forms, because power for a relationship depends on how often its two items are answered by the same student. The practical rule is to place the items you most need to analyse together — for example, an assessment-clarity item and the learning-gain item whose association you care about — in the common block that everyone answers, or in the same variable block, so their pairwise sample stays large. Items reported only as standalone marginals, where no cross-item correlation matters, are the safest to distribute across the variable blocks. Before fielding, run a quick power check: simulate the design with plausible correlations and confirm the standard errors on your headline estimates are acceptable. This up-front planning is what separates a defensible planned-missing design from an accidental mess of holes, and it is why the technique is a deliberate measurement choice rather than a convenient way to shorten a survey after the fact.
Related Resources
- How long should a course evaluation be? Length and data quality
- Missing data: MAR, MNAR and multiple imputation
- Satisficing and straightlining in course evaluations
- Non-response weighting and post-stratification
- What Cronbach's alpha tells you about reliability
- Single-item vs multi-item global ratings
References
- Graham, J. W., Taylor, B. J., Olchowski, A. E., & Cumsille, P. E. (2006). Planned missing data designs in psychological research. Psychological Methods, 11(4), 323–343. https://doi.org/10.1037/1082-989X.11.4.323
- Little, T. D., & Rhemtulla, M. (2013). Planned missing data designs for developmental researchers. Child Development Perspectives, 7(4), 199–204. https://doi.org/10.1111/cdep.12043
- Rhemtulla, M., & Hancock, G. R. (2016). Planned missing data designs in educational psychology research. Educational Psychologist, 51(3–4), 305–316. https://doi.org/10.1080/00461520.2016.1208094
- Enders, C. K. (2010). Applied Missing Data Analysis. Guilford Press.
Related articles
Cronbach's Alpha and Course-Evaluation Reliability: What It Does and Doesn't Tell You
A high Cronbach's alpha on your course-evaluation instrument is widely read as proof of a "reliable" survey. The psychometric literature says that reading is largely wrong. What alpha actually measures, where it misleads, and what to report instead.
Is Your Missing Course-Evaluation Data Random? MAR, MNAR, and What to Do About It
A low response rate is a missing-data problem. Rubin's MCAR/MAR/MNAR framework explains when a course-evaluation mean is biased and when multiple imputation or maximum likelihood can help.
Why Students Click Straight Down the Middle: Satisficing in Course Evaluations
A large share of students do not answer evaluation questions carefully — they straightline, speed through, and pick the easy option. Barge and Gehlbach (2012) showed this "satisficing" doesn't just add noise; it inflates your reliability and validity statistics, making bad data look good.
How Long Should a Course Evaluation Be? Questionnaire Length, Breakoff, and Answer Quality
What the survey-methodology evidence says about questionnaire length: longer instruments depress participation and degrade answers to later questions, but ruthless shortening is not automatically the answer. A research-grounded guide for designing course evaluations.