How to Cut a Course-Evaluation Survey in Half Without Dropping a Single Question: Planned Missing-Data Designs
Survey fatigue is usually treated as a reason to delete questions. There is a better answer from the measurement literature: planned missing-data designs. By randomly giving each student a subset of the items and reassembling the full picture statistically, you can ask everything you need while every respondent answers far less. Here is how the three-form design works, when it is valid, and where it breaks.
Koji Education Team
Product · August 22, 2026
Bottom line up front: The standard response to survey fatigue is to shorten the questionnaire by cutting questions — which means giving up information you wanted. There is a rigorous alternative that the measurement community has used for decades and most universities have never applied to course evaluation: planned missing-data designs (PMDD). The idea, formalised by Graham, Taylor, Olchowski, and Cumsille in Psychological Methods (2006), is to keep your full item set but give each respondent only a random subset, then reconstruct the complete-data estimates statistically. Their "three-form design" — a form of matrix sampling — lets you collect data on roughly a third more questions than any single respondent answers, while cutting each student's burden. Because the missingness is created by the researcher at random, it is missing completely at random (MCAR) by construction — the one benign, well-understood form of missing data. This is the constructive counterpart to a problem we have written about diagnostically in survey fatigue and over-evaluation: instead of accepting shorter, thinner instruments, you can design missingness deliberately.
Why "just delete questions" is the wrong default
When response rates fall and students complain the evaluation is too long, the reflex is to trim the instrument to its "most important" items. This has two costs. First, you lose the very items that provide texture — the assessment-specific, resource-specific, or pacing-specific questions that make a result actionable rather than a generic satisfaction number. Second, deciding which items to cut is itself a judgement that bakes in assumptions about what matters, often before you have the data to know. You end up with a short survey that is easy to complete and hard to act on.
Planned missingness refuses the trade-off. You do not shorten the instrument; you shorten each student's share of it.
The three-form design, concretely
Divide your evaluation items into four blocks. One block, X, contains the core items you want everyone to answer (say, the overall and the two or three questions you benchmark on). The other three blocks — A, B, and C — are split across three versions of the form:
- Form 1: X, A, B
- Form 2: X, A, C
- Form 3: X, B, C
Each student, assigned a form at random, sees the core block plus two of the three optional blocks — three quarters of the items, not all of them. Every optional block is answered by two-thirds of respondents, and every pair of items across blocks is observed in some subset of students, which is what lets modern estimation recover the relationships among all items. Graham and colleagues showed this design lets you field about 33% more items than any respondent completes. Variants scale further: a six-form design with five item sets, or a ten-form design with six, push the ratio higher at the cost of needing a larger sample.
The reconstruction is not guesswork. Because the data are MCAR by design, full-information maximum likelihood (FIML) or multiple imputation produce unbiased estimates of means, correlations, and scale scores across the whole item set — with a modest, quantifiable loss of statistical power that Graham's power tables let you plan for in advance. A 2019 applied study, Maximizing data quality and shortening survey time: Three-form planned missing data survey design, demonstrated the approach doing exactly what the title promises.
Why this is better missingness, not worse
Course evaluation's usual missing-data problem is the dangerous kind: students who did not respond differ systematically from those who did, producing non-response bias that is often missing not at random (MNAR) — the hardest case to correct. Planned missingness is the opposite. You chose which items each student skipped, using a random mechanism unrelated to any student characteristic or to how they would have answered. That is MCAR, the assumption every standard missing-data method is built to handle. In effect, PMDD converts an uncontrolled, biasing form of missingness into a controlled, benign one — and buys shorter surveys in the bargain.
This does not replace the tools that address unit non-response, such as post-stratification weighting. It addresses a different axis: how much you ask each student who does respond. The two are complementary. A well-designed cycle uses planned missingness to keep per-student burden low and weighting to correct for who shows up.
"But doesn't a smaller sample per item just make everything noisier — and isn't this too complex for a quality office?"
This is the strongest and most practical objection. Two honest concessions and one rebuttal.
Concession one: yes, each optional item is answered by fewer students, so estimates for those items carry wider confidence intervals than if everyone answered everything. Planned missingness is not free; it trades a little precision per item for lower burden and higher completion. Graham's power tables exist precisely so you can decide whether that trade is acceptable before fielding — and for the core block X, which everyone answers, you lose nothing.
Concession two: the analysis is more sophisticated than averaging a column. FIML and multiple imputation are not built into a typical spreadsheet, and a quality office running the design by hand would need statistical support. If the reporting layer cannot handle principled missing-data estimation, the design's advantages evaporate — you would be left with awkward gaps and no valid way to fill them.
The rebuttal: this is an argument for the right tooling, not against the method. The measurement community adopted planned missingness because for long instruments it is strictly better than the alternative of cutting content. The barrier has never been the statistics — those are 20 years mature — but whether your platform can assign forms at random and reconstruct estimates correctly. That is a solved problem for software; it is only unsolved for paper forms and naïve dashboards. And crucially, planned missingness suits structured items (scales, single-choice) far better than open text, so it complements rather than replaces the qualitative depth that makes feedback actionable.
Where Koji fits
Koji for Education sidesteps much of the fatigue problem at its root, and planned missingness is a natural extension of how it already works. Rather than presenting every student with a fixed, exhaustive form, Koji's AI-moderated conversational interview adapts what it asks — probing the areas a given student actually has something to say about, instead of marching everyone through an identical long questionnaire. That is a behavioural cousin of matrix sampling: each respondent contributes depth where it is informative and is not detained by items irrelevant to them. Its six structured question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) let you keep a compact benchmarked core while rotating optional structured items across respondents, and its automatic thematic analysis aggregates the open-ended material that planned missingness deliberately leaves to conversation rather than to fixed blocks. Reporting is at the programme and institution level, where the reassembled, cohort-wide picture belongs. The same adaptive interview engine powers general research on the main Koji platform. To be precise about the claim: adaptive questioning is not identical to a formal three-form PMDD, and where you need strict statistical reconstruction across a fixed instrument you still need FIML or imputation in your analysis layer — but the design goal is the same, and the conversational approach achieves much of the burden reduction without asking a quality office to run imputation by hand.
The takeaway
Survey fatigue does not force a choice between "long and ignored" and "short and useless." Planned missing-data designs — matrix sampling with principled reconstruction — let you keep the questions you need while asking each student far fewer of them, and they turn an uncontrolled, biasing kind of missing data into the one benign kind. It is one of the best-evidenced, least-used ideas in course-evaluation methodology. The only real prerequisite is a platform that can do the sampling and the statistics for you.
Frequently asked questions
What is a planned missing-data design in surveys?
A planned missing-data design deliberately gives each respondent a random subset of the full item set, then reconstructs complete-data estimates statistically. The best-known version, the three-form design (a form of matrix sampling), lets you field roughly a third more questions than any single respondent answers while reducing each person's burden. Because the missingness is created at random, it is missing completely at random (MCAR) and easy to handle.
How does the three-form design work?
You split items into a core block X that everyone answers and three optional blocks A, B, and C. Three forms are created — X+A+B, X+A+C, and X+B+C — and students are randomly assigned one. Each respondent sees three-quarters of the items; every optional block is seen by two-thirds of students. Full-information maximum likelihood or multiple imputation then recovers unbiased estimates across all items.
Is planned missingness statistically valid for course evaluation?
Yes, when analysed correctly. Because the researcher creates the missingness at random, the data are MCAR — the assumption standard missing-data methods are designed for. The trade-off is slightly wider confidence intervals for the optional items, which power tables let you plan for in advance. The core benchmarked block, answered by everyone, loses no precision.
How is this different from just shortening the survey?
Shortening deletes questions, so you permanently lose that information and must decide in advance which items matter. Planned missingness keeps the full instrument but gives each student only part of it, reconstructing the complete picture across the cohort. You reduce per-student burden without giving up content or actionable detail.
Does planned missingness fix non-response bias?
No — it addresses a different problem. Non-response bias comes from which students choose not to respond at all, which is often missing not at random and needs weighting or other corrections. Planned missingness controls how much each responding student is asked, converting that into benign MCAR missingness. The two approaches are complementary.
Can adaptive AI interviews achieve the same burden reduction?
Largely, yes, by a different route. An adaptive conversational interview probes where a student has something to say and skips what is irrelevant to them, reducing burden much as matrix sampling does. It is not a formal three-form design, so strict statistical reconstruction across a fixed instrument still requires FIML or imputation in the analysis layer, but it achieves the same goal of asking each respondent less while learning more overall.