Your 200 Responses Are Worth 120: Design Effects, Effective Sample Size, and the Finite-Population Correction
Clustering and weighting inflate variance so the effective sample size is below the headcount; a near-census of a small class earns a finite-population correction. Kish's design effect makes precision honest.
Koji Education Team
Product
In brief
A course-evaluation report that says "based on 200 responses" is quietly assuming those 200 responses carry the precision of 200 independent draws. They usually do not. When responses cluster within sections or tutors, and when you weight them to correct for uneven response rates, the variance of your estimates inflates — so the effective sample size is smaller than the headcount, often by a third or more. Pulling the other way, when you have collected answers from most of a small class, the finite-population correction means you actually know more than a simple-random-sampling formula assumes. The design effect, effective sample size, and finite-population correction are the three ideas that convert a raw response count into an honest statement of precision.
What the research says
Kish (1965), in Survey Sampling, named the design effect (DEFF): the ratio of the variance of an estimate under the actual, complex sampling design to the variance you would get from simple random sampling with the same number of respondents. A DEFF of 1 means the design costs nothing; a DEFF of 1.6 means your estimate is as precise as a simple random sample 1.6 times smaller — so 200 clustered, weighted responses might carry the precision of 125. Two mechanisms drive DEFF above 1. Clustering: when students within a section resemble each other, each additional response from the same cluster adds less new information; Kish's clustering formula is DEFF ≈ 1 + (m − 1)ρ, where m is the average cluster size and ρ the intraclass correlation. Unequal weighting: when you weight responses to compensate for differential non-response, the weights themselves inflate variance; Kish's weighting formula is DEFF ≈ 1 + CV²(w), the squared coefficient of variation of the weights. Kish (1992), in "Weighting for unequal Pi," set out the several reasons — non-response, frame problems, post-stratification — that produce those unequal weights in the first place.
Gabler, Häder and Lahiri (1999), in Survey Methodology, gave Kish's combined formula a model-based justification, deriving a general expression that captures clustering and weight variation simultaneously and showing that Kish's DEFF is a somewhat conservative — that is, an over-estimating — approximation under stated assumptions. The effective sample size follows directly: n_eff = n / DEFF, the number of independent observations that would give the same precision as your actual design. Working in the opposite direction, standard sampling theory (see, for example, Lohr's Sampling: Design and Analysis) supplies the finite-population correction (FPC), the factor √((N − n)/(N − 1)) that shrinks the standard error when the sample n is a large fraction of a finite population N. When you evaluate 45 of a 50-student class, the FPC is substantial and your interval is genuinely narrower than a naive formula would report.
Why it matters for course evaluation in practice
Evaluation data are almost never a simple random sample, yet almost every report treats them as one. Three practical consequences follow.
Precision is overstated for clustered, weighted data. A programme aggregates 200 responses across twelve tutorial groups and reports a confidence interval computed from n = 200. If students within a tutor's group rate similarly — the same tutor, room, and timetable slot — the intraclass correlation is non-zero and the true interval is wider. Reporting the naive interval manufactures false confidence, and small differences between programmes or years look significant when they are within noise. Correcting for the design effect widens the interval to what the data actually support.
Weighting is not free. Offices increasingly weight responses to counter the well-documented non-response bias in evaluations. That is the right instinct, but each round of weighting raises CV²(w) and therefore DEFF, trading bias reduction for variance. A defensible report states the effective sample size after weighting, not the raw count, so the reader sees the true precision cost of the adjustment.
Precision is understated for a near-census of a small class. The opposite error is just as common. A seminar of 30 returns 26 evaluations, and a bootstrap or textbook interval treats those 26 as a sample from an infinite population. With the finite-population correction, you have observed 87 percent of the class and know its mean far more precisely than the naive interval claims. Ignoring the FPC makes small-class results look flimsier than they are — which matters when those results carry weight in a tenure or programme-review file.
Limitations and honest caveats
The design effect is a summary, not a model, and it can mislead if used carelessly. DEFF and effective sample size are estimate-specific: the intraclass correlation for "overall satisfaction" may differ from that for "assessment clarity," so a single DEFF applied to every item is an approximation. Kish's weighting formula, DEFF ≈ 1 + CV²(w), assumes the weights are roughly uncorrelated with the outcome; when weights and responses are related, it can over- or under-state the true inflation, and Gabler and colleagues' more general derivation is needed. Estimating ρ reliably requires enough clusters — a handful of tutorial groups gives a noisy ρ and therefore a noisy DEFF. The finite-population correction assumes the responders are a random sample of the finite population; if the 26 who replied differ systematically from the 4 who did not, the FPC narrows an interval around a possibly biased point, so it must never be applied without first considering non-response bias. In short, these tools correct the precision statement, not the validity of the underlying estimate — a beautifully corrected interval around a biased mean is still centred in the wrong place.
How Koji incorporates this
Koji is built to treat evaluation data as the clustered, unequally weighted, finite-population data it actually is, rather than as a convenient simple random sample. Because Koji knows the structure of a collection — which responses belong to which section, tutor, or cohort — it can express precision in terms of effective sample size rather than a raw response count, so a programme director sees that 200 responses across twelve groups carry the precision of, say, 130 independent ones. Where Koji applies weighting to counter non-response, it can surface the variance cost of that weighting through the design effect, making the bias–variance trade-off visible instead of hidden. For the small, near-complete classes that dominate seminar and postgraduate teaching, Koji's reporting reflects the finite-population correction, so a near-census is credited with the precision it has earned rather than being penalised by an infinite-population assumption. And because a corrected interval is only as trustworthy as the estimate it surrounds, Koji pairs these precision statements with its non-response diagnostics and its AI-moderated qualitative evidence, so a narrow interval is never mistaken for a valid one. The same survey-statistical rigour underpins Koji's core research platform at koji.so, where product and customer-research samples are likewise clustered by account or segment and weighted for representativeness, and where honest effective-sample-size reporting keeps teams from over-reading a noisy result.
Frequently asked questions
What is a design effect in one sentence?
It is how many times larger the variance of your estimate is because of clustering and weighting, compared with a simple random sample of the same size — so a DEFF of 1.5 means your data are as precise as an SRS one-and-a-half times smaller.
How is effective sample size different from the response rate?
The response rate is the fraction of invited students who answered; the effective sample size is how many independent observations your responses are worth after accounting for clustering and weighting. You can have a high response rate and still a low effective sample size if responses cluster heavily.
Isn't this the same as needing enough responses for reliability?
No. Our note on how many responses you need is about reaching a reliability or power target. The design effect is about how much each response is worth once the design is non-simple — the two combine, because your target n must be inflated by the DEFF.
When does the finite-population correction actually help?
When your responders are a large fraction of a finite class — roughly a fifth or more. For a seminar of 30 with 26 replies the correction is large and narrows your interval; for one section of a 2,000-student cohort it is negligible.
Does weighting for non-response make my estimates better or worse?
Both. Weighting reduces bias from uneven response but raises variance through the design effect. The right report shows the effective sample size after weighting so the reader sees the trade-off rather than only the bias-reduction benefit.
How does this relate to multilevel models?
They are two views of the same nested structure. A multilevel model explicitly models the clustering; the design effect is the compact precision summary of that same clustering. If you are already fitting a multilevel model you have effectively accounted for the DEFF; if you are only reporting means, the DEFF is how you keep the intervals honest.
References
- Kish, L. (1965). Survey Sampling. New York: John Wiley & Sons.
- Kish, L. (1992). Weighting for unequal Pi. Journal of Official Statistics, 8(2), 183–200.
- Gabler, S., Häder, S., & Lahiri, P. (1999). A model based justification of Kish's formula for design effects for weighting and clustering. Survey Methodology, 25(1), 105–106.
- Lohr, S. L. (2019). Sampling: Design and Analysis (3rd ed.). Boca Raton: Chapman & Hall/CRC.
Related resources
- How Many Responses Do You Need for a Reliable Course Evaluation?
- Students Are Nested in Courses: Why Multilevel Models Beat Raw Averages
- Generalizability Theory and the Reliability of Student Ratings: Why One Class Is Not Enough
- Can You Report a Class Mean? ICC(1), ICC(2), and r_wg for Aggregating Student Ratings
- Can You Weight Your Way Out of a Low Response Rate? Post-Stratification and Nonresponse Adjustment
- Fair Confidence Intervals for Small Classes: The Bootstrap for Course-Evaluation Reporting
Related articles
How Many Responses Do You Need for a Reliable Course Evaluation?
Reliability of a class-average rating rises steeply with the number of students who respond. Generalizability theory — from Marsh''s syntheses to Dzakadzie (2026) — shows raters matter more than items, and a class of 12 cannot be evaluated like a class of 120.
Generalizability Theory and the Reliability of Student Ratings: Why One Class Is Not Enough
Reliability is not one number. Generalizability theory (Gillmore, Kane & Naccarato 1978; Marsh 1984) decomposes the variance in student ratings into student, teacher, course and occasion components — and shows that a single class can be reliable for the course yet a poor estimate of the teacher. What that means for fair evaluation.
Can You Report a Class Mean? ICC(1), ICC(2), and r_wg for Aggregating Student Ratings
Before you average student ratings into a class score, three organisational-psychology indices decide whether you can: r_wg (within-class agreement), ICC(1) (how much variance is between classes), and ICC(2) (the reliability of the class mean).
Students Are Nested in Courses: Why Multilevel Models Beat Raw Averages for Course Evaluation
Course-evaluation data has a nested structure — students within sections within instructors — and a flat average ignores it. This guide explains how multilevel (hierarchical linear) models partition variance, why ignoring clustering understates uncertainty, and what it means for fair reporting.