New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
analysis-reporting9 min read

Are Your Respondents Representative? Using Early-vs-Late Wave Analysis to Estimate Nonresponse Bias

A low response rate is not automatically biased — what matters is whether respondents differ from non-respondents. Armstrong and Overton (1977) gave us a cheap diagnostic: compare early and late responders. Here is how to use wave analysis on course-evaluation data and where it breaks down.

Koji Education Team

Product

Answer (BLUF): A 35% response rate is not a problem in itself — the problem is nonresponse bias, the risk that the students who answered differ systematically from those who did not. Armstrong and Overton (1977) showed you can estimate that risk cheaply: treat late responders (those who reply only after reminders) as a proxy for non-respondents, and compare them to early responders. If early and late groups give similar ratings, that is reassuring; if they diverge, you have evidence of likely bias and its direction. Wave analysis turns a bare response-rate number into an actual representativeness check — but it rests on the "continuum of resistance" assumption and cannot, on its own, prove your sample is unbiased.

What the research says

Quality teams routinely fixate on the response-rate percentage, as if 50% were automatically trustworthy and 30% automatically junk. Survey methodology says this is the wrong question. What threatens validity is not low response per se but nonresponse bias — a correlation between the propensity to respond and the thing you are measuring. A 90% response rate can be biased; a 30% rate can be representative. You have to look.

J. Scott Armstrong and Terry Overton (1977), in Estimating Nonresponse Bias in Mail Surveys (Journal of Marketing Research), provided the most widely used cheap diagnostic. Their extrapolation method rests on the continuum-of-resistance idea: respondents who require more prompting to answer are, in relevant respects, more like the people who never answered at all. So if you split your achieved sample into those who responded early (say, the first quartile or the wave before any reminder) and those who responded late (the last quartile, after follow-ups), the late group approximates the non-respondents. Comparing the two waves on your key measures estimates both whether bias is present and in which direction it runs. Analysing published mail-survey data, Armstrong and Overton found that extrapolation gave valid predictions of the direction of nonresponse bias and improved estimates of its magnitude relative to ignoring it — while being honest that the method is better at direction than at precise magnitude.

The approach was operationalised for the social sciences by Lindner, Murphy and Briers (2001) in the Journal of Agricultural Education, whose widely cited guidance recommended comparing early to late respondents (for example, the first 50% against the last 50%, or those responding before versus after a reminder) as a standard control for nonresponse, treating late respondents as the proxy for non-respondents when no external data on non-respondents exist.

That the concern is real for course evaluation specifically is shown by Adams and Umbach (2012) in Research in Higher Education. Studying roughly 135,000 online teaching evaluations from over 22,000 undergraduates, they found participation was systematically driven by salience and survey fatigue — students were more likely to respond when the course felt salient and less likely when over-surveyed. Because these drivers correlate with engagement, the responding sample is not a random slice of the class: the disengaged and over-surveyed are underrepresented. That is precisely the situation in which early-vs-late wave analysis earns its keep.

Why it matters for course evaluation in practice

  • It reframes the response-rate conversation. Instead of arguing whether 40% is "enough" (see how many responses you need), wave analysis asks the sharper question: do the late, reluctant responders rate the course differently from the eager early ones? That is a direct, data-driven test of representativeness.
  • It is nearly free. Every online evaluation system timestamps submissions. Splitting responses into early and late waves and comparing means on the key items costs nothing extra and needs no new data collection.
  • It catches directional bias before high-stakes use. If late responders are systematically harsher, your headline mean is likely flattering the course (because the most critical, least-engaged students are underrepresented) — important to know before the score feeds promotion or programme review.
  • It complements known-population checks. Where you hold roster data (grades, demographics), comparing the respondent profile to the full cohort — the Adams and Umbach validation style — is even stronger. Wave analysis is the fallback when you have no data on the non-respondents themselves.

Limitations & honest caveats

Wave analysis is a diagnostic, not a guarantee, and a methodologically literate reader should treat it cautiously.

  • The continuum-of-resistance assumption can be wrong. Late responders are only a proxy for non-respondents. Some non-respondents differ from everyone who answered, in ways no amount of prompting would surface. Armstrong and Overton were explicit that the method predicts direction better than magnitude.
  • Similar waves do not prove no bias. If early and late responders look alike, nonresponse bias could still exist on a dimension orthogonal to resistance. Convergence is reassuring evidence, not proof of representativeness.
  • Confounds in "lateness." Responding late may track procrastination, course load, or engagement for reasons unrelated to how a student evaluates teaching, muddying the proxy.
  • You need timestamps and a big-enough late wave. Small classes may not yield enough late responders to compare meaningfully, and the technique inherits all the small-sample reliability problems of course evaluation.
  • It diagnoses, it does not fix. Detecting likely bias does not correct your estimate; correction needs weighting or modelling, which brings its own assumptions. The right move is usually to report the diagnostic alongside the score, not to silently adjust.

The defensible practice is to run wave analysis as a routine flag, report it transparently, and pair it with roster-based representativeness checks wherever the data exist — while continuing to attack the root cause by raising response rates (see what actually raises response rates).

How Koji incorporates this

Koji is designed so that representativeness is something you can see, not assume.

  • Built-in response timestamps and wave comparison. Because every Koji interview is timestamped, early-vs-late wave analysis is available out of the box: the platform can compare students who responded before reminders against those who responded after, surfacing divergence on key items as a nonresponse-bias flag rather than burying it.
  • Bias-aware reporting, not just a mean. Koji reports the achieved distribution and can highlight when early and late responders diverge, so an evaluation committee receives a representativeness signal alongside the headline number — designed to mitigate over-confident interpretation of a score built on a skewed sample.
  • Roster-based representativeness where data permit. Where an institution supplies cohort information, Koji supports comparing the respondent profile against the full class — the stronger, Adams-and-Umbach-style check — to corroborate or challenge what the wave analysis suggests.
  • Attacking the root cause. The most reliable defence against nonresponse bias is fewer non-respondents. Koji''s conversational, mobile-friendly format and reminder workflows are built to lift completion and reduce the over-surveying that Adams and Umbach (2012) identified as a driver of dropout, shrinking the gap that extrapolation has to span.

The same diagnostics power Koji''s core research platform at koji.so, where customer and employee studies face identical nonresponse questions: a satisfaction score is only as trustworthy as the representativeness of the people who chose to answer.

How to run it: a worked example

Wave analysis needs nothing more than your existing response timestamps. A practical recipe:

  1. Order responses by submission time and define waves. A common split is the first quartile (early) versus the last quartile (late); alternatively, define "late" as everyone who responded only after the first reminder.
  2. Compare the waves on your key outcome items — overall rating, clarity, workload — using simple mean differences, plus any available demographics.
  3. Read the direction. Suppose early responders give an overall rating of 4.3 and late responders 3.8. Under the continuum-of-resistance logic, the unseen non-respondents likely sit at or beyond the late group, so your true cohort mean is probably below the 4.3 the eager students produced — the headline flatters the course.
  4. Report the flag; do not silently adjust. State the early-late gap alongside the score so the committee can weight it. Move to formal weighting only if you hold external population data to justify the model.

Two cautions: make sure the late wave has enough responses to compare (a handful tells you nothing), and remember that a null result — early and late looking alike — is reassuring but not conclusive, because non-respondents can differ on dimensions unrelated to how quickly someone replies.

Related Resources

References

  1. Armstrong, J. S., & Overton, T. S. (1977). Estimating Nonresponse Bias in Mail Surveys. Journal of Marketing Research, 14(3), 396–402. https://doi.org/10.1177/002224377701400320
  2. Lindner, J. R., Murphy, T. H., & Briers, G. E. (2001). Handling Nonresponse in Social Science Research. Journal of Agricultural Education, 42(4), 43–53. https://doi.org/10.5032/jae.2001.04043
  3. Adams, M. J. D., & Umbach, P. D. (2012). Nonresponse and Online Student Evaluations of Teaching: Understanding the Influence of Salience, Fatigue, and Academic Environments. Research in Higher Education, 53(5), 576–591. https://doi.org/10.1007/s11162-011-9240-5

Related articles