Some of Your Course Evaluations Were Answered in Ten Seconds: Response Latency as a Data-Quality Signal
Response times ('paradata') are a free, objective quality signal in online course evaluations. Evidence from Zhang & Conrad and others shows speeders straightline more — and how to use latency to screen data without deleting honest fast responses.
Koji Education Team
Product
Answer box (BLUF)
The time a student spends answering an online course evaluation — captured automatically as paradata — is one of the cheapest and most objective data-quality signals you have. Zhang and Conrad (2014) showed that speeding (answering unrealistically fast) is systematically associated with straightlining (giving the same answer down a whole grid), and both appear to share a common root: minimising effort. Used carefully, response latency lets you flag likely low-effort responses without discarding honest fast readers. Used carelessly — a blunt "delete anything under N seconds" rule — it introduces its own bias. The defensible approach is a relative threshold (e.g. faster than roughly one-third of the median completion time) combined with a second corroborating indicator, not latency alone.
What the research says
Speeding predicts straightlining. Zhang and Conrad (2014), in Survey Research Methods, analysed web-survey response times and found that respondents who answer very fast are markedly more likely to straightline — to select the same scale point for every item in a matrix. Speeding and straightlining, they argued, are two behavioural expressions of the same satisficing tendency: the respondent takes the cognitive shortcut rather than doing the full "comprehend → retrieve → judge → respond" work that a considered answer requires. Younger respondents (the bulk of a student population) were more prone to speed, which makes the signal especially relevant to course evaluations.
Speeders can distort estimates — but not always by much. Greszki, Meyer and Schoen (2015) examined how excluding speeders changes survey estimates. Their finding is nuanced and important for QA: applying a relative speeding threshold (defined against the median completion time) identified a meaningful group of low-effort respondents, but removing them often changed substantive estimates only modestly. The lesson is not "speeders ruin everything," but "speeders are a real, identifiable subgroup you can study and screen, and the impact should be measured rather than assumed."
Latency as paradata. Callegaro (2013) and Kreuter (2013) situate response times within the broader concept of paradata — data about the process of answering (timestamps, clicks, device, page revisits) generated automatically as a by-product of an online instrument. Paradata is attractive precisely because it is non-reactive: unlike an attention-check question, it does not interrupt or annoy the respondent, and it is collected for every response at zero marginal cost. Response latency is the most widely used paradata quality indicator.
Feedback can reduce speeding. Conrad, Tourangeau, Couper and Zhang (2017) found that giving respondents immediate feedback when they answer implausibly fast ("You seem to be answering quickly — please take your time") reduced subsequent speeding and improved response quality — evidence that speeding is partly a corrigible behaviour, not a fixed trait.
Why it matters for course evaluation in practice
Course evaluations are a near-perfect habitat for satisficing: they are low-stakes for the student, often mandatory or nagged into existence, frequently completed on a phone between classes, and dominated by long Likert grids. That is exactly the design that Zhang and Conrad associate with speeding and straightlining. Three practical implications follow:
- A mean of 4.3 can be partly an artefact of straightliners. If a nontrivial fraction of your responses are "all 4s" produced in twelve seconds, your central-tendency statistics are contaminated by non-differentiated data. This is invisible in the mean but detectable in the paradata.
- Latency screening is fairer than gut feeling. Reviewers already discount responses they suspect are careless. Paradata replaces that subjective, potentially biased judgement with an objective, pre-registered rule applied equally to everyone.
- You can measure the damage before acting. Following Greszki et al., the right move is to compute your key statistics with and without flagged speeders and report the difference. If it is negligible, you keep the data and note the robustness; if it is large, you have found a genuine quality problem worth addressing at the collection stage.
Limitations and honest caveats
A careful methodologist will raise several objections, and they are right to:
- Fast is not the same as careless. A student who genuinely read the syllabus questions before, has strong settled opinions, and reads quickly can legitimately finish fast. A naïve absolute cutoff ("under 30 seconds = delete") will disproportionately discard fluent readers and, because reading speed correlates with language proficiency, can systematically remove non-native speakers or particular cohorts — introducing exactly the bias you were trying to remove.
- Thresholds are arbitrary and must be relative. The literature favours thresholds defined against the median completion time for that specific instrument (e.g. faster than one-third of the median), not fixed seconds, because instruments differ enormously in length.
- Latency alone is a weak flag. Best practice is convergent evidence: flag a response only when short latency co-occurs with another indicator (straightlining, failed attention item, impossible response patterns). Zhang and Conrad's whole point is that these indicators validate each other.
- Device and interruption confound timing. A mobile respondent, or one who paused to answer a message and returned, produces long or bimodal times that are not "effort." Page-level timestamps and revisit paradata help disambiguate, but no timing rule is clean.
- Screening is not deletion. The goal is transparent sensitivity analysis, not silently dropping inconvenient data — which, unreported, is itself a research-integrity problem.
How Koji incorporates this
Koji for Education treats data quality as something to be engineered at collection and made visible at reporting, rather than patched by deleting rows afterwards:
- The conversational format is a structural antidote to straightlining. Because Koji's AI-moderated interviews ask one thing at a time and follow up on the specific answer, there is no long matrix grid to run a finger down — the format that Zhang and Conrad tie most tightly to speeding is simply not present. A respondent cannot "straightline" a conversation that reacts to what they just said.
- Adaptive probing raises the cost of a non-answer. When a student gives a thin or contradictory reply, the moderator asks a clarifying follow-up. This is the collection-time analogue of Conrad et al.'s "immediate feedback" finding — gentle, in-context prompting that nudges a satisficing respondent back toward genuine engagement.
- Quality scoring uses response substance, not just speed. Koji's quality signals draw on the richness and internal consistency of open-text answers, so a fast-but-substantive interview is not penalised while an empty, low-effort one is flagged — directly addressing the "fast ≠ careless" caveat.
- Bias-aware reporting supports sensitivity analysis. Rather than silently dropping flagged responses, Koji is designed to let QA teams see results with and without low-quality responses, echoing the Greszki et al. principle of measuring the impact of screening before acting on it.
- Structured items still exist where needed (
scale,single_choice,ranking,yes_no), but they are interleaved with probing rather than stacked into the long grids that manufacture straightlining.
The claim is bounded: Koji is designed to mitigate satisficing and to make quality screening transparent — it cannot guarantee every respondent engages fully, and no format eliminates low effort entirely. Teams applying the same engine to customer and product research beyond the classroom can do so through Koji's core platform at koji.so.
Related Resources
- Why Students Click Straight Down the Middle: Satisficing in Course Evaluations
- Some of Your Course-Evaluation Responses Are Careless: Detecting Insufficient-Effort Responding
- How Long Should a Course Evaluation Be? Questionnaire Length, Breakoff, and Answer Quality
- Do Students Completing Evaluations on Their Phones Give Worse Data? The Device-Effects Evidence
- Grid or One Question at a Time? Matrix Formats and Straightlining in Course Evaluations
- Survey Fatigue: Why Over-Surveying Students Quietly Wrecks Your Response Rates
A defensible latency-screening protocol
Turning the evidence into a repeatable, auditable procedure matters more than any single threshold. A protocol a critical reader would accept looks like this:
- Capture, do not act, first. Record page-level timestamps and total completion time as paradata for every response. Collecting is non-reactive and costless; decisions come later.
- Set a relative threshold. Following Greszki, Meyer and Schoen, define "too fast" against the median completion time for that specific instrument — a common choice is faster than one-third of the median — rather than an absolute number of seconds that ignores instrument length.
- Require convergence. Flag a response as likely low-effort only when short latency co-occurs with a second indicator: straightlining down a block, a failed instructional-manipulation check, or impossible patterns (e.g. selecting "not applicable" and then rating the same item). This is the mutual-validation logic Zhang and Conrad established.
- Run a sensitivity analysis, do not silently delete. Compute your key statistics — means, top-box percentages, theme frequencies — both with and without flagged responses, and report the difference. If it is negligible, keep the data and note the robustness; if it is large, you have found a collection-stage problem worth fixing.
- Pre-register the rule. Decide the threshold and the convergence criteria before looking at the results, and apply them identically to every response, so screening cannot become a way to remove inconvenient but honest feedback.
The point of the protocol is fairness and transparency. Paradata replaces a reviewer's private hunch that "this one looks careless" with a rule applied equally to everyone and disclosed openly — which is exactly what a QA process should be able to defend to an external panel.
References
- Zhang, C., & Conrad, F. G. (2014). Speeding in web surveys: The tendency to answer very fast and its association with straightlining. Survey Research Methods, 8(2), 127–135. https://doi.org/10.18148/srm/2014.v8i2.5453
- Greszki, R., Meyer, M., & Schoen, H. (2015). Exploring the effects of removing "too fast" responses and respondents from web surveys. Public Opinion Quarterly, 79(2), 471–503. https://doi.org/10.1093/poq/nfu058
- Conrad, F. G., Tourangeau, R., Couper, M. P., & Zhang, C. (2017). Reducing speeding in web surveys by providing immediate feedback. Survey Research Methods, 11(1), 45–61. https://doi.org/10.18148/srm/2017.v11i1.6304
- Callegaro, M. (2013). Paradata in web surveys. In F. Kreuter (Ed.), Improving Surveys with Paradata: Analytic Uses of Process Information (pp. 261–279). Wiley. https://doi.org/10.1002/9781118596869.ch11
- Kreuter, F. (Ed.). (2013). Improving Surveys with Paradata: Analytic Uses of Process Information. Wiley. https://doi.org/10.1002/9781118596869
Related articles
Survey Fatigue: Why Over-Surveying Students Quietly Wrecks Your Response Rates
Porter, Whitcomb & Weitzer (2004) showed that administering multiple surveys in one year suppresses later response rates. A research-grounded guide to survey fatigue in course evaluation — what causes it, what the evidence shows, and how to design around it.
Why Students Click Straight Down the Middle: Satisficing in Course Evaluations
A large share of students do not answer evaluation questions carefully — they straightline, speed through, and pick the easy option. Barge and Gehlbach (2012) showed this "satisficing" doesn't just add noise; it inflates your reliability and validity statistics, making bad data look good.
How Long Should a Course Evaluation Be? Questionnaire Length, Breakoff, and Answer Quality
What the survey-methodology evidence says about questionnaire length: longer instruments depress participation and degrade answers to later questions, but ruthless shortening is not automatically the answer. A research-grounded guide for designing course evaluations.
Do Students Completing Evaluations on Their Phones Give Worse Data? The Device-Effects Evidence
Most students now answer course evaluations on a smartphone. Does the device degrade the data? The evidence says ratings stay stable across devices, but participation and the length of open-text comments differ — with clear design implications.