Can You Trust Your Course Evaluation Data? Bots, Duplicates and Fraud in Online Student Feedback
The integrity threat in student evaluation has shifted from who stayed silent to whether the responses you received are authentic at all. A methodology deep-dive on bots, duplicate submissions and careless responding — and how to defend your data.
Koji Education Team
Product ·
Bottom line up front: The integrity of course evaluation used to be a question of non-response — who stayed silent. In the online, open-link era it is increasingly a question of authenticity: whether the responses you did receive came from the enrolled students they claim to, submitted once, in good faith. Bots, duplicate submissions, ballot-stuffing and careless "insufficient-effort" responding are well-documented threats across online survey research, yet almost no student-evaluation system audits for them. If you cannot vouch for who — or what — produced a datum, you cannot defend any decision built on it.
The threat model nobody audits
Quality-assurance offices spend enormous energy on response rates and non-response bias, and rightly so. But a datum can arrive and still be worthless, or worse, actively misleading. Three failure modes are largely invisible in standard student-evaluation of teaching (SET) reporting:
- Machine-generated responses. Automated scripts and, increasingly, large language models can complete an open-link survey in seconds, generating plausible Likert patterns and even fluent open-text comments.
- Duplicate and ballot-stuffed submissions. A single motivated actor — a disgruntled student, or an instructor coaching a cohort — can submit multiple times when links are shared and identity is not bound to a single session.
- Careless or insufficient-effort responding. Real students who straightline, speed through, or answer without reading. This is not fraud, but it degrades data in the same way.
None of these are exotic. They are the daily reality of applied survey methodology. What is unusual is that course evaluation, uniquely among high-consequence survey instruments, rarely screens for any of them.
What survey methodology already knows
The evidence base outside higher education is sobering. Industry analyses by CloudResearch estimate that a large share of responses to open online surveys are fraudulent or unusable, with organised human "click-farm" networks — not just bots — driving much of the bad data. The Association for Psychological Science has argued bluntly that the biggest threat to online data collection is humans, not bots: coordinated, incentivised humans who defeat naive screening.
On the machine side, Storozuk and colleagues' 2020 paper Got Bots? Practical Recommendations to Protect Online Survey Data from Bot Attacks documented how quickly an open survey link can be overrun and set out concrete countermeasures. And a substantial methodological literature — synthesised in the Annual Review of Psychology on careless responding — shows that even attentive, well-meaning respondents produce low-quality data at non-trivial rates, and that single attention checks are insufficient on their own.
The uncomfortable implication: if you run an open-link Likert survey and do nothing to verify authenticity, you are almost certainly ingesting some contaminated data. The only open question is how much, and whether it is randomly or systematically distributed — the latter being far more dangerous for the department- and instructor-level comparisons SET is used to justify.
Why course evaluation is unusually exposed
Course evaluation combines several risk factors that most commercial surveys do not:
- Stakes without safeguards. For adjunct and probationary faculty, evaluation scores can drive contract renewal and promotion. That creates a motive to game — in either direction — while the low-stakes framing means few institutions invest in fraud controls proportionate to the actual consequences.
- Low base rates and small n. Many modules have a handful of respondents. A single stuffed or bot response moves the mean visibly. Statistical noise is already the enemy of small-class reporting; authenticity failures compound it.
- Open or weakly-authenticated links. Where evaluation links are shared by email or posted in a learning-management system, the binding between a response and a unique enrolled student is often loose.
- Emotional intensity. Evaluation invites strong feeling, which is exactly the condition under which a motivated actor is most likely to submit repeatedly, and under which negativity bias already distorts the open-text record.
The generative-AI escalation
Two years ago, a fraudulent course-evaluation comment looked obviously fake. Today a language model can produce feedback that is coherent, specific and tonally appropriate — indistinguishable, on the page, from a thoughtful student. This cuts two ways. First, it lowers the cost of manufacturing fake open-text feedback to essentially zero. Second, it corrupts the very signal — rich, specific qualitative comment — that institutions have started to trust more than the numbers. If committees are now reading open text as the "authentic voice" of students, that voice is precisely what is now cheapest to fake.
But doesn't anonymity plus low stakes mean nobody bothers?
This is the strongest objection, and it deserves a direct answer. The reasoning goes: course evaluations are anonymous and mostly consequence-free for students, so why would anyone commit fraud, and why worry about bots that have nothing to gain?
Three reasons the objection under-weights the risk. First, the incentive does not have to sit with the student. Instructors and, occasionally, programme leads have a documented interest in the numbers, and the fraud literature repeatedly finds that opportunity plus low detection risk is sufficient — motive need not be large. Second, bots do not need a rational payoff; open links get scraped and hit by automated traffic indiscriminately, as the Got Bots? work shows. Third, and most importantly, the low-stakes framing is a reason detection is absent, not a reason fraud is absent — it produces exactly the unguarded environment in which contamination goes unmeasured. "We have never found fraud" is not evidence of clean data when you have never looked.
What defensible practice looks like
Authenticity is a solvable problem, but only if it is treated as a first-class part of instrument design rather than an afterthought:
- Bind each response to a single authenticated session so duplicate and ballot-stuffed submissions are structurally prevented, not detected after the fact.
- Capture behavioural signals — response timing, coherence, internal consistency — and use them to flag insufficient-effort and machine-generated responses, following the multi-signal approach the methodological literature recommends over single attention checks.
- Score data quality explicitly and report it alongside results, so a committee knows how much of a module's evaluation rests on high-confidence responses.
- Prefer interaction over a static form. A one-shot Likert grid is trivial to auto-complete; a genuine back-and-forth that probes and follows up is far harder to fake convincingly at scale.
How Koji addresses authenticity
Koji for Education was built around a conversational, AI-moderated interview rather than a static survey form — and that architecture changes the authenticity equation. Each evaluation is a single, authenticated session, which structurally blocks the duplicate and ballot-stuffing routes that open links leave wide open. Because the interaction is a real dialogue — the AI moderator asks follow-up probes, seeks specifics, and reacts to what the student actually says — it is markedly harder for a bot or a bored respondent to maintain a coherent, specific thread than to tick a grid of scales. Koji applies automatic quality scoring to every response, surfacing low-effort or incoherent submissions rather than silently averaging them into a departmental mean, and its thematic analysis works on comments that have already passed those quality checks. None of this eliminates fraud — no system honestly can — but it moves authenticity from an unaudited blind spot to an explicit, reportable property of the data.
The same interview engine powers general user and customer research on the main Koji platform, where response authenticity is an equally hard problem — a reminder that student evaluation is not a special case, but a high-stakes instance of a universal data-quality challenge.
The takeaway
Non-response bias taught quality-assurance offices to ask "who is missing?" The online era demands a second, harder question: "is what I have real?" Until an evaluation system can answer that — with authentication, behavioural screening and transparent quality scoring — the confident decimal points on a course-evaluation dashboard are resting on an assumption no one has tested.
Ready to evaluate on data you can actually vouch for? Explore Koji for Education.