New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Sector trends10 min read

Are Students Using AI to Write Course Evaluation Feedback? The New Data-Quality Threat

Generative AI does not only help institutions analyse feedback — it lets students generate it. New evidence on AI-contaminated open-text responses raises an uncomfortable question for course evaluation, and points to where instrument design has to go next.

Koji for Education

Research & Editorial Team · June 10, 2026

Bottom line up front: There is now direct evidence that a substantial share of respondents use generative AI to write open-ended survey answers — in one 2024 study, 34.3% of online participants admitted to it. Course evaluation is not immune. AI-generated feedback is fluent, plausible, and passes the bot checks that legacy survey tools rely on, which means the static open-comment box is becoming a less trustworthy source of evidence. The response is not to abandon open feedback but to redesign how it is collected and verified.

A threat the sector has not priced in

For two years the conversation about AI and course evaluation has run in one direction: institutions using large language models to analyse open-text comments faster. That use is real and valuable — a 2024 study in BMC Medical Education found ChatGPT could surface themes from course-evaluation comments with reasonable agreement against human coders. But the same technology runs in the other direction too: students can use AI to generate the feedback in the first place. The analytical conversation has almost entirely ignored the data-integrity one.

The evidence that this is already happening comes from survey methodology, where the stakes are identical. A 2025 paper in PNAS, "The potential existential threat of large language models to online survey research", reported that in a 2024 sample of online research participants, 34.3% admitted to using AI to help answer open-ended questions. The authors warn that LLMs could turn survey fraud "from a labor-intensive, low-margin cottage industry into a potentially lucrative and scalable" operation. The data-quality firm Research Defender, cited in that work, estimates that roughly 31% of raw survey responses are problematic or fraudulent across all causes. Separate work presented at NeurIPS 2024 showed how readily LLMs produce coherent, human-passing survey answers.

Why detection is failing

The instinct is to filter AI-written responses out. The PNAS authors are blunt that this is no longer reliable: autonomous AI agents "operating from a simple prompt can evade current detection methods," and the heuristic that logically coherent responses come from humans is now "untenable." Legacy course-evaluation platforms inherit exactly this weakness. Their defences — reCAPTCHA, attention checks, timing thresholds, gibberish detectors — were built to catch bots and inattentive clickers, not a student pasting "write three sentences of polite but critical feedback about a statistics course" into a chatbot. AI text-detectors, meanwhile, carry false-positive rates high enough to be unsafe for any consequential use, and they disproportionately flag non-native English writers — so "screen the comments for AI" risks penalising exactly the students whose voice institutions most want to protect.

Why students would do it — and why it matters

Some AI use here is benign: a non-native speaker drafting a comment in their language and translating it, or a student tidying grammar. But the methodological problem is the same regardless of motive. When a comment is AI-generated, it stops being a faithful report of that student's experience and becomes a plausible-sounding average of how courses are generally discussed. It launders away the specific, surprising, course-changing detail — the lab that always ran out of time, the feedback that arrived after the resit deadline — that makes open text worth collecting at all. Worse, survey fatigue is a documented driver of low-effort responding, and an end-of-term comment box that students experience as pointless box-ticking is precisely the context where reaching for a chatbot is most tempting.

But isn't this just the old satisficing problem with a new name?

This is the strongest objection, and it is half right. Students have always written low-effort, copy-paste, or strategically vague evaluation comments — the satisficing and straightlining literature documents it well. So one could argue AI changes nothing fundamental. Two things make the new threat genuinely different. First, scale and quality: previous low-effort responses were detectable because they were short, repetitive, or empty; AI-generated ones are long, varied, and superficially excellent, defeating both human readers and automated quality flags. Second, directionality: satisficing produces obviously thin data, whereas AI produces confidently wrong data that looks like signal. A comment box full of articulate, on-topic, entirely synthetic feedback is more dangerous to a quality-assurance process than one full of "N/A," because it will be believed and acted upon. The old problem degraded data visibly; the new one degrades it invisibly.

What actually helps: design for authenticity, not detection

If detection cannot be the front line, instrument design has to be. Several principles follow from the evidence. Ask for specifics that are costly to fabricate — concrete moments, examples, and consequences rather than general impressions, which is precisely what a generic AI prompt produces. Use interactive, adaptive follow-ups rather than a single static box, because a real follow-up question ("you mentioned the labs felt rushed — which lab, and what would have helped?") is far harder to satisfy with pre-generated text than an open prompt answered offline. Collect feedback in a single contextual session rather than a link a student opens in a separate tab beside a chatbot. And measure engagement quality, not just completion, so that thin or evasive responses are visible.

How Koji fits

Koji for Education is built around exactly the design principles the AI-contamination evidence points to. Its AI-moderated conversational interviews replace the static comment box with an adaptive exchange: when a student raises an issue, Koji probes for the specific example and its consequence, which both yields richer evidence and makes wholesale AI fabrication far less practical than answering a one-shot prompt offline. Because the interview happens in a single guided session rather than a take-home link, it is harder to outsource end to end. Koji's automatic quality scoring flags thin, evasive, or non-substantive responses for what they are, so quality-assurance staff can weight the evidence accordingly — without resorting to unreliable, discriminatory AI-text detectors. And its thematic analysis works on the substance that survives this process, surfacing genuine patterns rather than synthetic ones. To be precise about the claim: this reduces the practicality and impact of AI-generated feedback and surfaces low-quality responses — it does not guarantee every answer is human-authored, and no honest vendor should say otherwise. (The same conversational interview engine powers customer and user research on the main Koji platform, where AI-contaminated survey panels are an even more acute problem.)

What this means for benchmarking and trend data

The contamination problem is most corrosive where institutions least expect it: longitudinal and cross-programme comparison. If the share of AI-assisted responses grows year on year — and every indication is that it will — then a programme's open-text feedback becomes progressively less comparable to its own history, even when nothing about the teaching changed. A rise in articulate, on-theme comments could reflect better courses, or simply wider chatbot use. Trend lines built on open-text sentiment are quietly losing their baseline.

There is also an equity dimension that cuts directly against naive screening. Because AI-text detectors over-flag non-native English writers, an institution that responds to the threat by filtering suspected-AI comments risks systematically discarding the feedback of international and multilingual students — the very voices already most marginal in evaluation data. The contamination problem and the inclusion problem therefore have to be solved together: any response that improves data integrity by silencing legitimate non-native voices has made the evaluation worse, not better. That is a strong argument for designing authenticity into collection rather than policing it after the fact, and for capturing feedback in a respondent's stronger language so that fluent, AI-polished text is not the only route to being heard.

The static open-comment box was already a weak instrument. Generative AI has quietly made it weaker — not by filling it with nonsense, but by filling it with fluent, confident, untraceable plausibility. Institutions that take the evidence seriously will stop trying to detect their way out and start designing feedback that is worth trusting.

See how conversational, adaptive course evaluation holds up where static surveys don't — explore Koji for Education.