New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods11 min read

When Students Use ChatGPT to Write Their Course Feedback: AI-Generated Open-Text and Data Integrity

Generative AI lets anyone produce fluent, coherent open-text survey answers, breaking the old assumption that a coherent response is a human response. What the evidence says about AI-written feedback, why detection is failing, and how to protect the integrity of qualitative course evaluation.

Koji Education Team

Product

Bottom line: Generative AI now produces open-text survey answers that are fluent, on-topic, and internally consistent, which breaks the assumption course evaluations have always relied on: that a coherent comment is a human comment. Recent evidence finds a substantial minority of survey respondents already paste AI-written answers, and that synthetic responses sail through standard attention and quality checks. The defensible response is not a magic detector — they do not work reliably — but evaluation design that makes effortful, AI-assisted faking pointless: authenticated low-stakes collection, adaptive conversational probing, and treating open-text as evidence to be triangulated, not counted.

Why this is suddenly a problem

Open-text comments have always been the part of a course evaluation that institutions trust most. A specific, articulate paragraph reads as authentic in a way a 4-of-5 tick never could. That trust rested on an implicit assumption: producing a coherent, relevant comment required a human who had actually taken the course. Large language models have dissolved that assumption.

Westwood (2025), in Proceedings of the National Academy of Sciences, frames this bluntly as an existential threat to online survey research. The paper reports that 34.3% of respondents in a 2024 Prolific sample said they used AI to answer open-ended survey questions, and demonstrates that LLM-generated "respondents" achieve a 99.8% pass rate on standard attention checks while maintaining a consistent demographic persona across an entire questionnaire. The old quality-control toolkit — attention-check items, gibberish filters, response-time thresholds, consistency indices — was built to catch bots and inattentive humans. Reasoning-capable LLMs are neither: they read the question, stay on topic, and answer coherently. Detection methods premised on incoherence are, in Westwood's framing, obsolete.

What the research says

Three strands of evidence define the problem and the limits of the response.

The scale of AI-assisted responding. Beyond Westwood's 34.3% self-report figure, the broader survey-integrity literature documents a collapse in usable data. Pinzón and colleagues (2024), analysing 31 fraud-detection strategies across two surveys in Frontiers in Research Metrics and Analytics, describe a fall in usable online-survey responses "from 75% to 10%" in recent years as bots and coordinated fraud have proliferated. Crucially, they find that open-ended questions are among the more useful detection tools — bots and templated responders historically struggle to produce meaningful, experience-grounded answers — but they also note that AI-augmented fraud increasingly defeats exactly that defence.

Careless responding predates AI. Even before generative models, Meade and Craig (2012), in Psychological Methods, showed that roughly 10-12% of undergraduates completing a lengthy survey for course credit were careless responders, falling into two distinct patterns (random and non-random) that require different indices to detect. The lesson is that data quality was always partly an illusion; AI lowers the effort of faking and raises the realism of the fake, but it intensifies a pre-existing problem rather than inventing a new one.

Detection is an arms race the defender is losing. The combined message of these sources is that signature-based detection (perplexity scores, watermarks, stylometry) is fragile: it produces false positives that wrongly accuse genuine students, false negatives as models improve, and it degrades the moment a student lightly edits the AI output. No current method reliably separates AI-written from human-written short comments at the level of an individual response.

Why it matters for course evaluation in practice

It would be easy to assume this is a paid-panel problem that does not touch authenticated classroom surveys. That is half right, and the distinction is the key to a sensible policy.

  • The incentive structure is different — and milder. Prolific respondents are paid per completion, so there is a direct financial motive to mass-produce fake answers; a student evaluating their own course is not. This materially lowers the expected rate of adversarial faking in institutional SET. But it does not remove the driver that matters most here: effort reduction. A student who finds the form tedious may paste "write three sentences of polite feedback about a statistics course" into ChatGPT not to defraud anyone but to finish faster. That produces fluent, generic, content-free comments that pollute thematic analysis.
  • It corrupts the one channel you cannot quantify-check. A skewed Likert distribution can at least be modelled. AI-generated open-text is more insidious because it looks like the richest data you have. Generic AI prose ("the instructor was knowledgeable and the course was well-structured, though pacing could improve") is plausible for almost any course, so it survives a human reader's skim and dilutes genuine signal in any text-analytics pipeline.
  • It raises the stakes for how comments feed decisions. If open-text increasingly feeds programme review or promotion cases, a growing fraction of synthetic, course-agnostic prose silently lowers the evidential value of the whole corpus.

Limitations and honest caveats

Intellectual honesty requires resisting both panic and complacency:

  • The 34.3% figure is from a paid panel, not a classroom. It is the strongest available data point but it almost certainly overstates the rate in authenticated, unpaid, low-stakes course evaluations. Treat it as an upper-bound warning, not a measured prevalence for SET. We currently lack good prevalence estimates specific to institutional course feedback.
  • AI use is not automatically illegitimate. A non-native speaker who writes genuine feedback and asks AI to polish the grammar has not faked anything; the content is real. Any policy that treats all AI involvement as fraud will wrongly penalise legitimate accessibility use and erode trust. The target is content-free responses, not the presence of a tool.
  • Detectors cause real harm. Because no detector is reliable at the individual level, accusing a student of submitting AI feedback based on a classifier score is indefensible — the false-positive rate is too high and the reputational cost too great. Detection belongs at the aggregate, signal-quality level, not the disciplinary level.
  • This is a moving target. Every claim here is provisional against model improvement. The durable strategies are the ones that do not depend on detecting the AI at all.

How Koji incorporates this

Koji's design predates this threat but is unusually well-matched to it, because its core method is conversational rather than form-based.

  • Adaptive probing raises the cost of a one-shot paste. A static text box can be answered by pasting a single AI-generated paragraph. Koji's AI-moderated interview is a turn-by-turn dialogue: it asks a question, then follows up on the specific answer given, requesting concrete examples ("you said the labs were rushed — which lab, and what would have helped?"). Course-agnostic AI prose breaks down under this specificity, because each follow-up demands grounded detail the generic answer cannot supply. This is designed to mitigate, not eliminate, effort-saving faking — a determined student can still run the conversation through an assistant.
  • Authenticated, low-stakes collection. Koji is built for identified institutional cohorts rather than open, paid panels, which removes the financial fraud incentive that drives the worst panel behaviour and lets representativeness be checked against the enrolled class.
  • Quality scoring at the aggregate, not accusatory, level. Koji can flag low-information or templated responses for analytic down-weighting and surface them to evaluators as signal-quality metadata, rather than producing an individual "AI/human" verdict that no method can justify. The aim is to protect the integrity of the synthesis, not to police students.
  • Triangulation so no single channel is load-bearing. Because Koji combines structured items (scale, single_choice, ranking, yes_no) with probed open_ended responses and reports across cohorts, a pocket of generic comments cannot quietly drive a conclusion on its own. Open-text is treated as evidence to be corroborated, exactly the posture the integrity literature recommends.
  • Designing the question to resist genericness. Specific, experience-anchored prompts (the kind good open-ended design already calls for) are harder to satisfy with course-agnostic AI text than vague "any other comments?" boxes, so Koji's question design doubles as an integrity control.

The same engine guards commercial research: Koji's core platform at koji.so applies adaptive conversational probing and aggregate quality scoring to customer and product studies, where AI-generated and incentivised panel responses are an even sharper threat than in the classroom.

Related Resources

References

  • Westwood, S. J. (2025). The potential existential threat of large language models to online survey research. Proceedings of the National Academy of Sciences, 122(47), e2518075122. https://doi.org/10.1073/pnas.2518075122
  • Pinzón, N., Koundinya, V., Galt, R. E., Dowling, W. O., Baukloh, M., Taku-Forchu, N. C., Schohr, T., Roche, L. M., Ikendi, S., Cooper, M., Parker, L. E., & Pathak, T. B. (2024). AI-powered fraud and the erosion of online survey integrity: an analysis of 31 fraud detection strategies. Frontiers in Research Metrics and Analytics, 9, 1432774. https://doi.org/10.3389/frma.2024.1432774
  • Meade, A. W., & Craig, S. B. (2012). Identifying careless responses in survey data. Psychological Methods, 17(3), 437-455. https://doi.org/10.1037/a0028085