New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Sector trends8 min read

Synthetic Students: Why LLM-Simulated Feedback Cannot Replace Course Evaluation

"Silicon sampling" promises survey results without surveys — prompt an LLM with student personas and skip the response-rate problem entirely. The methodological evidence explains why that is exactly backwards for course evaluation, and where AI genuinely belongs in the feedback loop.

Koji Education Team

Product ·

A seductive idea is circulating wherever survey fatigue meets AI enthusiasm: if large language models can convincingly role-play a demographic profile, why not generate student feedback synthetically? Prompt the model with your course description and a distribution of student personas, and receive instant "evaluation data" with a 100% response rate, no fatigue, no GDPR headaches. The political-science literature that first took this idea seriously has now tested it carefully, and the verdict is specific: LLM "silicon samples" can approximate averages while systematically failing at variance, minority viewpoints, and stability — which are precisely the properties course evaluation exists to capture. Simulation has legitimate uses in the evaluation workflow, but substituting for student voice is not one of them, and quality-assurance frameworks could not accept it anyway.

Where the idea came from — and it is not a straw man

The intellectual origin is respectable. Argyle et al. (2023), in Political Analysis, showed that GPT-3 conditioned on thousands of real respondents' socio-demographic backstories could reproduce response distributions of human subgroups with surprising fidelity — coining the "silicon sample." Marketing firms now sell synthetic panels; product teams A/B-test copy against persona-prompted models. The extension to education writes itself, and versions of it are already being pitched to institutions: simulate the student body, "pre-evaluate" the course, skip the survey.

What the evidence actually shows

The most careful stress-test to date is Bisbee et al. (2024), also in Political Analysis, who prompted an LLM to adopt personas from a major survey's real respondent profiles and compared synthetic against human answers. Four findings matter for education:

  1. Averages roughly right, variance badly wrong. Synthetic responses showed markedly smaller variance than real human data. The model produces the central tendency of a persona and shaves off the disagreement, ambivalence and outliers.
  2. Overconfident and biased at the tails. Estimates for minority opinions and less-typical subgroups deviated most — the model regresses everyone toward the stereotype of their demographic label.
  3. Relationships between variables distort. Regression coefficients estimated on synthetic data differed from those on human data, sometimes substantially and systematically. Synthetic data does not merely add noise; it changes what you would conclude.
  4. Unstable to prompt and time. Responses shifted with prompt wording and with when the model was queried (as underlying models updated) — a reproducibility problem with no analogue in human surveying.

Translate each finding into course evaluation and the problem becomes vivid. The value of evaluation is almost never the mean — means hide the bimodality that signals a course working for one group and failing another. It is the surprising minority signal: the three students who could not follow the second half, the placement student for whom the scheduling failed, the safeguarding disclosure in a free-text box. A method that is competent at averages and systematically wrong about variance, tails and subgroups deletes exactly the information quality assurance exists to find. And a synthetic student, prompted from last year's course description, cannot report the thing evaluation most needs to catch: what changed — the new instructor's pacing, the broken lab kit, this cohort's particular struggle.

There is a deeper epistemic point beneath the psychometrics. An LLM's persona-response is drawn from the distribution of text about students in its training data. Your students' experience of your course this semester is not in that distribution. Simulation can interpolate what students-in-general plausibly say; it cannot observe what your students actually experienced. Course evaluation is measurement, not plausibility generation.

The regulatory dead end

Even if the psychometrics improved, European quality assurance has no slot for synthetic voice. The ESG build student participation into quality processes — Standard 1.3 expects student-centred learning and 1.9 expects information from students to feed ongoing monitoring. An accreditation panel asking for evidence of student feedback will not accept a model's estimate of what students would probably have said, any more than a clinical trial regulator would accept simulated patients. And presenting synthetic feedback as student voice to any internal committee is, bluntly, fabrication — the institutional-integrity equivalent of inventing survey responses, a problem institutions are otherwise spending real money to detect.

The steelman: what simulation is legitimately for

Intellectual honesty requires the other side of the ledger, because simulation does have defensible uses around the measurement act:

  • Instrument piloting. Running draft questions past an LLM is a cheap first-pass check for ambiguity and double-barreled phrasing before human cognitive pretesting — a complement to, not substitute for, writing better questions.
  • Analysis rehearsal. Synthetic data can stand in for real data while building dashboards and pipelines, avoiding unnecessary processing of personal data — a genuinely GDPR-friendly use.
  • Interviewer training and red-teaming. Simulated respondents help test that an AI moderator handles edge cases — hostile answers, disclosures, off-topic tangents — before real students meet it.
  • Power planning. Plausible synthetic distributions can inform decisions about sample needs for small-course reporting.

Notice the pattern: every legitimate use treats simulation as scaffolding for better measurement of real students. The moment synthetic output is reported as feedback, the line is crossed.

"But the models will get better" deserves a direct answer: the tail-compression problem has proven persistent across model generations, and the epistemic problem is structural. No future model gains access to what happened in your seminar room last Tuesday. Capability growth improves the plausibility of the interpolation; it cannot convert interpolation into observation.

A test for vendor pitches

Synthetic-respondent offerings rarely arrive labeled as such; they arrive as "AI-powered insight generation" or "predictive feedback analytics." Three questions separate legitimate tooling from fabrication-by-another-name. First: does any reported number originate from a model rather than a respondent? Summaries and theme labels derived from real answers are analysis; response values a model generated are not data. Second: can every claim be traced to identifiable (if pseudonymous) student responses? If the audit trail ends at a prompt, so does the evidence. Third: would you describe the method, in those words, in your self-evaluation report? An institution comfortable writing "projected student satisfaction was estimated by a language model" for an accreditation panel has at least been honest; one that would not write that sentence should not buy the product. The same discipline applies internally: pilot studies, dashboards and training environments built on synthetic data should carry the label all the way to any committee that sees them, because a synthetic number that survives two forwarding steps becomes indistinguishable from evidence.

The irony: AI's real leverage is on the other side of the microphone

The silicon-sampling pitch gets the technology's role in feedback precisely inverted. The expensive, failure-prone parts of course evaluation are asking good follow-up questions at scale and analyzing thousands of open-text answers — and those are the parts modern AI genuinely transforms, without replacing the respondent.

This is the architecture Koji for Education is built on. AI as interviewer: Koji's AI-moderated conversational evaluations probe each real student's answers — asking for the example, the moment, the specific difficulty — so institutions get interview-depth evidence at questionnaire scale, addressing the fatigue and shallow-response problems that make synthetic shortcuts tempting in the first place. AI as analyst: automatic thematic analysis organizes what real students said into themes with quality scoring, and response-quality signals help flag inauthentic submissions rather than manufacture them. Humans as the only source of ground truth: six structured question types, formative mid-cycle collection, closing-the-loop action tracking and programme-level reporting — all of it anchored to actual student voice, handled with GDPR/AVG-appropriate EU data practices.

Research teams tempted by synthetic panels for user or market research face the same trade-offs; the main Koji platform applies this same real-respondent, AI-moderated architecture beyond education.

The takeaway

Silicon sampling is a real method with real findings — and those findings are the argument against using it as evaluation. Synthetic respondents compress variance, blur minority signal, distort relationships and drift with prompts and model versions, while the entire point of course evaluation is to detect specific, local, often-minority experience that no training corpus contains. Use simulation to pilot instruments, rehearse pipelines and train moderators. Then put AI where it earns its keep: interviewing and analyzing real students, rigorously.

See how Koji for Education delivers interview-depth evidence from real students at survey scale — book a demo.