New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods10 min read

Should Your Course Evaluation Ask How Students Used AI? What the Learning Evidence Says

Students now study with generative AI in nearly every course, and it changes what a satisfaction rating means. The evidence on AI and learning explains why — and what a course evaluation should and should not try to ask about it.

Koji Education Team

Product

In short: Yes — but ask about use, not about approval. A field experiment with nearly a thousand high-school mathematics students found that unrestricted GPT-4 access improved performance by 48% while students had it, and left them performing 17% worse than students who never had access once it was withdrawn (Bastani et al., 2025). AI use can raise the perceived ease of a course while lowering the learning it produces, which decouples satisfaction from outcomes further than it already was. Course evaluations should therefore capture behaviour (which tools, at which task, in place of what) as descriptive context — never as a compliance check.

Why this is now an evaluation-design question

For decades, the central critique of student evaluation of teaching has been that ratings reflect perceived rather than actual learning. Generative AI widens the gap between the two, and it does so unevenly across students within the same cohort.

Consider two students in one module. One uses an AI assistant to generate worked solutions to every problem set. The other uses it to check reasoning after attempting problems unaided. Both may report that the course was well paced and the materials clear. They will have learned very different amounts. A course evaluation that captures only their satisfaction records the same signal from both — and the signal has become less informative about the teaching than it was before either of them had the tool.

This is not a reason to make evaluations into an AI-policing instrument. It is a reason to collect enough context to interpret what the ratings mean.

What the research says

The anchor: help while it is there, harm once it is gone

Bastani and colleagues ran a field experiment with nearly a thousand students at a large high school in Turkey during the 2023–24 autumn semester, published in PNAS. Students practised mathematics with one of two AI tutors: GPT Base, an interface approximating standard ChatGPT, and GPT Tutor, prompted with teacher-designed safeguards that give hints rather than answers.

The results are the clearest natural statement of the problem:

  • With AI access during practice, performance improved substantially — a 48% grade improvement for GPT Base and 127% for GPT Tutor.
  • With access subsequently removed, GPT Base students performed 17% worse than students who had never had access at all.
  • GPT Tutor's safeguards largely eliminated that harm. The design of the tool, not its presence, determined the outcome.

The authors' interpretation is that without guardrails students use the model as a crutch during practice and never do the retrieval work that produces durable learning. The finding is a direct instance of the desirable-difficulties principle: the version that felt easier produced worse learning.

The wider literature is more positive — and that is the point

Deng, Jiang, Yu, Lu and Liu (2025) systematically reviewed 69 experimental studies of ChatGPT in education published 2022–2024, meta-analysing 62. They report that ChatGPT interventions enhanced academic performance and higher-order thinking propensities and reduced mental effort, with no significant effect on self-efficacy. Most studies were at university level, concentrated in language education, and involved direct student use in classroom settings.

The apparent contradiction with Bastani et al. resolves cleanly once you notice what each measures. Deng et al. largely aggregate outcomes assessed with the intervention in place; Bastani et al. measured what happened after it was taken away. Both are true. "AI improves performance during the task" and "unguarded AI reduces retained capability" are compatible claims — and the distinction is precisely what a satisfaction rating cannot see.

That reduced mental effort finding deserves particular attention from anyone reading course evaluations. Lower cognitive effort reliably feels like better teaching. It is not reliably better learning.

Why it matters for course evaluation in practice

1. AI use is now a confound you should measure, not assume away. Workload and difficulty items are among the most-used in any evaluation instrument, and both are now mediated by tool use. A module whose perceived workload has fallen year-over-year may have been redesigned, or its cohort may have adopted a tool. Without a use question you cannot distinguish these, and you will attribute the change to teaching. This is the same confounding logic as our analysis of emergency remote teaching.

2. Ask behaviourally and non-judgementally. "Did you use AI inappropriately?" measures willingness to self-incriminate. "Which of the following did you use an AI assistant for in this module?" with concrete task options — understanding a concept, checking work, drafting text, generating solutions — measures behaviour. The distinction is the standard social-desirability problem covered in our guide to anonymity and honest feedback.

3. Keep it strictly separate from academic-integrity processes. If students suspect evaluation responses feed misconduct proceedings, you lose both the AI data and the honesty of every other item. Say plainly, in the instrument, that responses are not used for integrity purposes. This is a precondition, not a nicety.

4. Ask about the course's AI design, not just the student's AI use. The actionable finding in Bastani et al. is about tool design — teacher-designed hints beat raw answers. The evaluable question for a module is therefore whether the course gave students usable guidance on when AI helps and when it substitutes for learning. That is a teaching-quality question with a real answer.

5. Expect the ceiling to move. If AI use raises perceived ease across a portfolio, satisfaction scores drift upward for reasons unrelated to teaching. Year-over-year comparisons that ignore this will misattribute a sector-wide shift to local improvement. See ceiling effects and regression to the mean.

Limitations and honest caveats

  • The anchor study is high-school mathematics in one country. Bastani et al. studied a single large Turkish high school in one semester with a well-defined problem-solving task. Mathematics practice is close to the ideal case for crutch behaviour; an essay-based humanities seminar may behave differently. Generalisation to European higher education requires argument.
  • The models have changed. These experiments used 2023–24 model generations. Both capability and default pedagogical behaviour have moved since. Directional conclusions about guardrails are more durable than the specific percentages.
  • This evidence base has a quality problem. A widely circulated meta-analysis of ChatGPT's effect on learning performance and higher-order thinking, published in Humanities and Social Sciences Communications, has been retracted. That is a reason for real caution about how quickly claims in this area are being made, and a reason to prefer registered field experiments and carefully bounded reviews.
  • Publication bias plausibly favours positive findings. Studies showing AI helps are easier to publish and to fund than studies showing it harms. The Uttl et al. re-analysis of the SET literature is a reminder of how much a corrected literature can differ from an uncorrected one.
  • Self-reported tool use is imperfect. Students under-report use they suspect is disapproved of, and misremember what they used a tool for weeks later. Treat these items as descriptive context, never as measurement of academic conduct.
  • This is a fast-moving area. Any institutional policy built on the current evidence should carry an explicit review date.

How Koji incorporates this

Koji's position is that AI use belongs in course evaluation as interpretive context — data that helps you read the ratings — and not as a compliance instrument.

  • Structured question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) support the behavioural item design this evidence calls for: multiple-choice task inventories of what AI was used for, ranking items on which resources students relied on most, and yes/no items on whether the course gave guidance about AI use.
  • AI-moderated conversational interviews go where a checkbox cannot. When a student indicates they used an assistant for problem sets, the interview can probe whether they attempted problems first, what they did when the model's answer was wrong, and whether they felt they understood the material afterwards. That is the difference between knowing a tool was used and knowing how it displaced or supported the learning.
  • Automatic thematic analysis of open text lets a programme see how AI use is described across a whole cohort, surfacing patterns — such as substitution concentrated in a specific assessment — that individual comments do not reveal.
  • Quality scoring flags low-effort and implausible responses, which matters more in this domain than most; our guide to AI-generated open text and data integrity covers the related risk that the feedback itself is model-written.
  • Mid-cycle collection surfaces a substitution pattern while the module is still running, when the teaching team can still change the task design — which is the intervention the Bastani findings actually support.
  • Bias-aware reporting keeps year-over-year comparisons honest about the confound rather than presenting a rising satisfaction trend as evidence of improvement.

Koji is designed to mitigate the interpretive blind spot AI use creates in conventional instruments. It does not detect AI use, does not verify student self-reports, and should not be used as an academic-integrity tool. The same conversational research engine, applied to product and customer research, is available at koji.so.

Related resources

References

  • Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences. https://doi.org/10.1073/pnas.2422633122
  • Deng, R., Jiang, M., Yu, X., Lu, Y., & Liu, S. (2025). Does ChatGPT enhance student learning? A systematic review and meta-analysis of experimental studies. Computers & Education, 227, Article 105224. https://doi.org/10.1016/j.compedu.2024.105224
  • Humanities and Social Sciences Communications (2025). RETRACTED ARTICLE: The effect of ChatGPT on students' learning performance, learning perception, and higher-order thinking: Insights from a meta-analysis. https://www.nature.com/articles/s41599-025-04787-y (cited here as an example of retraction risk in this literature)
  • Uttl, B., White, C. A., & Gonzalez, D. W. (2017). Meta-analysis of faculty's teaching effectiveness: Student evaluation of teaching ratings and student learning are not related. Studies in Educational Evaluation, 54, 22–42. https://doi.org/10.1016/j.stueduc.2016.08.007