New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Sector trends9 min read

How Do You Evaluate a Course That Now Uses Generative AI?

Generative AI has quietly rewritten what happens inside thousands of European courses — yet most evaluation instruments still ask the questions they asked in 2019. Here is what course evaluation needs to measure once a course is taught, and learned, with AI.

Koji for Education

Research & Editorial Team · June 13, 2026

Answer first: When a course uses generative AI — whether the instructor uses it to teach or students use it to learn — the old satisfaction questionnaire stops measuring the right things. Evaluation needs to ask new questions: whether AI use was clear and fair, whether it supported or short-circuited learning, whether assessment still measured the student rather than the model, and whether students built durable skills or rented them from a chatbot. None of this fits a five-point "the course was well organised" item. It requires evaluation that can probe, follow up, and surface what actually happened.

The course changed; the survey did not

Generative AI adoption in higher education is no longer a frontier story; it is the baseline. In the UK, the Higher Education Policy Institute's 2025 Student Generative AI Survey found that 92% of students now use AI tools in some form, up from 66% the previous year — one of the fastest behavioural shifts higher education has ever recorded. On the staff side, the same survey found the proportion of students who think university staff are "well-equipped" to work with AI rose from 18% to 42% in a single year, while only 29% felt their institution actively encouraged AI use.

Meanwhile, institutional governance lags badly. UNESCO, in its 2023 Guidance for Generative AI in Education and Research, reported that fewer than 10% of schools and universities had any formal institutional guidance on generative AI. The classroom has been transformed; the policy scaffolding and — crucially for quality assurance — the evaluation scaffolding has not caught up.

This matters because course evaluation is the instrument universities use to know whether teaching is working. If a course is now half-taught through an AI tutor, or if a third of the cohort completed the formative tasks with a chatbot, an instrument that only asks about lecture clarity and assessment fairness is measuring a course that no longer exists.

What evaluation now needs to surface

A course that uses generative AI introduces at least four questions that traditional Student Evaluation of Teaching (SET) instruments were never designed to answer.

1. Was the AI policy clear, consistent, and fair? Students report enormous variation — sometimes within the same programme — about what is permitted. One module bans AI outright; the next requires it. Evaluation should ask whether expectations were explicit and whether they were applied even-handedly. Ambiguity here is a genuine teaching-quality failure, not a footnote.

2. Did AI support learning or substitute for it? This is the pedagogical crux. AI can scaffold understanding (explaining a proof three different ways) or it can short-circuit it (producing the essay the student never learned to write). A single satisfaction score cannot tell these apart — and a happy student is not evidence of learning, a point we develop in measuring skill development, not just student happiness. Evaluation needs to ask how students used AI and what they could do afterwards without it.

3. Did assessment still measure the student? If the assessment was vulnerable to AI completion, the course's grades may be measuring the tool. Students often know this better than anyone — they can tell you which tasks were "AI-proof" and which were theatre. Surfacing that intelligence is one of the most valuable things evaluation can now do.

4. Did the instructor's own AI use help? Increasingly, instructors use AI to generate materials, feedback, or examples. Students notice when AI-generated feedback is generic, and they notice when it is genuinely faster and more useful. Evaluation should capture the difference rather than pretend the instructor's workflow is unchanged.

But isn't this just an academic-integrity problem, not an evaluation problem?

This is the most common objection, and it is half right. Most institutions have responded to generative AI through the integrity and assessment-design route: detection tools, redesigned assignments, honour-code updates. That work is necessary. But treating AI purely as a cheating risk misses what evaluation is uniquely positioned to capture.

Integrity policy tells you what is allowed. Evaluation tells you what happened — and whether what happened helped students learn. Those are different questions. A course can be fully compliant with the integrity policy and still fail pedagogically because the AI policy confused students, or because the "AI-resistant" assessment was so artificial it taught nothing. Conversely, a course might integrate AI brilliantly in ways no integrity framework anticipated. Only evaluation, asked well, reveals this. Outsourcing the entire AI question to the integrity office leaves the teaching-and-learning question unanswered.

A second objection: "students are not qualified to judge whether AI helped them learn." Partly true — students are unreliable narrators of their own learning gains, a limitation that applies to all self-reported learning on evaluations. But they are excellent witnesses to process: what they did, what was confusing, where the assessment felt gameable. Good evaluation asks students about observable experience, not about constructs they cannot assess, and then triangulates with other evidence — the approach we argue for in triangulation in teaching evaluation.

Why static surveys struggle here

The problem with bolting "AI questions" onto a legacy Likert form is that AI use is heterogeneous and conditional. Whether AI helped depends entirely on how it was used — and a fixed scale item cannot branch. Ask "Did AI tools support your learning? (1–5)" and you get an average that blends the student who used AI to deepen understanding with the one who used it to avoid the work. The mean is meaningless because it sums two opposite phenomena. This is the same averaging trap that undermines ordinal evaluation generally, which we examine in why averaging Likert scores misleads.

What this topic needs is follow-up: when a student says AI was central, ask how; when they say assessment felt gameable, ask which task. Static surveys cannot do this. Human interviews can, but not at the scale of every course, every term — and human interviewers introduce their own inconsistency.

There is also a moving-target problem. The tools, the norms, and students' fluency with generative AI are changing every term, so a static question bank written this year will be measuring last year's practice by the next cycle. Evaluation of AI-augmented teaching therefore has to be revisable at the speed the technology moves — closer to continuous sensing than to a fixed annual instrument carved in stone. An evaluation that cannot be re-pointed quickly will keep producing confident answers to questions that no longer match what is happening in the room.

Where Koji fits

Koji for Education was built for exactly this kind of conditional, probing evaluation. Its AI-moderated conversational interviews can follow a student's answer wherever it leads: when a respondent mentions using generative AI, the interview asks how, for which tasks, and what they could do without it — the follow-up a paper form can never ask. Because the moderation is standardised and bias-aware, every student gets the same depth of probing without the inconsistency of dozens of different human interviewers, and without a tired teaching assistant skipping the hard questions at 11pm.

Koji's automatic thematic analysis then turns hundreds of these conversations into a structured map: where AI policy confused students, which assessments felt gameable, where AI genuinely supported learning. Its six structured question types (open-ended, scale, single-choice, multiple-choice, ranking, yes/no) let evaluation teams combine a few comparable metrics with the open, conditional probing that AI-augmented teaching demands. And because Koji is GDPR/AVG-compliant and EU-appropriate in its data handling, institutions can ask these sensitive questions — about student AI use and instructor AI use — within a defensible governance frame, complementing rather than cutting across the integrity and regulatory work we discuss in whether AI-assisted evaluation is high-risk under the EU AI Act.

The same conversational engine underpins the main Koji platform, where teams researching how their own users adopt AI products face a strikingly similar measurement problem: heterogeneous, conditional behaviour that a fixed survey flattens into a useless average.

Generative AI has already changed what your courses are. The question is whether your evaluation can still see them clearly. Explore how Koji for Education evaluates AI-augmented teaching.