New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Sector trends10 min read

AI in European Higher-Education Quality Assurance: A Sober Guide

AI is arriving in quality assurance just as the EU AI Act, the ESG, and GDPR set the terms of engagement. A clear-eyed look at where AI genuinely helps QA, where it is legally high-risk, and how to tell the difference.

Koji for Education

Research & Editorial Team · June 1, 2026

The short answer: AI can meaningfully improve quality assurance in European higher education — chiefly by making student feedback richer, faster to analyse, and more consistently moderated — but only inside a regulatory frame that the EU AI Act, the ESG, and GDPR now define. The most important distinction for QA leaders is this: using AI to gather and analyse student feedback about a course is a fundamentally different, and far lower-risk, activity than using AI to assess individual students' learning outcomes. Conflating the two is the fastest way to either over-regulate a useful tool or under-regulate a dangerous one.

Why this conversation is happening now

European QA does not operate in a vacuum. The Standards and Guidelines for Quality Assurance in the European Higher Education Area (ESG 2015), maintained under ENQA, expect institutions to monitor and periodically review their programmes so they continue to meet the needs of students and society (Standard 1.9), and to take student voice seriously as evidence. National agencies — NVAO in the Netherlands and Flanders, and equivalents elsewhere — translate these expectations into accreditation requirements. The pressure, in short, is to produce credible, acted-upon evidence of teaching quality. That is precisely the work AI is now being marketed to help with, which is why every QA office is fielding the question.

Where AI genuinely helps QA

Strip away the hype and a few concrete, defensible use cases remain:

  • Analysing open-text feedback at scale. The single biggest waste in QA is the unread free-text box. Thematic analysis that reliably clusters thousands of comments into recurring issues turns qualitative feedback from a token gesture into usable evidence.
  • Consistent, bias-aware moderation. Human interviewers and focus-group facilitators vary — in framing, in follow-up, in patience. A standardised AI moderator asks every cohort the same way, removing one source of inconsistency that undermines comparability across programmes.
  • Probing beyond the number. AI-moderated conversation can ask why a student rated something low and follow the thread, recovering the diagnostic detail that a Likert form discards.
  • Faster closing of the loop. Quicker synthesis means programmes can respond within a cycle rather than a year later, which is both better practice and a stronger accreditation story.

None of this requires AI to judge anyone. It is feedback infrastructure, not a verdict machine — and that distinction is the heart of the regulatory question.

The line that matters: feedback vs. assessment

The EU AI Act (in force from 2024, with obligations phasing in) classifies certain education uses as high-risk under Annex III — specifically AI used to determine admissions, to evaluate learning outcomes (including where those outcomes steer the learning process), to assign the appropriate level of education, and to monitor for prohibited behaviour during tests. These designations exist because algorithmic decisions about a student's progression affect fundamental rights and life chances.

Course evaluation — collecting and analysing students' feedback about teaching and a course — is a different activity. It does not grade the student, gate their progression, or determine their certification. Recognising this distinction is not a loophole; it is the correct reading of why the high-risk category exists. The responsible posture for a QA office is therefore twofold: do not quietly let a feedback tool drift into territory where it scores or ranks individual students (which would change its risk classification), and do not treat a teaching-feedback tool as if it carried the obligations of an automated grading system. Precision protects both students and the institution.

GDPR is the constant, not the afterthought

Whatever the AI Act classification, GDPR (the AVG in the Netherlands) applies to every student response. Course feedback can be sensitive — open text often reveals identifiable details, complaints about named staff, or information about a student's circumstances. The non-negotiables: a lawful basis and clear purpose limitation, genuine data minimisation, transparency about how AI processes responses, appropriate retention limits, and EU-appropriate data handling. An AI QA tool that cannot articulate its GDPR posture in plain terms is not ready for institutional use, regardless of how impressive its analysis looks.

The strongest counterargument — taken seriously

"Given documented bias in human-collected evaluations, isn't adding AI just stacking a second opaque bias on top of the first?" This is the objection QA professionals should press hardest, and it is partly right. AI systems can encode and amplify bias, and a poorly governed model could systematise unfairness at scale rather than reduce it. Opacity is a real cost: a black-box score is arguably worse than a flawed-but-legible average.

The honest response is not to wave this away but to set conditions. AI earns its place in QA only when it is (1) transparent about how it moderates and analyses, (2) bias-aware by design, with standardised framing rather than ad hoc prompting, (3) auditable, so institutions can inspect how themes were derived, and (4) used to surface and structure evidence for human judgement, not to automate personnel decisions. Under those conditions AI replaces an opaque, inconsistent human process with a more consistent and inspectable one. Without them, the critic is correct and the tool should be refused. The technology is not self-justifying; the governance is what makes it defensible.

A practical checklist for QA leaders

Before adopting any AI tool in evaluation, ask:

  1. Does it assess students, or gather feedback about teaching? Only the latter sits clearly outside the AI Act's high-risk grading category — keep it there.
  2. Can the vendor state its GDPR/AVG basis, retention, and data-residency plainly?
  3. Is the moderation standardised and bias-aware, and can you audit how themes are produced?
  4. Does it keep humans in the loop for any consequential decision?
  5. Does it help you close the loop, evidencing action for ESG and accreditation, rather than just generating more data?

Where Koji fits

Koji for Education is designed to sit on the right side of every line above. It is a feedback and evaluation platform, not a student-grading system: it uses AI-moderated conversational interviews to gather and analyse students' views about their courses, with automatic thematic analysis of open-text responses, quality scoring, and standardised, bias-aware moderation that removes human-moderator inconsistency. It supports formative, mid-cycle collection and closing-the-loop action tracking, producing exactly the acted-upon evidence the ESG expects, with programme- and institution-level reporting for accreditation. Data handling is GDPR/AVG-compliant and EU-appropriate.

To be precise about the claims: Koji mitigates and surfaces bias through consistent, transparent moderation and structured analysis; it does not eliminate bias, and it deliberately does not assess or rank individual students. That restraint is the point — it keeps the tool clearly outside the EU AI Act's high-risk assessment category while still doing real QA work. Compared with legacy SET tools that simply average a Likert score, this is the AI-native, evidence-aware way to meet modern European QA expectations.

Institutions whose research offices also run user or stakeholder studies beyond teaching quality use the same underlying interview engine in the main Koji platform.

The competitive context: from static SET to AI-native QA

It is worth being fair about where legacy tools succeed and where they stop. Established systems such as EvaSys, generic Qualtrics surveys, and paper forms are reliable at one thing: administering a fixed questionnaire to a large population and tabulating the results. For straightforward logistics — distribution, collection, basic reporting — they work. The limitation is conceptual, not technical: they are built around the static, average-a-Likert-score model, and they inherit every weakness that model carries. They cannot ask a follow-up question, cannot probe why a rating was low, and turn rich open-text into an unread overflow field. Their headline output is the very averaged number whose validity the methodology literature has questioned for decades.

The AI-native approach does not discard what these tools do well; it changes what evaluation is. Instead of administering a frozen form, it holds a responsive conversation; instead of tabulating ordinal codes, it analyses meaning at scale; instead of one moderator's inconsistency or none at all, it applies standardised, bias-aware framing to every interaction. That is the substantive difference between the static past and the evidence-aware present — and it is why institutions reviewing their QA toolchain are increasingly unwilling to treat "we have a survey tool" as the end of the conversation.

The bottom line

AI in European QA is neither a silver bullet nor a compliance trap — provided you hold one distinction firmly: gathering and analysing feedback about teaching is low-risk, useful work; assessing students with AI is high-risk and governed accordingly. Demand transparency, bias-awareness, auditability, human oversight, and a clear GDPR posture, and AI becomes a genuine upgrade to how institutions listen, act, and evidence quality. Skip those conditions, and the skeptics are right to refuse it.