New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
accreditation12 minutes

The EU AI Act and Course Evaluation Software: What European Universities Actually Have to Do

Course evaluation is usually not an Annex III high-risk use of AI. The obligations that genuinely bite are Article 50 transparency (2 August 2026, not deferred), Article 4 AI literacy, and the GDPR. A procurement-ready guide with a requirement-to-evidence mapping.

Koji Education Team

Product

Most course evaluation software is not a high-risk AI system under the EU AI Act. That single sentence resolves the question most quality assurance directors are actually asking in 2026 — and it is the opposite of what a lot of vendor marketing currently implies.

The obligations that genuinely apply to a university running AI-supported course evaluation are narrower, earlier, and more mundane than the high-risk regime: transparency under Article 50, which applies from 2 August 2026 and was not postponed; AI literacy under Article 4, which has been in force since 2 February 2025; and the GDPR, which never went away. The dramatic high-risk conformity-assessment machinery — risk management systems, technical documentation, conformity assessment, registration in the EU database — mostly does not attach to students telling you what they thought of a module.

This guide sets out where the line actually falls, which adjacent uses do cross into high-risk territory, and what to put in your procurement documents. It is written for QA directors, deans of teaching and learning, institutional research leads, and procurement officers at European institutions.

Scope note. This is a practitioner's reading of the Regulation for procurement purposes, not legal advice. Classification depends on your specific configuration and intended purpose. Your Data Protection Officer and legal counsel own the final call.

The timeline, as it stands in July 2026

The AI Act — Regulation (EU) 2024/1689 — entered into force on 1 August 2024 and phases in over several years. The phasing changed materially in mid-2026, and it is worth being precise, because a lot of guidance written in 2024 and 2025 is now out of date.

ObligationApplies fromStatus
Prohibited practices (Art. 5)2 February 2025In force
AI literacy (Art. 4)2 February 2025In force
General-purpose AI model obligations2 August 2025In force
Transparency for certain AI systems (Art. 50)2 August 2026Not deferred
High-risk obligations, standalone Annex III systemsDeferred to 2 December 2027Deferred by the Digital Omnibus
High-risk obligations, Annex I embedded productsDeferred to 2 August 2028Deferred by the Digital Omnibus

The deferral came through the so-called Digital Omnibus on AI. Following a provisional political agreement on 6 May 2026, the European Parliament endorsed the package on 16 June 2026 and the Council gave its final approval on 29 June 2026, with publication in the Official Journal following shortly after. The practical effect: the date on which non-compliance with the high-risk regime begins to bite moved by roughly sixteen months. The underlying substantive obligations did not change — they were postponed, not softened.

The critical detail for evaluation buyers is the row in bold. Article 50 transparency was carved out of the deferral. If your institution is deploying a conversational AI system that students interact with, 2 August 2026 is a live date, and it is weeks away at the time of writing.

Why course evaluation is usually not high-risk

Annex III point 3 lists the high-risk education use cases. Read the actual text, because the wording does the work:

  • 3(a) — systems used "to determine access or admission or to assign natural persons to educational and vocational training institutions at all levels"
  • 3(b) — systems used "to evaluate learning outcomes, including when those outcomes are used to steer the learning process of natural persons in educational and vocational training institutions at all levels"
  • 3(c) — systems used "for the purpose of assessing the appropriate level of education that an individual will receive or will be able to access"
  • 3(d) — systems used "for monitoring and detecting prohibited behaviour of students during tests"

Every one of these points is about the AI system forming a judgement about a student. Admission, learning outcomes, educational level, exam misconduct. The subject of the assessment is the learner.

Course evaluation inverts that relationship. In a student evaluation of teaching, the student is the respondent, not the subject. The AI system is collecting and analysing what a student says about a module, a teaching approach, an assessment design, or a learning environment. It is not evaluating that student's learning outcomes, it is not deciding their educational level, and it is not deciding whether to admit them. On a plain reading of Annex III point 3, standard course evaluation falls outside all four sub-points.

This is not a loophole — it is the regulation working as designed. The high-risk regime exists because AI decisions about individuals create risks to their fundamental rights and life chances. A student saying the seminar reading list was too heavy does not implicate their life chances.

Where the line genuinely does get crossed

Being honest about the boundary is more useful than a blanket reassurance. Four configurations move you into scope, and a serious QA function should check all four:

1. Feeding evaluation output into staff decisions. This is the big one, and it is not in point 3 — it is in point 4. Annex III point 4(b) covers AI systems used "to make decisions affecting terms of work-related relationships, the promotion or termination of work-related contractual relationships, to allocate tasks based on individual behaviour or personal traits or characteristics or to monitor and evaluate the performance and behaviour of persons in such relationships."

If your institution uses AI-generated teaching evaluation scores or summaries as an input to promotion, tenure, probation, contract renewal, or performance management for academic staff, you are plausibly inside Annex III point 4(b) — and the system is high-risk, notwithstanding that point 3 does not apply. This is the single most commonly missed classification risk in higher education evaluation, and it is a governance question, not a software question. Many institutions already prohibit this use for sound methodological reasons; the AI Act gives them a second, harder reason.

2. Blending evaluation with learning-outcome assessment. If the same platform also scores student work, produces attainment judgements, or steers a student's learning pathway, point 3(b) attaches to that functionality.

3. Proctoring or behaviour monitoring. Point 3(d), squarely.

4. Admissions or placement uses. Points 3(a) and 3(c).

What actually applies: transparency and AI literacy

For an institution running AI-moderated course evaluation and nothing else, two obligations are live now.

Article 50 — transparency. Providers of AI systems intended to interact directly with people must ensure those people are informed that they are interacting with an AI system, unless it is obvious to a reasonably well-informed person in the circumstances. For a conversational interview, the disclosure has to reach the student before or at the very beginning of the conversation — not buried in a privacy policy, not disclosed afterwards. Article 50 also carries marking obligations for AI-generated synthetic content, with a slightly later date for generative systems already on the market before 2 August 2026.

In procurement terms: ask any vendor of a conversational evaluation tool to demonstrate, on screen, where and when the AI disclosure appears to a student. If they cannot show you, that is a finding.

Article 4 — AI literacy. Providers and deployers must take measures to ensure a sufficient level of AI literacy among staff and others operating AI systems on their behalf, proportionate to their role, expertise, and context of use. This has applied since 2 February 2025. For evaluation, that means the QA officers configuring interview protocols and the programme leaders reading AI-generated thematic summaries need documented, role-appropriate training. This is an institutional obligation you cannot fully outsource to a vendor — though a vendor should be supplying the training materials that make it achievable.

The GDPR, unchanged. The AI Act sits alongside data protection law, it does not replace it. Lawful basis, data minimisation, retention, transfer mechanisms, student anonymity thresholds, and the Article 22 rules on automated decision-making all continue to apply on their own terms. For most institutions this remains the heavier compliance burden of the two regimes. Our GDPR guidance for evaluation data covers how this intersects with quality assurance evidence.

Requirement-to-evidence mapping

What a procurement panel needs is not a legal essay but a column of evidence it can file. This maps each live obligation to the concrete artefact you should be able to produce.

ObligationWhat the deployer must showConcrete artefact from Koji
Art. 50 — AI interaction disclosureStudents are told they are talking to an AI before the conversation startsStandard pre-interview disclosure screen; screenshot and configuration record exportable for the compliance file
Art. 50 — synthetic content markingAI-generated summaries are identifiable as suchAI-generated thematic summaries and reports are labelled as machine-generated, with the underlying verbatim quotations traceable
Art. 4 — AI literacyRole-appropriate training for staff operating the systemDocumented onboarding for QA administrators and report readers; method documentation describing how thematic analysis is produced
Annex III p.4(b) — staying out of scopeEvaluation outputs are not used for individual staff performance decisionsInstitution-level and programme-level reporting; documented policy position on permitted use of outputs
Traceability of analysisAnalytical claims can be audited back to sourceEvery theme links to the verbatim student responses that generated it
GDPR — data location and transfersWhere personal data is processed and storedEU data residency and processing documentation, DPA, retention configuration
GDPR — anonymity in reportingSmall-cohort responses cannot re-identify studentsConfigurable minimum-response thresholds before results release

The traceability row deserves emphasis for a QA audience. Whether or not a system is formally high-risk, an evaluation platform whose conclusions cannot be traced back to source data is difficult to defend in an accreditation panel — and impossible to defend if a member of staff challenges a summary. Koji's thematic analysis is built so that each theme resolves to the specific student responses behind it, which serves the compliance requirement and the methodological one at the same time.

When a different tool is the better choice

Three honest cases where you should not buy an AI-native platform for this:

  • Your institution has a moratorium on AI processing of student data. Some institutions and works councils have one, and it is a legitimate position. A traditional Likert-scale platform with no AI component sits outside Article 50 entirely and involves a shorter governance conversation. You will pay for that simplicity in the depth and quality of what you learn.
  • Your evaluation instrument is genuinely settled and purely quantitative. If you run a validated, unchanging quantitative instrument and your only need is administration and distribution at scale, established survey-automation platforms do that competently and the AI question does not arise.
  • You need it inside an existing suite for reasons that outrank capability. If integration and consolidated procurement dominate your decision, LMS-native or incumbent suite tooling may win on grounds that no feature comparison will overturn.

Where AI-moderated evaluation earns its place is the qualitative side: standardised follow-up probing that a static form cannot do, thematic analysis across thousands of open responses that no QA office has capacity to code by hand, and the resulting evidence base being genuinely defensible in a review panel. The same interview engine underpins the main Koji platform for user and customer research outside higher education.

What to do before 2 August 2026

A short, concrete checklist:

  1. Classify your actual use. Write down, in one paragraph, what your evaluation AI does and who it forms judgements about. File it. This is your classification record.
  2. Check the point 4(b) exposure. Confirm in policy — not just in practice — that evaluation outputs are not inputs to individual staff performance or promotion decisions. If they are, escalate to legal now.
  3. Verify the Article 50 disclosure. Look at the actual student-facing screen. Confirm the AI disclosure appears before the conversation begins.
  4. Document AI literacy measures. Record what training your QA staff and report readers have had, and when.
  5. Refresh the DPIA. The AI Act does not replace your GDPR documentation; make sure it reflects the current processing.
  6. Put all of the above in the RFP. Ask vendors to evidence each row of the mapping table above rather than to assert compliance in prose.

Institutions that do this now will find the December 2027 high-risk deadline largely irrelevant to their evaluation stack — which is exactly the outcome a well-run classification exercise should produce.

Related Resources

Ready to see how this works in practice? Book a walkthrough and we will show you the Article 50 disclosure flow, the traceability model, and the EU data residency documentation on a live account.