New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Sector trends10 min read

Is Your AI Course-Evaluation Tool a High-Risk AI System? Article 14, Human Oversight, and the Honest Answer

The honest answer is usually no — an AI that analyses feedback about your teaching is not the same as an AI that grades your students. But there is a route where it does become high-risk, and the human-oversight standard behind Article 14 is worth adopting either way. A precise reading for European quality offices.

Koji Education Team

Product · August 13, 2026

Bottom line up front: An AI tool that analyses student course feedback — summarising open text, clustering themes, scoring sentiment about teaching — is, in most configurations, not automatically a high-risk AI system under the EU AI Act. The Act’s education high-risk category targets AI that evaluates learners, not AI that evaluates teaching. But there is a specific route by which a course-evaluation tool crosses into high-risk territory, the compliance clock has just been reset by a 2026 amendment, and the human-oversight logic of Article 14 is the right governance standard to adopt regardless of how your tool is legally classified. Precision here matters, because both over-claiming and under-claiming compliance carry real risk.

What the Act actually classifies as high-risk in education

Under Regulation (EU) 2024/1689, high-risk education AI is enumerated in Annex III, point 3. It covers AI systems intended to determine access or admission; to evaluate learning outcomes, including when those outcomes are used to steer the learning process; to assess the appropriate level of education a person will receive; and to monitor and detect prohibited behaviour during tests. Read those carefully: every item is about an AI acting on a student — admitting them, grading them, streaming them, or proctoring them.

An AI that reads what students wrote about a module or a lecturer and hands a themed summary to a teaching-and-learning committee is doing none of those things. It is analysing feedback about the provision, not evaluating the learner. On a plain reading of Annex III(3), a standard AI course-evaluation tool is not high-risk — a point we make in the specific context of course feedback in our analysis of the AI Act and student feedback.

The route by which it does become high-risk

Here is the nuance most vendor marketing skips. A course-evaluation output changes legal character depending on what decision it feeds.

If the AI-generated analysis of student feedback is used to make or materially inform decisions about staff — promotion, contract renewal, non-renewal of a fixed-term lecturer, performance management — then you are no longer only in education. You are in Annex III, point 4: employment and worker management, which classifies as high-risk AI systems used to evaluate performance and behaviour or to inform promotion and termination decisions. The same summary of student comments is low-stakes when it drives curriculum improvement and high-stakes when it drives someone’s job. This is precisely why we argue elsewhere that student evaluations should be handled with extreme care in tenure and promotion decisions, and why automated decision-making about faculty also engages GDPR Article 22. The determinant is not the tool; it is the consequence.

The clock just moved — but not the direction of travel

If you were bracing for high-risk obligations to bite in August 2026, note the change. The Digital Omnibus on AI (Regulation (EU) 2026/1744, published in the Official Journal on 24 July 2026) postponed the applicability of Annex III high-risk obligations to 2 December 2027 (and Annex I product-embedded high-risk AI to 2 August 2028). Two things did not move: the Article 50 transparency duty — telling people when they are interacting with an AI system — and the Article 4 AI-literacy obligation both remain in force. So even if your configuration is not high-risk, or benefits from the delay, you still owe students transparency that an AI is moderating or analysing their feedback, a point we treat in AI moderation disclosure. The heaviest regime is deferred; the baseline duties are not.

Article 14 is the standard to adopt anyway

Article 14 requires that high-risk AI systems be designed so that natural persons can effectively oversee them — understand their capacities and limits, watch for automation bias, correctly interpret output, and decide not to use it or to override it. Even where a course-evaluation tool is not legally high-risk, this is simply good governance, and the failure mode it guards against is real in evaluation contexts: automation bias, the documented tendency of committees to defer to a confident-looking machine summary, is a live danger when an AI condenses hundreds of student comments into three bullet points (our deep-dive here). If a human cannot trace a themed claim back to the underlying comments, they cannot oversee it — they can only trust it. Provenance and traceability, not just a tidy summary, are what make oversight possible (more).

"So you are admitting your own product might be high-risk?"

Yes — and that candour is the point. Any honest vendor should tell you that whether an AI course-evaluation system is high-risk depends on how the institution uses the output, not on a compliance badge the vendor prints. A tool used to improve teaching is very likely outside Annex III; the same tool wired into personnel decisions plausibly falls under Annex III(4). Claiming blanket "AI Act compliant, not high-risk" status regardless of deployment is exactly the kind of over-claim a methodologically literate quality office should distrust. The defensible position is: build to the human-oversight standard by default, be transparent under Article 50 now, and treat the December 2027 date as breathing room to get governance right, not permission to defer thinking about it.

What human oversight means in practice for a quality office

Article 14 can read as abstract, so it helps to translate it into what a teaching-and-learning committee would actually do. Effective oversight of an AI feedback tool means, concretely: that the people reading the output understand the tool can miss, over-weight or mischaracterise themes, and do not treat its summary as ground truth; that they can trace any AI-generated claim back to the underlying student comments and spot-check it; that they retain a genuine ability to disregard or override the analysis rather than rubber-stamping it; and that at least one person in the loop has enough AI literacy — the Article 4 duty — to know what the system can and cannot do. A summary no one can audit is not overseen; it is merely trusted.

Provider and deployer are not the same role

One further distinction shapes who owes what. The AI Act separates the provider (who builds or places the system on the market) from the deployer (the institution using it). Many obligations, including practical human oversight and appropriate use, fall on the deployer — the university — not only on the vendor. This is why "our vendor is compliant" is never a complete answer: even a perfectly built tool can be deployed in a way that creates risk, for instance by wiring its output straight into staff decisions with no human in between. Getting deployment right — who reads the output, what it may and may not inform, and how oversight is exercised — is the institution’s own responsibility, and no procurement checkbox transfers it away.

How Koji fits

Koji for Education is designed so that human oversight is possible by construction, not bolted on. Every AI-generated theme in Koji is traceable to the underlying student responses, so a committee can drill from a summary back to the verbatim evidence and exercise genuine Article-14-style oversight rather than deferring to a black box. Moderation and analysis are standardised, bias-aware and disclosed to respondents, supporting Article 50 transparency and mitigating — never eliminating — the biases of both human and automated reading. Koji reports feedback as evidence for human decision-makers, and our guidance is explicit that course-evaluation output should inform teaching improvement and be handled with great caution before it ever touches an individual staff decision — the boundary that determines your Annex III(4) exposure. Data handling is GDPR/AVG-compliant and EU-appropriate throughout. Teams running wider research on the shared AI interview engine can see the same design principles at koji.so.

Getting the classification right — and building for oversight either way — is the difference between a defensible AI evaluation programme and a compliance surprise. See how Koji for Education approaches AI course evaluation with traceability, transparency and human oversight built in. This article is general information, not legal advice; confirm your own classification with counsel.