New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Regulation & compliance10 min read

The AI Act Article Everyone Skips: Accuracy, Robustness, and the Feedback Loop Inside Your Feedback Tool

Most AI Act coverage is about literacy, transparency and human oversight. Article 15 is the substantive-quality article — and its feedback-loop clause describes almost exactly what an AI that themes course-evaluation feedback can do to itself.

Koji Education Team

Product · August 25, 2026

Bottom line up front: If your university uses AI to theme or score student course-evaluation feedback, and those outputs inform decisions about teaching staff, the system is plausibly high-risk under the EU AI Act's employment limb — and Article 15 then imposes substantive technical duties most coverage ignores: declared accuracy metrics in the instructions for use, robustness and redundancy, protection against feedback loops in systems that keep learning, and cybersecurity against data and model poisoning. These duties fall on the provider, but as a deployer you are chained to them through Article 26 — and the feedback-loop clause describes, almost word for word, a failure mode a feedback-analysis tool can create on its own.

Why course-evaluation AI lands in the employment limb, not the education limb

The instinctive classification is wrong. When people hear "AI in education," they reach for Annex III(3) — the education limb — which covers AI used to evaluate learning outcomes and steer the learning of students. But an AI that themes or scores course feedback is not grading learners. If its outputs feed decisions about promotion, contract renewal, task allocation or performance review of teaching staff, it falls under Annex III(4)(b): systems used "to monitor and evaluate the performance and behaviour of persons" in work relationships. The data subject is the lecturer, not the student — the same reframing that puts automated faculty decisions under GDPR Article 22. Get the limb right and the high-risk obligations, including Article 15, come into view.

(If your use is purely formative — feedback goes only to the individual teacher for their own development and never into a personnel decision — you may fall outside the high-risk trigger. The classification turns on what the output is used for, which is exactly why the high-risk analysis for student-feedback systems has to be done deliberately, not assumed.)

What Article 15 actually requires

Article 15 is the AI Act's accuracy, robustness and cybersecurity article — the substantive-quality provision that the popular focus on literacy, transparency and human oversight tends to skip. Four duties matter here:

  1. Appropriate accuracy, robustness and cybersecurity, consistent across the lifecycle (Art 15(1)). Not a one-off validation — sustained performance.
  2. Declared accuracy metrics in the instructions for use (Art 15(3)). The provider must state the levels of accuracy and the relevant accuracy metrics. This is the clause a procurement team can hold a vendor to.
  3. Robustness and redundancy (Art 15(4)). Resilience to errors, faults and inconsistencies, achieved through technical and organisational measures and, where appropriate, redundancy such as backup or fail-safe plans.
  4. Cybersecurity against AI-specific attacks (Art 15(5)). Measures against data poisoning, model poisoning, adversarial examples and confidentiality attacks.

Note that Article 15(2) only obliges the Commission to encourage the development of benchmarks and measurement methodologies — it does not fix a metric. As of now no harmonised accuracy benchmark for this kind of system is established, which means "appropriate accuracy" for thematic analysis of free-text feedback is genuinely contestable. That is a reason to demand the provider's declared metrics under 15(3), not to assume a settled standard exists.

The feedback-loop clause describes your tool's own failure mode

Article 15(4) contains a clause written as if for this exact use case. High-risk systems that continue to learn after deployment must, in the Regulation's words, "eliminate or reduce as far as possible the risk of possibly biased outputs influencing input for future operations (feedback loops)," with those loops "duly addressed with appropriate mitigation measures."

Consider the mechanism. An AI themes course feedback and surfaces, say, "assessment clarity" as a recurring concern. Those outputs shape which teaching gets flagged, praised or rewarded. Rewarded behaviour changes how staff teach and how they solicit feedback. The next round of feedback — now shaped by the model's own earlier outputs — becomes training or reference data for the same continuously-learning model. Left unaddressed, the system starts confirming its own priors: the themes it amplified last cycle look more prevalent this cycle, and documented student-evaluation biases get laundered into an authoritative-looking trend. This is precisely the feedback loop 15(4) targets, and it is not hypothetical for any tool that learns from the data its own outputs influence.

The practical defences are unglamorous but real: don't let a production model silently retrain on post-deployment feedback without controls; keep a stable, human-audited thematic taxonomy rather than one the model drifts on its own; and separate "surfacing a theme" from "acting on it" so the model's output is evidence for a human decision, not an input to its own next training set — the provenance-and-traceability discipline in regulatory form.

Whose duty is it — and the honest deflation

Article 15 duties fall on the provider of the high-risk system, via Article 16. So the honest answer to "is this our problem?" is: not primarily. But two things keep the deployer on the hook. First, Article 26(1) requires deployers to use the system in accordance with the instructions for use — which means you are bound to the provider's declared accuracy conditions and cannot use the tool outside them. Second, if you fine-tune the model on your own data or otherwise substantially modify it, you can become a provider yourself and inherit the Article 15 obligations directly.

The realistic value for a university is therefore procurement leverage: require the Article 15(3) accuracy metrics in writing before you buy, ask how the vendor addresses feedback loops under 15(4), and confirm cybersecurity measures under 15(5). Breaches of provider obligations sit in the middle penalty tier — up to EUR 15 million or 3% of worldwide annual turnover — so a serious vendor should already have answers.

On timing: the standalone high-risk obligations were due to apply from 2 August 2026, but the Digital Omnibus (Regulation (EU) 2026/1744, in force 27 July 2026) postponed Annex III standalone high-risk application to 2 December 2027. Article 15 duties on these systems move with that date. More runway is not a reprieve — the systems being procured now will still have to meet these requirements, so the sensible posture is to build the governance into procurement today.

But isn't this just a provider problem we can ignore?

The strongest counterargument is that Article 15 is a manufacturer's duty, so a university that merely buys a tool has nothing to do. That is half right, and worth saying plainly rather than inflating the risk. But it fails on three points. You cannot outsource the classification of your own use — whether your deployment hits Annex III(4) depends on what you do with the outputs. You are contractually and legally tied to the provider's declared accuracy conditions through Article 26. And the feedback-loop risk is created partly by your operating choices — whether you let the model retrain on data its outputs influenced. "The vendor is responsible" is not a governance plan; it is a procurement question you have not yet asked.

A second, fairer objection: accuracy for thematic analysis is not well-defined, so demanding "accuracy metrics" is naive. True in part — 15(2) benchmarks are not established. But that argues for transparency about how the vendor measures and validates its outputs, human verification of AI-surfaced themes, and against treating an AI summary as ground truth — not for ignoring the article. Uncertainty about the metric is a reason for more scrutiny, not less.

Where Koji fits

Koji is an AI-native course-evaluation platform built with these duties in mind. Its thematic analysis is designed to be auditable and human-verifiable — themes are evidence for a human decision, not an unaccountable verdict — which supports the separation of surfacing from acting that keeps feedback loops in check. Koji uses standardized, bias-aware AI moderation rather than a model that silently retrains on the data its own outputs shaped, and provides provenance and traceability from a theme back to the underlying responses. Its six structured question types, quality scoring, formative mid-cycle collection and closing-the-loop action tracking give you a documented, defensible evidence trail, and data handling is GDPR/AVG-compliant and EU-appropriate.

Koji does not make you compliant on its own, and it does not eliminate bias — no tool does. It mitigates the feedback-loop and accuracy risks Article 15 is concerned with, and surfaces the provenance you need to keep a human in the loop. Teams running wider research beyond course evaluation use the same AI interview engine on the main Koji platform.

Before you procure an AI tool that touches staff decisions, ask the Article 15 questions. Explore Koji for Education to see evaluation built for auditable, human-verifiable feedback analysis.

Frequently asked questions

Does AI that analyses course feedback count as high-risk under the AI Act? It can. If the AI's outputs inform decisions about teaching staff — promotion, contract renewal, performance review — it plausibly falls under Annex III(4)(b), the employment limb (monitoring and evaluating the performance of persons), rather than the education limb about learners. Purely formative use that never feeds a personnel decision may fall outside the high-risk trigger.

What does Article 15 require? Appropriate accuracy, robustness and cybersecurity, sustained across the lifecycle (15(1)); declared accuracy metrics in the instructions for use (15(3)); resilience to errors with technical redundancy (15(4)); protection against feedback loops in systems that keep learning (15(4)); and defences against data poisoning, model poisoning and adversarial attacks (15(5)).

What is the feedback-loop clause and why does it matter for feedback tools? Article 15(4) requires high-risk systems that keep learning after deployment to eliminate or reduce, as far as possible, the risk of biased outputs influencing future inputs (feedback loops). A tool that themes feedback, shapes which teaching is rewarded, then learns from the next round of feedback its own outputs influenced, can confirm its own priors — exactly the failure this clause targets.

Whose responsibility is Article 15 — the vendor's or ours? Primarily the provider's, via Article 16. But as a deployer you must use the system per the instructions for use (Article 26), you cannot outsource the classification of your own use, and if you fine-tune the model you may become a provider yourself. In practice, treat Article 15(3) accuracy metrics as a procurement requirement.

Has the timeline changed? Yes. Standalone Annex III high-risk obligations were due from 2 August 2026, but the Digital Omnibus (Regulation (EU) 2026/1744) postponed application to 2 December 2027. Article 15 duties on these systems move with that date, so the extra time is runway to build governance into procurement, not a reprieve.