New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Sector trends10 min read

Italy Standardised Course Evaluation Nationally. What the ANVUR/OPIS Model Teaches Europe

Italy is a natural experiment in "one questionnaire for everyone": student course evaluation is mandatory by law and the instrument is standardised nationally by ANVUR. It buys comparability and coverage - and pays for it in validity and local relevance. Here is the honest ledger.

Koji Education Team

Product ยท August 18, 2026

Most European countries let each university design its own course-evaluation form. Italy took the opposite path: a single, nationally standardised instrument, collected mandatorily, governed by the national quality agency. That makes Italy the closest thing Europe has to a controlled experiment in centralised student feedback - and a superb case study in what you gain, and what you give up, when you standardise a questionnaire for an entire higher-education system.

The short answer (BLUF): In Italy, collecting students' opinions on courses is mandatory - required of universities under Article 1, paragraph 2 of Law 370/1999 - and the questionnaire itself is standardised nationally, established by ANVUR (the national agency) in 2013 as part of the AVA accreditation system, so results are comparable across institutions. This model delivers three real goods - universal coverage, cross-institution comparability, and protection against instrument-shopping - and carries three real costs - construct compression into a fixed set of Likert items, weak local and disciplinary relevance, and the familiar validity problems of any single-number rating amplified to national scale. The lesson for the rest of Europe is not "copy Italy" or "avoid Italy," but that comparability and validity are in tension, and you should choose your instrument knowing which you are buying.

What the Italian system actually mandates

Two features make Italy distinctive. First, the obligation: the collection of students' opinions is carried out mandatorily by universities for attending students, pursuant to Art. 1, para 2 of Law 370/1999. This is not a nudge or an internal-quality aspiration; it is a statutory requirement, and it flows into the national accreditation regime.

Second, the standardisation: the questionnaire was established by ANVUR in 2013 and, apart from minor changes permitted at local level, is adopted by all Italian universities "in order to allow comparisons at national level." A typical OPIS (Opinioni Studenti - student opinions) instrument asks attending students about study load, course organisation, the lecturer (teaching, availability, flexibility), course materials, and overall interest and satisfaction. The data feed institutional quality assurance and the AVA cycle, where student opinion is treated as "one of the central aspects" of the system.

In short: near-universal coverage of attending students, a common core of items, and a direct line from the classroom questionnaire to national accreditation. Very few systems in the world can say that.

The genuine strengths

It is easy for methodologists to sneer at a standardised national Likert form. That would miss what the design gets right.

Coverage and coordination. A statutory mandate plus a common instrument means the whole system collects feedback, not just the conscientious departments. It closes the equity gap where some students are heard and others never asked.

Comparability. Because the core items are shared, an external reviewer, a national agency, or a prospective student can - in principle - compare like with like across faculties and institutions. Home-grown instruments make this impossible; every university measuring something slightly different is a recipe for incomparable dashboards, a problem we examine in validated instruments versus home-grown forms.

Resistance to instrument-shopping. When each unit picks its own form, there is a quiet incentive to choose flattering questions and easy scales. A nationally fixed core removes that degree of freedom - a real integrity benefit, and a structural defence against the Goodhart dynamics that plague locally-gamed metrics.

The genuine costs

Now the ledger's other side, which a PhD audience will insist on.

Construct compression. A fixed national item set must be generic enough to fit a philology seminar and a chemistry lab, a 12-student masterclass and a 400-student lecture. Generic items buy comparability by sacrificing the specific, actionable detail that makes feedback useful to a particular teacher. The result is data that are comparable and shallow - and shallow feedback rarely changes teaching.

The single-number problem, at scale. Standardising the instrument does not repair the underlying measurement issue: averaging ordinal Likert responses into a mean and treating it as an interval quantity is statistically fraught, as we argue in why averaging Likert scores misleads. Nationalising the form propagates that flaw uniformly. Comparability across institutions is only meaningful if the thing being compared is valid in the first place - and cross-unit benchmarking of raw means imports every confound (class size, discipline, cohort) along with it.

Local and disciplinary relevance. Signature pedagogies - the studio crit, the clinical placement, the problem-based tutorial - are poorly captured by a one-size questionnaire. Italian scholarship on the national instrument has explored exactly these tensions; an exploratory analysis of students' course evaluations in an Italian case study illustrates the interpretive care the OPIS data demand (Students' evaluation of academic courses, 2021, Studies in Educational Evaluation). The "minor changes allowed at local level" are an implicit admission that pure standardisation is too blunt.

The lesson for Europe: comparability and validity pull against each other

The Italian case crystallises a trade-off every quality system faces under the European Standards and Guidelines. Push toward standardisation and you gain comparability, coverage and integrity - but you compress the construct and lose local validity. Push toward local, tailored instruments and you gain relevance and depth - but you lose comparability and open the door to instrument-shopping. Italy sits near the standardisation pole; the UK's institution-by-institution approach and much of the system-accreditation model elsewhere sit closer to the local pole. Neither pole is "correct." The mistake is to adopt one while assuming you also get the other pole's benefits.

"But critics argue..." - the strongest objections

"Standardisation is obviously right - without it you cannot compare anything." Comparability is valuable, but only of a valid measure. A nationally-comparable invalid number is nationally-comparable noise; it can even be worse than local data, because its apparent comparability invites high-stakes cross-institution rankings the instrument cannot support. Comparability is a property to be earned after validity, not a substitute for it.

"Italy proves mandates work - everyone gets surveyed." The mandate does secure coverage, which is a genuine equity win. But a legal obligation to collect is not an obligation to act, and the persistent risk in any accountability-driven system is that feedback is gathered for compliance and filed, not used - the action gap that makes students cynical and depresses future participation.

"So national standardisation is a mistake?" No. For coverage, integrity and system-level steering, a common core is defensible and arguably necessary. The error is treating the standardised core as the whole of course evaluation. The sophisticated position is a layered one: keep a lean, valid, comparable national core for accountability, and pair it with a richer, locally-relevant, depth-oriented layer for genuine improvement.

Where Koji fits

That layered model is exactly what Koji for Education is designed to enable. Koji does not ask institutions to abandon a required national core - the OPIS items, or any statutory instrument, can sit alongside it. What Koji adds is the improvement layer the standardised form structurally cannot provide: an AI-moderated conversational interview that adapts to the discipline and the individual response, probing beyond a generic Likert item to surface why a course worked or failed. Its six structured question types and automatic thematic analysis turn open student voice into comparable themes without flattening it to a single mean, and programme- and institution-level reporting can aggregate that richer evidence for the AVA cycle. Because moderation is standardised by the AI rather than by inconsistent human administration, Koji preserves the comparability benefit Italy prizes while restoring the depth a fixed form sacrifices - and it is built for GDPR/AVG-compliant, EU-appropriate data handling from the ground up. The same conversational engine powers general customer and user research on the main Koji platform.

Koji does not claim to eliminate the tension between comparability and validity - that tension is real and permanent. It claims to let institutions stop choosing one pole and losing the other, by making the standardised core and the improvement layer complementary rather than rival.

Italy standardised course evaluation before almost anyone else. The value of studying it is not to declare the model right or wrong, but to see - unusually clearly - the price of every design choice a quality system makes.

Related reading