The OECD Tried to Measure What Graduates Actually Learn. What AHELO Teaches Course Evaluation
Between 2010 and 2013 the OECD ran AHELO — an ambitious attempt to directly measure higher-education learning outcomes across countries. It reached 23,000 students in 17 countries and was then quietly shelved. Its failure is the most instructive thing that ever happened to the case for course evaluation.
Koji Education Team
Product ·
Bottom line: The most ambitious attempt ever made to directly measure what university graduates learn — the OECD's AHELO project — was declared feasible, then abandoned. Understanding why it stalled tells you something uncomfortable and useful: if a well-funded international consortium could not build a comparable, cross-institutional measure of learning outcomes, then your course evaluation certainly is not one either. That is not a reason to give up on course evaluation. It is a reason to be precise about what it can and cannot claim — and to stop pretending a satisfaction survey is evidence of learning.
The experiment most quality-assurance offices have never heard of
For a decade, policymakers worried about the same thing employers and accreditors still worry about: rankings measure inputs and reputation (research output, selectivity, spending per student), not whether students actually learn anything. So the OECD attempted something audacious. Its Assessment of Higher Education Learning Outcomes (AHELO) feasibility study set out to measure learning outcomes directly — a "PISA for higher education" — and to do so comparably across very different national systems.
This was not a think-piece. Fieldwork took place in 2012 and, according to the OECD and the Australian Council for Educational Research (ACER, which ran the assessment design), involved roughly 23,000 students and 5,000 faculty across about 250 higher-education institutions in 17 countries. It tested three strands: generic skills (critical thinking, analytic reasoning, problem-solving, written communication), plus discipline-specific strands in economics and civil engineering.
And then it stopped. After the feasibility study reported in 2012–13, OECD member countries declined to commit to a full-scale AHELO. The project did not proceed. For a body with the OECD's convening power and budget, that is a remarkable outcome — and a deeply informative one.
Why AHELO stalled — and why it matters for your evaluation form
The reasons AHELO was shelved are not gossip; they are the same measurement problems that sit, unacknowledged, inside every course-evaluation programme.
1. Comparability is brutally hard. AHELO's central promise was cross-national, cross-institutional comparison. But defining a "learning outcome" that means the same thing in a Japanese engineering faculty and a Mexican one — and building an instrument that measures it equivalently — proved enormously contested. This is the same measurement-invariance problem that undermines your own benchmarking when you compare a 4.1 in Engineering to a 4.4 in History. If the OECD could not solve comparability with a purpose-built instrument, a Likert item certainly does not solve it by accident.
2. Direct measurement is expensive and intrusive. Actually testing what students can do — as AHELO tried to — requires real assessment instruments, student time, and institutional buy-in on a scale that turned out to be politically and financially unsustainable. This is precisely why institutions fall back on indirect measures like satisfaction surveys: they are cheap. But cheap indirect measures buy affordability by giving up the very thing AHELO was chasing — evidence of actual learning.
3. Stakes distort participation. Institutions were wary of a comparison that might rank them unflatteringly. Low-stakes for students meant questionable effort; high-stakes for institutions meant defensiveness. The incentive tangle that dogged AHELO is a cousin of the response-rate and motivation problems that dog every voluntary course evaluation.
The signalling shadow over the whole enterprise
There is a deeper reason measuring "what graduates learned" is so slippery, and it comes from labour economics. The human-capital account of education (associated with Becker) says a degree raises productivity by building skills. The rival signalling account (Spence's 1973 model of job-market signalling, for which he shared the 2001 Nobel Memorial Prize) says a degree can raise wages even if it taught nothing, by credibly signalling pre-existing ability to employers.
Both are partly true, and that is the point: graduate salaries and employment rates — the outcome data universities love to cite — are contaminated by signalling. A programme can produce excellent employment outcomes while adding modest skill, because it selected able students who would have done well anyway. This is why graduate-destination statistics are a weak proxy for teaching quality, and why AHELO tried to measure skill directly rather than infer it from wages. AHELO's difficulty and destination data's ambiguity are two sides of the same coin: learning is genuinely hard to observe.
So what should course evaluation honestly claim?
If direct measurement defeated the OECD and outcome data is confounded by signalling, where does that leave the humble module survey? In a more modest but still valuable place.
Course evaluation cannot measure learning outcomes, and it should stop implying that it does. What it can do is measure the proximal conditions and behaviours that the evidence links to learning: whether students found the workload manageable, whether feedback was timely and usable, whether the teaching prompted them to engage actively, whether assessment aligned with what was taught. These are not learning itself — they are the observable levers that produce it. A well-designed evaluation is an instrument for improving those levers, not a graduate-outcome oracle.
That reframing is liberating. It frees the module survey from a burden it was never going to carry, and points it at questions it can actually answer well.
Consider a concrete example. A programme director who asks students "did this module improve your critical thinking?" is asking for a self-report of an outcome — precisely the kind of judgement AHELO showed is hard to measure and, from the self-assessment literature, one students are poorly calibrated to give. But a director who asks "which specific activities in this module forced you to defend a claim with evidence, and where did that break down?" is asking about an observable behaviour the student genuinely witnessed. The second question is answerable, actionable, and honest about its own altitude. AHELO's lesson, translated into a survey design principle, is simply this: ask about the visible mechanism, not the invisible outcome.
But doesn't AHELO's failure prove all outcome measurement is futile?
The strongest counterargument cuts the other way: if even the OECD gave up, why measure anything beyond satisfaction? Three responses.
First, AHELO failed as a global comparability project, not as a measurement idea. Within a single programme, where the curriculum and context are shared, measuring learning-related behaviours and even direct competence (through embedded assessment, capstones, portfolios) is far more tractable than across 17 countries. The lesson is "comparability degrades with distance," not "measurement is impossible."
Second, AHELO's partial success is often forgotten: the feasibility study did show it was technically possible to build and administer common instruments and get institutions to participate. What failed was the political will to fund a permanent version, not the psychometrics outright.
Third, and most importantly, the honest conclusion is not "measure nothing" but "measure the right thing at the right altitude." Triangulate: combine indirect student feedback with direct evidence (assessment data, work-integrated-learning performance, graduate tracer studies), and be explicit about which is which. No single instrument, AHELO included, was ever going to be sufficient alone.
Where Koji fits: measuring the levers, not faking the outcome
Koji for Education is built for exactly the modest-but-rigorous role the AHELO story argues for. It does not claim to measure learning outcomes — it surfaces the proximal, improvable conditions that produce them.
- Instead of a single satisfaction number, Koji runs AI-moderated conversational interviews that probe why something worked or did not — turning "I'd recommend this course" into specific, actionable insight about workload, feedback, and alignment.
- Six structured question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) plus automatic thematic analysis let you ask directly about the learning-linked behaviours that matter, then read the open text at programme scale.
- Programme- and institution-level reporting supports the triangulation the AHELO lesson demands — feeding student-experience evidence alongside your direct assessment and destination data rather than pretending to replace them.
- Bias-aware, standardised AI moderation and GDPR/AVG-compliant handling keep the indirect evidence as clean as indirect evidence can be.
The same conversational engine powers the main Koji platform for organisations measuring hard-to-observe outcomes in customer and user research — the same discipline of measuring the levers you can see.
Koji will not hand you a graduate-learning score. Nothing honestly can. What it gives you is a rigorous read on the conditions that produce learning — which, after AHELO, is exactly the claim worth making.
Measure what you can actually improve. Explore Koji for Education to move your programme from satisfaction scores toward evidence about the conditions that produce learning.