New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Methodology10 min read

Stop Asking Whether Students Were Satisfied. Ask Whether They Met the Standard

Generic satisfaction items ("the course was well organised") float free of any disciplinary standard. Mapping evaluation onto external reference points — QAA Subject Benchmark Statements, the Tuning competences, the EQF — turns feedback into evidence a quality office and an accreditor can actually use.

Koji Education Team

Product ·

Short answer: Most course-evaluation instruments ask whether students were satisfied — a construct that, on the best available evidence, barely correlates with how much they learned. A more defensible design references evaluation to the external standards a discipline already publishes: QAA Subject Benchmark Statements in the UK, the Tuning competences across Europe, and the learning-outcomes language of the European Qualifications Framework. Done honestly, outcome-referenced evaluation does not claim students can self-certify their competence; it channels student perceptions toward the specific outcomes a programme promised, so the results feed the quality cycle the ESG 2015 require rather than floating free of it.

The problem with "the course was well organised"

The standard evaluation form is a grid of generic satisfaction items scored on a Likert scale. They are cheap, comparable across programmes, and easy to benchmark — and that is their entire appeal. What they are not is evidence of learning.

The uncomfortable finding here is now hard to dismiss. Re-analysing decades of "multisection validity" studies, Uttl, White and Gonzalez concluded that student evaluation ratings and student learning are essentially unrelated once you account for small-sample artefacts and publication bias (Studies in Educational Evaluation, 2017). A follow-up meta-analysis restricted to conflict-of-interest-free studies found the correlation between ratings and learning was near zero, around r = .06 (Uttl, Cnudde & White, PeerJ, 2019). If satisfaction ratings track learning that weakly, a form built entirely from satisfaction items cannot tell a programme director whether the course did its job.

What "referencing to a standard" actually means

Disciplines in Europe do not lack standards — they lack the habit of pointing evaluation at them. Three reference systems are already in place.

QAA Subject Benchmark Statements. The UK's Quality Assurance Agency publishes statements that "describe the nature of study and the academic standards expected of graduates in specific subject areas" — what a graduate "might reasonably be expected to know, do and understand" by the end of their studies. The current suite covers roughly 47 subject areas, most revised between 2022 and 2025 (QAA Subject Benchmark Statements). They are explicit, discipline-specific, and public.

Tuning Educational Structures in Europe. Tuning, a European Commission-funded project coordinated by the University of Deusto and the University of Groningen, set out to identify "points of reference for generic and subject-specific competences" of first- and second-cycle graduates. Its definition of a learning outcome is the one worth adopting: "what a learner knows or is able to demonstrate after the completion of a learning process." Crucially, Tuning insists competences are "points of reference for curriculum design and evaluation, not straightjackets" (Tuning executive summary, EHEA). Evaluation is written into the model from the start.

The European Qualifications Framework. The EQF is an eight-level, learning-outcomes-based framework defining each level by knowledge, skills, and responsibility and autonomy — a common vocabulary, established by a 2008 Recommendation and revised in 2017, for describing what a qualification means across borders.

The pedagogy that ties these together is John Biggs's constructive alignment (Higher Education, 1996): intended learning outcomes, teaching activities, and assessment should all point at the same target. If your teaching and assessment are aligned to disciplinary outcomes but your evaluation asks only whether students liked the room, you have left the last link in the chain unaligned.

What outcome-referenced evaluation looks like in practice

The shift is concrete. Instead of "The course was intellectually stimulating (agree/disagree)", an outcome-referenced item is anchored to a specific competence the benchmark names — for example, in a benchmark that expects graduates to critically evaluate primary evidence: "Describe a point in this module where you had to weigh conflicting evidence. What did you do, and where did you feel unsure?"

Two things change. First, the item names the outcome the programme committed to, so a low result points at a specific curricular gap rather than a diffuse mood. Second, the phrasing invites evidence of doing, not a global satisfaction verdict. This is the difference between evaluation that decorates a report and evaluation that survives contact with an accreditor.

Why this maps cleanly onto European quality assurance

The Standards and Guidelines for Quality Assurance in the European Higher Education Area (ESG 2015) practically ask for it. Standard 1.2 requires that programmes be designed "to meet the objectives set for them, including the intended learning outcomes." Standard 1.3 requires student-centred learning and assessment. Standard 1.7 requires institutions to "collect, analyse and use relevant information for the effective management of their programmes." Standard 1.9 requires on-going monitoring and periodic review so that programmes "achieve the objectives set for them." Evaluation referenced to learning outcomes feeds every one of these; a satisfaction grid feeds none of them well. We have argued the same from the constructive-alignment and Kirkpatrick angles — that most evaluation never escapes the "reaction" level.

But doesn't this just move the self-report problem around? The strongest objection

Here is the objection that should worry anyone proposing this, and it is serious. Asking a student "did you attain outcome X?" is still a self-report. The very literature that undermines generic satisfaction items — Uttl and colleagues — also undermines confidence that students can accurately judge their own competence. Perceived learning and actual learning correlate weakly. Relabelling "the course was well organised" as "I can now critically evaluate primary sources" may not escape the self-report ceiling; it may just repaint it. Genuine outcome measurement requires direct evidence — assessed work, capstones, portfolio review — which is more costly and arguably a different instrument entirely. We have made this exact distinction in direct versus indirect measures of learning.

Two further objections deserve airing. Disciplinary standards are deliberately non-prescriptive — Tuning calls competences reference points "not straightjackets", and there is genuine disagreement within any field about which outcomes matter and how to weight them, so an outcome-referenced instrument encodes a contestable view of the discipline. And bespoke, competence-mapped instruments fragment by subject, breaking the cross-programme comparability and cheap benchmarking that generic items provide.

These are real, and they set the honest boundary of the claim. Outcome-referenced evaluation is not a self-report substitute for measuring learning. It is a better-aimed indirect measure whose job is to route student perception toward specific outcomes and to feed direct evidence into the quality cycle — a complement to assessed work and triangulated evidence, never a replacement for them. Positioned that way, it improves on satisfaction items without over-claiming.

Starting without boiling the ocean

You do not need to rebuild every instrument at once, and the comparability objection above is a reason to be incremental rather than a reason to do nothing. A pragmatic first step keeps the short generic core that gives you longitudinal, cross-programme comparability, then adds two or three outcome-referenced prompts drawn from the relevant Subject Benchmark Statement or the programme's own validated learning outcomes. Pilot them in a single programme; check whether the free-text responses genuinely localise problems to specific outcomes rather than restating global satisfaction; and expand only where they earn their place. This preserves the institution-wide dashboard while giving programme teams the diagnostic detail a satisfaction grid can never produce.

Where conversational evaluation fits

Outcome-referenced items are harder to run well on a static form. A single fixed question mapped to a competence gets you a number and a shrug; the value is in the follow-up — the example, the moment of difficulty, the specific gap. That is exactly what a fixed Likert grid cannot do.

Koji for Education is built for this. Its six structured question types — including open_ended and scale — let you anchor each prompt to a named learning outcome or benchmark competence, and its AI moderation probes for the evidence behind the rating: "you said you can apply this — walk me through where." Automatic thematic analysis then maps free-text responses back to the outcomes and to ESG monitoring categories, and closing-the-loop action tracking documents what changed as a result, which is precisely the evidence Standard 1.9 reviews demand. Formative, mid-cycle collection means the outcome signal arrives while the cohort can still benefit. Koji does not claim students can self-certify competence — it surfaces perception against the standard and channels it into the direct-evidence cycle where the real judgement is made. The same interview engine runs general research at koji.so. Legacy tools that only average a satisfaction Likert leave the discipline's own published standards sitting unused a click away. Point the evaluation at them.