New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Graduate outcomes11 min read

Your Medical Course Evaluation Measures Satisfaction. WFME and Programmatic Assessment Ask for Something Else.

Medical education has moved to competence-based, outcomes-oriented quality assurance — WFME accreditation and programmatic assessment. Yet many medical schools still run a happy-sheet as their course evaluation. Here is why satisfaction alone does not meet the standard, and what does.

Koji Education Team

Product · August 6, 2026

Bottom line up front: Medical education has spent two decades moving toward competence-based, outcomes-oriented evaluation — driven by the World Federation for Medical Education (WFME) accreditation standards and by the theory of programmatic assessment. Both frameworks treat a single satisfaction score as, at best, one small and weak data point. Yet many medical schools still run an end-of-rotation happy-sheet as their "course evaluation" and report the mean. That instrument does not meet what WFME asks for, and it cannot see what a clinical programme most needs to know: whether the learning environment is actually producing competent, safe practitioners.

For medical educators, the stakes are unusually concrete. Since the ECFMG accreditation requirement took effect in 2024, physicians seeking ECFMG certification must have graduated from a medical school accredited by an agency whose process uses criteria comparable to WFME standards (WFME/ECFMG). Programme evaluation is not a box to tick; it is part of what makes a school's award portable.

What WFME actually requires of programme evaluation

The WFME Basic Medical Education Global Standards (2020 revision) devote an entire domain — Area 7, Programme Evaluation — to how a school monitors and improves its programme. Read closely, it asks for far more than student contentment. A compliant programme-evaluation system must establish a mechanism for programme monitoring; it must seek and act on feedback from both teachers and students; and crucially, it must relate that feedback to student performance and outcome data and involve a wide range of stakeholders. Satisfaction is one strand, deliberately triangulated against evidence of what students can actually do.

That is a fundamentally different design from a Likert survey averaged into a number. WFME is asking whether the programme works — measured against competence outcomes — not whether the cohort enjoyed the rotation.

Programmatic assessment makes the point sharper

The dominant assessment philosophy in modern medical education, programmatic assessment as developed by van der Vleuten and Schuwirth, is built on a single organising idea: no individual data point carries a high-stakes decision alone. Multiple low-stakes observations are collected longitudinally, triangulated across methods, used principally to drive feedback, and only aggregated by a competence committee for consequential judgements.

Apply that logic to course evaluation and the conclusion is immediate. A student-satisfaction rating is exactly the kind of single, low-stakes data point programmatic assessment refuses to over-read. Treating "4.2 out of 5 on this clerkship" as a verdict on teaching quality commits, at the programme level, the same error the discipline has spent years training its assessors to avoid at the individual level. Satisfaction belongs in the portfolio of evidence — as one input, triangulated — not on the cover as the headline.

What satisfaction can and cannot see in a clinical setting

Place a happy-sheet against Miller's pyramid — knows, knows how, shows how, does — and it sits below the bottom rung. It does not measure knowledge, competence, or performance; it measures reaction, which is Kirkpatrick's first and weakest level. This is the same gap between direct and indirect measures of learning we have written about across disciplines, and it is precisely what assurance-of-learning in business schools and EuroPsy in psychology demand you close.

But there is a subtler point that rescues course evaluation from irrelevance. Students genuinely cannot rate their own competence reliably — self-assessment is weak, as the Dunning-Kruger problem shows. What they can report on, and often are the only witnesses to, is the clinical learning environment: the quality of supervision, whether feedback actually happened, psychological safety on the ward, whether they were treated as a learner or free labour, whether mistreatment occurred. These are process variables that WFME cares about and that predict whether the competence outcomes will follow. The right question is not "did you like the rotation?" but "what did your supervision, feedback, and workload actually look like?" — and that is a question a number cannot answer.

The counterargument, taken seriously

"Students are not qualified to evaluate clinical teaching, so why ask them at all?" This conflates two things. They are not qualified to certify their own competence or to judge the scientific accuracy of instruction — correct. They are, however, the primary and often sole observers of the learning environment, and their reports on supervision quality, feedback frequency, and mistreatment are valid and actionable. WFME requires teacher and student feedback for exactly this reason. The answer to unreliable self-rated learning is not silence; it is asking students the questions they are positioned to answer, and triangulating with performance data for the rest.

"We already collect outcome data, so the survey is just a courtesy." Then design it to earn its place. A survey that only yields a satisfaction mean adds noise; a conversational evaluation that surfaces why a placement is failing students adds a strand of evidence your outcome data cannot supply on its own.

What competence-aligned course evaluation looks like

The move is from a rating to an account. Koji for Education runs AI-moderated conversational interviews that probe beyond the number — following up on a vague comment about "not much supervision" until it becomes a specific, actionable description of the clinical environment. Automatic thematic analysis turns hundreds of such conversations into programme-level themes (supervision, feedback culture, psychological safety, workload), with every theme traceable to the quotes behind it. That output is designed to be exactly what WFME Area 7 asks for: structured teacher and student feedback that can be triangulated against performance data and put in front of a programme-evaluation or competence committee — not a happy-sheet mean masquerading as quality assurance.

The consistent AI moderation matters in a clinical context, where placements are dispersed across many sites and human survey administration is uneven. And because the same interview engine powers the main Koji platform, education teams that also run wider stakeholder or workforce research get one method across both.

Medical education already decided, at the level of individual assessment, that no single data point should carry a high-stakes decision. Programme evaluation deserves the same rigour. Ask what students are positioned to tell you, gather it at depth, triangulate it with what they can actually do — and stop reporting a satisfaction average as if it were evidence of a competent graduate.

Frequently asked questions

Does WFME accreditation require more than a student satisfaction survey? Yes. WFME Basic Medical Education standards (Area 7, Programme Evaluation) require a monitoring mechanism, feedback from both teachers and students, and that this be related to student performance and outcome data with broad stakeholder involvement. Satisfaction alone does not meet it.

Why does the ECFMG 2024 requirement matter for programme evaluation? From 2024, applicants for ECFMG certification must have graduated from a medical school accredited by an agency using criteria comparable to WFME standards — so a WFME-aligned programme-evaluation system affects the international portability of a school's degree.

What is programmatic assessment and how does it relate to course evaluation? Programmatic assessment collects many low-stakes data points, triangulated over time, so that no single point drives a high-stakes decision. Applied to course evaluation, it means a satisfaction rating is one input among many, never the verdict.

Can medical students validly evaluate their teaching? They cannot reliably rate their own competence, but they are the primary observers of the learning environment — supervision, feedback, psychological safety, workload, mistreatment — which they can report validly and which predicts competence outcomes.

What should a clinical course evaluation ask instead of satisfaction? Questions about the actual learning process: how much supervision occurred, whether usable feedback was given, whether the environment was psychologically safe, and how workload affected learning — not simply whether students liked the rotation.

How does Koji support WFME-aligned evaluation? Koji uses AI-moderated conversational interviews and thematic analysis to produce structured, quote-traceable student and teacher feedback on the learning environment, designed to be triangulated with performance data for a programme-evaluation or competence committee.