New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Graduate outcomes9 min read

Beyond Satisfaction: Measuring Skill Development, Not Just Student Happiness

Most course evaluations measure whether students enjoyed the course. But satisfaction barely correlates with how much they learned — and tells employers and accreditors almost nothing about skills gained. The case for evaluating competency development instead, and how to do it.

Koji Education Team

Product · June 11, 2026

The short answer: The standard course evaluation measures satisfaction — did students like the course, the instructor, the workload. But satisfaction is a weak proxy for what universities actually exist to produce: capability. The best meta-analytic evidence finds student-evaluation ratings explain at most about 1% of the variance in measured learning. If programmes want evidence that holds up to accreditors and means something to employers, they need to evaluate skill and competency development — what students can now do that they could not before — and align that evaluation with the intended learning outcomes the course was designed around. This is harder than a happiness score, but it is the only kind of evaluation that answers the question that matters.

The uncomfortable finding: satisfaction is not learning

For decades, the implicit assumption behind student evaluations was that highly-rated teaching and effective teaching are roughly the same thing. The evidence does not support it. Uttl, White and Gonzalez's 2017 meta-analysis re-examined the multisection studies that had been used to validate student ratings and found the earlier optimistic correlations were largely an artefact of small samples and publication bias. Once those were corrected, student-evaluation ratings explained at most 1% of the variance in student learning, with a best estimate of the correlation around r = 0.08 — down from the r = 0.43 reported in Cohen's influential earlier work (Uttl, White & Gonzalez, 2017, Studies in Educational Evaluation).

Read that carefully. It does not say evaluations are worthless; we have argued they are a valid measure of the student experience but not of learning. It says that how satisfied students are is nearly disconnected from how much they learned. A course can delight students and teach them little; a demanding course can frustrate students and develop them enormously. If your evaluation only captures satisfaction, you cannot tell these apart — and you may reward the wrong one.

Why satisfaction became the default

Satisfaction won because it is easy to ask and easy to count. "Overall, how satisfied were you with this course?" produces a tidy number you can average, benchmark, and put on a dashboard. Skill development is messier: it requires knowing what the course was supposed to develop, and asking students to reflect on their own capability against that standard. The path of least resistance led institutions to measure the thing that was convenient rather than the thing that was meaningful — a familiar failure mode we also see in the misuse of single-number metrics like Net Promoter Score for course quality.

Anchoring evaluation in intended learning outcomes

The alternative starts with a concept most quality officers already know: constructive alignment. John Biggs's principle holds that teaching activities and assessment should be designed to address the course's intended learning outcomes (ILOs) — what a student should be able to do on completion (Biggs, constructive alignment). The insight for evaluation is simple but rarely applied: evaluation should be aligned to the same outcomes. If a module's ILO is "students can critically appraise a research design," the evaluation should ask about that capability — whether students feel able to do it, where they got stuck, what helped — not merely whether they enjoyed the seminars.

This reframing changes the questions entirely. Instead of "Rate the instructor's clarity," you ask "Which of the course's intended skills do you now feel confident applying, and which still feel shaky?" Instead of a global satisfaction score, you get a competency map: where the programme is developing capability and where it is leaving gaps. That is evidence a programme director can act on and an accreditor can use.

The employer and accreditation payoff

This is also where evaluation connects to the outside world. Employers do not care whether last year's cohort enjoyed module 3; they care whether graduates can do the things the degree claims to produce — and the skills gap between what programmes teach and what employers need is precisely a competency question, not a satisfaction one. Accreditation frameworks across Europe increasingly ask for evidence of outcome achievement and of closing the loop between feedback and improvement — not evidence that students were content. A satisfaction-only evaluation cannot supply this; a competency-aligned one can, and it sits naturally alongside the other outcome evidence (assessment results, graduate destinations) in a triangulated picture, which is the responsible way to evaluate teaching.

"But isn't self-reported skill just satisfaction in disguise?"

This is the serious objection, and it deserves an honest answer. Self-reported competence is imperfect: students can be overconfident, and a "halo" of enjoyment can inflate self-ratings of skill. Self-report is not a substitute for direct measurement of learning through assessment. But it is not the same as satisfaction either. Satisfaction asks "did you like it?"; self-reported capability asks "what can you now do, and where did you struggle?" — a question that yields specific, actionable, outcome-anchored information even when the confidence calibration is imperfect. The methodological discipline is to treat self-reported skill development as one triangulated source — read alongside assessment data and, where available, employer and graduate-outcome data — never as the sole proof of learning. Used that way, it adds a dimension satisfaction scores simply do not have.

How Koji supports competency-focused evaluation

Moving beyond satisfaction requires an instrument that can ask outcome-anchored questions and probe the answers — which is hard for a static Likert form. Koji for Education is built for it. Its six structured question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) let programmes map evaluation directly onto intended learning outcomes — for example, ranking which course-defined skills students feel most and least confident in. Its AI-moderated conversational interviews then probe beyond the rating: when a student says a skill still feels shaky, the moderator asks where it broke down and what would have helped, the same calibrated follow-up for every student. Automatic thematic analysis turns those open responses into a competency-gap map tied back to verbatim evidence, and programme- and institution-level reporting lets you see development patterns across a whole degree rather than module by module. To be precise: Koji surfaces richer evidence of perceived skill development; it does not replace direct assessment of learning, and we would never claim a feedback tool measures competence the way an exam does. It complements that evidence with the student voice on capability. The same conversational engine powers the main Koji platform for customer and user research, where the gap between "satisfied" and "able to succeed" is just as decisive.

What competency questions actually look like

The shift from satisfaction to capability is concrete, not abstract, and it shows up in the wording. A satisfaction form asks: "Rate the quality of teaching on this module" and "Overall, how satisfied were you with this course?" A competency-aligned evaluation, built from the module's stated outcomes, asks instead:

  • "Which of these course skills do you now feel confident applying independently?" (a ranking or multiple-choice item drawn directly from the intended learning outcomes)
  • "Think of a skill this course was meant to develop that still feels shaky. What was it, and where did it break down for you?" (open-ended, diagnostic)
  • "Compared with the start of the module, how would you describe your ability to [specific outcome]?" (a gain-oriented scale rather than a static rating)

Notice that none of these ask whether the student liked anything. Each ties back to something the course promised to develop, and each yields information a programme can act on: not "students were 4.1/5 satisfied" but "two-thirds feel confident appraising a research design, but most still struggle to justify a sampling strategy — fix the sampling teaching."

The payoff compounds when this is tracked longitudinally across a programme rather than module by module. Asking the same competency questions at multiple points lets a programme director see where capability is genuinely built and where it plateaus or is assumed but never developed — the kind of programme-level evidence that satisfaction scores, which reset every module, can never assemble. This is also what makes the evidence durable for accreditation: a competency map across a degree, triangulated with assessment and graduate outcomes, tells a far more convincing story of outcome achievement than a wall of module satisfaction averages.

The bottom line

Universities exist to develop capability, but most evaluate satisfaction — and the two are nearly uncorrelated. The fix is to anchor evaluation in the same intended learning outcomes that shape the curriculum, ask students what they can now do rather than whether they enjoyed it, and read that self-reported development as one triangulated source alongside assessment and graduate-outcome data. It is more demanding than a happiness score. It is also the only kind of evaluation that answers the question employers, accreditors, and students themselves are actually asking: did this course make me more capable?