New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Graduate outcomes10 min read

Course Evaluation Can't Measure Employability — But It Can Measure What Produces It

Asking students to rate how "employable" a course made them is asking them to judge an outcome they cannot yet see. The fix is not to abandon evaluation but to measure the evidence-based mediators of employability — and to triangulate them with the graduate data that arrives years later.

Koji Education Team

Product ·

Bottom line up front: Graduate employability is real, it matters to accreditors, and programme directors are under growing pressure to evidence it. But end-of-course evaluation is the wrong instrument to measure it directly. Employability is a distal outcome — it depends on the labour market, the economy, and everything that happens after graduation — and students at the end of week twelve simply cannot judge whether a module made them hireable. What course evaluation can do, and do well, is measure the proximal mediators of employability: the well-evidenced conditions inside a course that downstream research links to skill development and graduate success. Measure the mediators, triangulate them with the graduate data that arrives years later, and you get a defensible employability evidence base. Ask a satisfaction survey to do it alone and you get a vanity metric.

Why "rate how employable this made you" is a broken question

Three things go wrong when a course evaluation asks students to score employability directly.

First, the timing is impossible. Employability reveals itself in the job market months or years after a module ends. A student finishing a second-year statistics course has no information about whether it will help them land or hold a job. They can report a feeling about employability; they cannot report the fact.

Second, students are poor judges of their own learning in the moment — and feeling is exactly what they will report. The most striking evidence comes from Deslauriers and colleagues at Harvard (PNAS, 2019), who randomly assigned students to active-learning or passive-lecture versions of identical physics content. Students in active classrooms learned measurably more but felt they had learned less; actual and perceived learning were strongly anticorrelated. If perceived learning inversely tracks real learning, then perceived employability — an even fuzzier self-judgement — is a treacherous target. This is the same dynamic behind the active-learning penalty in course evaluations.

Third, the courses students rate highest are not reliably the ones that build durable skills. Carrell and West's study of random student-instructor assignment at the US Air Force Academy (NBER, 2010; 10,534 students) found that instructors whose students rated them highly and scored well contemporaneously tended to have students who performed worse in follow-on courses — and student evaluations were positive predictors of current-course achievement but poor predictors of later achievement. Uttl, White and Gonzalez's 2017 meta-analysis reinforces the point: once small-sample bias is corrected, SET ratings are essentially unrelated to how much students actually learn. A high satisfaction score is not a proxy for skill formation, let alone employability.

So the direct question fails on timing, on self-report validity, and on construct validity all at once. The honest move is to stop asking it.

Measure the mediators, not the miracle

Employability is distal, but it is not magic. Decades of higher-education research identify proximal features of a learning experience that plausibly cause skill development and, downstream, graduate success. These are things a current student can report on accurately because they describe their actual experience, not a forecast:

  • Authentic, applied assessment — did the work resemble the tasks of professional practice rather than only recall?
  • Feedback quality and feedback uptake — did students receive actionable feedback and get the chance to act on it?
  • Active and collaborative learning — were students doing, building, and working in teams, not only listening?
  • Work-relevant and real-world tasks — projects, cases, clients, placements — the territory of work-integrated learning evaluation.
  • Development of transferable skills — communication, problem-solving, teamwork, self-management.
  • Constructive alignment — coherence between stated learning outcomes, teaching, and assessment.

Each of these is a mediator: a present-tense, observable condition that the literature ties to the outcomes employers care about. And the labour-market case for caring is not abstract. Cedefop's European Skills and Jobs Survey reports that roughly a quarter of EU adult employees have significant skill deficits and that overeducation and skills under-utilisation remain widespread — a persistent mismatch between what graduates bring and what work requires. Programmes that can show they cultivate the mediators of skill are responding to a documented sector problem, not a marketing brief. This reframes the goal from measuring satisfaction to measuring skill-building conditions and connects directly to the graduate employability skills gap.

Then triangulate across time

Mediators are leading indicators; they are not proof. The complete evidence chain runs in three stages, on three different clocks:

  1. In-course evaluation captures the mediators while the experience is fresh and changeable — early enough to act on, via formative collection.
  2. Exit and early-career feedback captures graduates' retrospective judgement once they have some workplace perspective.
  3. Graduate tracer studies capture actual destinations and outcomes years later — the territory of tracer studies in European course evaluation and the employer feedback loop.

The point of triangulation is to test the chain: if your in-course mediators look strong but tracer data shows weak outcomes, the mediator model needs revisiting; if both align, you have a genuinely defensible, accreditation-grade employability narrative rather than a single self-reported number.

"But isn't this just dodging accountability for outcomes?"

The sharpest objection: accreditors and ministries increasingly want outcome evidence — employment rates, graduate earnings — so isn't "measure the mediators instead" a way of dodging the hard question?

No, and the distinction matters. Nobody is proposing to ignore outcome data; tracer studies and employment statistics stay firmly in the framework. The argument is narrower and entirely defensible: don't ask a course evaluation to measure something it structurally cannot, and don't mistake an outcome you can only observe years later for something a current student can rate. Measuring mediators is what makes outcomes improvable — they are the levers a programme team can actually pull this semester, whereas an employment rate is a lagging result you can only watch. Good outcome accountability requires good mediator measurement; the two are complements, not substitutes.

A second fair objection: aren't these "mediators" just proxies that might not actually cause employability? Partly, yes — which is exactly why the model is falsifiable through triangulation. You treat mediators as hypotheses about what produces graduate success and check them against downstream data, refining the set over time. That is more intellectually honest than a satisfaction score that is never tested against anything.

Where Koji fits

Measuring mediators well is harder than averaging a satisfaction item, because mediators are behavioural and specific. This is where instrument design decides everything. Koji for Education is built to capture exactly this kind of evidence. Its AI-moderated conversational interviews probe for concrete experience — describe a time you got feedback you could act on; what did the project ask you to actually produce — rather than soliciting a global guess about employability. Six structured question types (open-ended, scale, single- and multiple-choice, ranking, yes/no) let a programme team operationalise each mediator precisely, and ranking items can force trade-offs (which skills did this course actually develop most?). Automatic thematic analysis turns open-text accounts of authentic tasks and feedback into structured, programme-level evidence; quality scoring filters low-information responses; and formative, mid-cycle collection means the mediators are measured early enough to change a running course rather than only autopsying a finished one. Programme- and institution-level reporting lets you line up in-course mediators against tracer and employer data in one place — the triangulation that turns leading indicators into an accreditation-ready story — all within a GDPR/AVG-compliant, EU-appropriate pipeline. The same conversational engine drives the main Koji platform for customer and user research, where measuring the drivers of an outcome rather than the outcome itself is the everyday discipline.

Course evaluation will never see a graduate's career. But it can see, clearly and early, the conditions that make that career more likely — if you ask the right questions and triangulate the answers. That is the employability evidence base worth building.

Build an evidence base accreditors trust. See Koji for Education.