New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Graduate outcomes9 min read

Evaluating Placements, Internships, and Work-Integrated Learning: Why the End-of-Module Survey Doesn't Fit

Work-integrated learning is one of the strongest levers for graduate employability in European higher education — and one of the worst-served by standard course evaluation. A placement has two supervisors, no lecture hall, and outcomes that arrive months later. Here is why the end-of-module survey breaks, and what fits instead.

Koji Education Team

Product ·

The short answer: Work-integrated learning (WIL) — placements, internships, sandwich years, clinical and practicum experiences — is among the most powerful tools European universities have for graduate employability, yet it is routinely evaluated with an instrument designed for a lecture hall. A placement has distributed responsibility (an academic tutor and a workplace supervisor), unconventional timing, and outcomes that only become visible months or years later. The standard end-of-module satisfaction survey cannot see any of that. Evaluating WIL well means rethinking who is evaluated, when, and against what — and that is a job for flexible, conversational, formative evaluation rather than a fixed five-point form.

Why WIL deserves serious evaluation

The case for WIL rests on evidence, not enthusiasm. A 2025 systematic review of work-integrated learning across the European Higher Education Area found WIL consistently associated with improved work readiness, professional networks, and smoother school-to-work transitions — while also flagging that inconsistent design, assessment and evaluation are exactly where its promise leaks away (Curto-Reverte et al., 2025, Review of Education).

Crucially, the effect survives the most obvious objection — that high-calibre students simply self-select into placements. Using propensity-score matching to control for that selection, Jones, Green and Higson (2017) found that a work placement still improved final-year academic performance by roughly 2% at Aston and 4% at Ulster — a real causal increment, not just a reflection of who chooses to go (Jones, Green & Higson, 2017, Studies in Higher Education). If WIL genuinely moves both employability and attainment, evaluating it badly is not a minor administrative gap; it is failing to learn from one of the highest-leverage things the institution does.

Why the end-of-module survey breaks

The standard course-evaluation instrument carries four assumptions, and WIL violates every one.

1. One teacher to rate. A lecture has an instructor. A placement has at least two responsible parties — an academic placement tutor and a workplace/host supervisor — whose contributions are different in kind. Asking a student to "rate the teaching" collapses a mentor, an employer, and a university coordinator into a single nonsensical average. This is the attribution problem we examine for team-taught courses, in an even sharper form, because one of the "teachers" is not an academic at all.

2. A shared classroom experience. Standard items ("lectures were well organised", "materials were clear") presume a common, designed learning environment. On placement, every student is in a different workplace with different tasks, supervision quality, and culture. There is no common stimulus to rate, so cohort averages on classroom-style items are close to meaningless.

3. End-of-term timing. WIL learning is experiential and unfolds over months; the most important moments — a difficult first week, a supervisor who went quiet, a project that fell through — need catching while they happen, not in a retrospective survey after the placement ends. A summative form at the close cannot support the student who is struggling in week three. This is the formative-versus-summative gap, writ large; see our piece on formative vs summative evaluation.

4. Satisfaction as the outcome. A standard survey measures whether students enjoyed the experience. But the point of WIL is skill development, professional identity, and employability — outcomes that a happiness score barely touches and that often only become legible after graduation, through graduate tracer and destination data. Measuring WIL by satisfaction is measuring the wrong construct, a problem we explore in beyond satisfaction: measuring skills, not just happiness.

What good WIL evaluation actually requires

If the standard survey fails, what replaces it? The literature and practice point to four design shifts.

  • Multi-source, role-aware feedback. Capture the student voice, but also the workplace supervisor's perspective and the academic tutor's — and keep them distinct rather than averaged. This is triangulation applied to placements: no single source sees the whole experience.
  • Formative, in-placement check-ins. Lightweight mid-placement touchpoints that can flag a struggling student or a poor host in time to act, rather than a single post-mortem.
  • Outcome- and skills-oriented questions. Ask what the student can now do that they could not before, what professional behaviours they developed, and how the experience connected to their programme — not just whether they were satisfied.
  • A loop back to programme design and employers. WIL feedback should feed the employer feedback loop and programme-level review, because the lessons are about curriculum alignment and host quality, not just individual modules.

"But isn't placement evaluation too bespoke to standardise?"

This is the strongest objection, and it has real force. Placements are heterogeneous almost by definition — a nursing practicum, a software internship, and a humanities work placement share little operational DNA. Critics reasonably argue that any standardised instrument either becomes so generic it is useless, or so specific it cannot compare across a programme. And WIL cohorts are often small and staggered across the year, which compounds the small-sample reliability problems we describe elsewhere: you rarely get a clean, simultaneous cohort to survey.

The objection is correct that a rigid standardised form cannot handle this. But it conflates "standardised process" with "standardised questionnaire." You can hold the process consistent — every student gets a structured, comparable evaluation conversation — while letting the content adapt to a software placement versus a clinical one. That is precisely the capability a fixed survey lacks and an adaptive conversational interview provides. The heterogeneity is an argument against the static form, not against evaluating WIL rigorously.

How Koji fits work-integrated learning

Koji for Education is built for exactly this kind of non-standard, experience-based evaluation. Its AI-moderated conversational interviews adapt to each student's actual placement — probing the real supervisor relationship, the specific project, the skills developed — instead of forcing a lecture-hall questionnaire onto a workplace. Because the moderation is standardized and bias-aware, you get comparable, structured insight across wildly different placements without flattening them into the same ten Likert items. Formative, mid-cycle collection means a check-in can surface a struggling student or an unsuitable host while there is still time to intervene. Automatic thematic analysis turns hundreds of individual placement stories into programme-level themes — "students consistently wanted earlier employer expectations" — and closing-the-loop action tracking plus programme- and institution-level reporting route those themes to the people who design placements and manage employer relationships. Where you do want comparable metrics, the six structured question types (open-ended, scale, single-choice, multiple-choice, ranking, yes/no) let you keep a consistent skills-and-outcomes backbone.

The same conversational engine powers general stakeholder and customer research on the main Koji platform — useful when WIL evaluation extends to the employers and host organisations themselves.

To be precise about the claim: Koji does not turn the irreducible variety of placements into tidy comparability, and it does not replace formal assessment of placement learning. It reduces the mismatch between a classroom instrument and a workplace experience, and surfaces the formative, skills-oriented, multi-source signal that an end-of-module survey structurally cannot.

The host organisation is a data source too

One perspective that classroom evaluation never has to consider is the employer's. In WIL, the host organisation is both a co-educator and a stakeholder with a uniquely informed view: supervisors see whether the student arrived with the skills the programme claims to develop, where the gaps were, and how the placement could be designed better. Capturing that voice does two jobs at once. It evaluates the placement experience from the side the student cannot see, and it feeds the employer feedback loop that keeps a programme aligned with real labour-market needs — the same loop that underpins accreditation-relevant evidence of graduate employability. Ignoring the host means evaluating a two-sided experience from one side only. A conversational approach can interview supervisors as readily as students, using the same engine, so the placement is assessed by everyone who actually shaped it.

The bottom line

Work-integrated learning is too valuable — to students' employability and, on the evidence, even to their attainment — to evaluate with a form built for a different purpose. Stop asking placement students to "rate the teaching." Start asking, conversationally and while it still matters, what they are learning, who is supporting them, and what their host and tutor see.

See how adaptive, formative evaluation handles non-standard learning at Koji for Education, and pair this with our work on graduate employability and the skills gap and closing the employer feedback loop.