Evaluating Teacher Education: The Placement, Not the Lecture
An initial-teacher-education programme is judged on one thing a satisfaction survey never asks about: are its graduates ready to run a classroom, and do they stay? That evidence lives in the placement, the mentor relationship and early-career retention. What TALIS 2018 reveals about where preparation actually breaks.
Koji Education Team
Product ยท August 2, 2026
Initial teacher education (ITE) is the vertical where the gap between "students were satisfied with the course" and "the programme worked" is widest. An ITE programme is ultimately judged on one thing a standard course-evaluation survey never asks about: are its graduates ready to run a classroom on day one, and do they stay in the profession? The evidence that answers that question does not live in a lecture-hall satisfaction score. It lives in the school placement, in the mentor relationship, and in what happens to graduates in their first years of teaching.
The thesis in one line: evaluate teacher education around its clinical core - the practicum and the mentor triad - and treat classroom-readiness and retention as the outcomes, or you will collect ratings that look fine while missing where the programme actually succeeds or fails.
Why teacher education is a distinct evaluation problem
Most degree programmes have a straightforward structure: courses, taught by academics, on a campus. ITE does not. Its centre of gravity is the school placement, where a trainee works in a real classroom, supervised by a mentor triad - the trainee, a school-based mentor, and a university tutor. The scholarship on this is clear: Linda Darling-Hammond has argued for two decades that effective teacher education must be organised around clinical practice, tightly integrating coursework with supervised classroom experience rather than treating the practicum as an add-on (Darling-Hammond, 2006, Journal of Teacher Education). If the placement is the programme, then an evaluation that samples only the taught modules is evaluating the wrong object.
And the outcome is not "satisfaction." It is preparedness that holds up under the reality of a classroom, followed by whether the new teacher survives the early years at all.
What TALIS 2018 shows - and why a mean hides it
The OECD's Teaching and Learning International Survey gives us hard, comparable data on where preparation breaks, and each finding is an argument against the single-satisfaction-number approach.
Preparation is uneven by domain. Across TALIS countries, a large majority of teachers report having been prepared in subject content and general pedagogy - but the use of ICT for teaching was included in the training of only 56%, and teaching in a multicultural setting for just 35% (TALIS 2018 Results, Volume I, OECD). A programme-level satisfaction mean of 4.1 tells you nothing about which domains left trainees exposed. Only disaggregated, probing evaluation does.
Mentoring - the heart of the placement - is scarce and variable. On average across OECD countries in TALIS, only 22% of novice teachers have an assigned mentor (TALIS 2018 Results, Volume II, OECD). Mentor quality is precisely the thing that varies most from placement to placement, and precisely the thing a central end-of-year survey averages into oblivion. "My mentor was excellent" and "my mentor was absent for three weeks" cannot be summarised by a number.
The programme's effect is realised - or lost - after it ends. Only 38% of teachers participate in any formal or informal induction in their first school, and 66% of beginning teachers with up to five years' experience report having had none. Novice teachers feel least confident in classroom management and in deploying a range of instructional practices. The payoff of good ITE, in other words, plays out in an early-career window that almost no programme evaluation ever looks at.
Evaluate the placement, not the timetable
The practical implication is that ITE evaluation has to do three things a conventional survey does not. It has to reach into the placement, per-site, because that is where quality varies. It has to run mid-placement, not just at the end, so a failing placement can be rescued while the trainee is still in it. And it has to probe preparedness by domain in enough depth to learn why a cohort feels unready to teach in multicultural classrooms - not merely that they do.
The strongest objections
"Self-reported preparedness is unreliable - new teachers underestimate themselves." True, and TALIS bears it out: novices are systematically less confident. That is exactly why self-reported readiness must be triangulated - with mentor assessment against professional standards and, over time, with retention and destination data - rather than used alone (our piece on triangulating evidence makes the general case). Self-report is one strand, not the verdict.
"The real test is pupil outcomes, which ITE evaluation cannot measure." Also fair. Pupil learning is a distal outcome with enormous confounds - school context, intake, resourcing - and attributing it back to a teacher's training programme is a value-added minefield of the kind we have flagged around earnings and value-added proxies. The honest response is not to fake a causal claim. It is to evaluate the proximal, controllable things well (placement quality, mentor feedback, domain-specific preparedness) and to track distal lagging indicators like early-career retention without over-claiming causation.
"We already pass our national ITE inspection." Meeting a periodic accreditation or inspection requirement is not the same as running an evaluation that improves the programme. Accreditation is summative and infrequent; your internal evaluation should be formative and continuous. The two answer different questions.
Where Koji fits
Koji for Education is built for exactly this kind of multi-site, practice-based programme. Its AI-moderated conversational interviews probe the placement and mentor experience with a depth a Likert cannot reach - when a trainee mentions a mentoring problem, the interviewer follows up on what happened and when. Because interviews can run mid-placement as a formative check, a weak partner school surfaces while there is still time to intervene. Open-ended, domain-specific probing tells you why preparedness is thin in a given area, not just that it is. Automatic thematic analysis reads across dozens of placements to identify the partner schools and mentoring patterns that need attention, and programme-level reporting lets you map evidence to your professional standards rather than to generic course items. The same conversational engine that runs a mid-placement check can run a longitudinal early-career follow-up - the retention window TALIS shows is decisive - and it powers alumni and workforce research on the main Koji platform too.
Framed honestly: Koji surfaces and structures the placement and readiness evidence that a satisfaction survey flattens. It mitigates the depth, timing and attribution gaps in ITE evaluation. It does not measure a graduate's future pupils' results, because no evaluation credibly can.
If you run initial teacher education and want evaluation that reaches into the placement and the early-career window - not just the lecture hall - talk to the Koji for Education team.
Frequently asked questions
Why not just use a normal course-evaluation survey for teacher education? Because the centre of an ITE programme is the school placement and the mentor relationship, not the taught modules, and the outcome is classroom-readiness and retention rather than satisfaction. A conventional survey samples the wrong object and averages away exactly the placement-to-placement variation that matters.
What does TALIS 2018 tell us about teacher preparation? That preparation is uneven by domain (ICT included for only 56% of teachers, multicultural teaching for 35%), that mentoring is scarce (only 22% of novices have an assigned mentor), and that induction is rare (38% participate; 66% of early-career teachers report none) - so the programme's real effect plays out in a window most evaluations ignore.
Can course evaluation measure whether a teacher-education programme "works"? It can measure proximal, controllable things well - placement quality, mentor feedback, domain-specific preparedness - and track lagging indicators like early-career retention. It cannot credibly attribute future pupil outcomes to the programme, and honest evaluation avoids that claim.
What is the mentor triad and why evaluate it? The mentor triad is the trainee, the school-based mentor and the university tutor who supervise a placement. Because mentor quality varies enormously between schools and is central to the trainee's development, evaluating the placement and the mentoring relationship - per site, mid-placement - is far more informative than an end-of-programme average.
How does Koji help evaluate ITE? Through conversational interviews that probe placement and mentor experiences in depth, mid-placement formative checks that catch failing placements in time, domain-specific open-ended questions, thematic analysis across many placements, programme-level reporting mapped to professional standards, and longitudinal early-career follow-up using the same engine.