Evaluating Doctoral Supervision: Why the 89% Satisfaction Headline Hides the Students Who Need You Most
Supervision is consistently the highest-rated part of the doctoral experience — and also the one whose failures are most catastrophic for the individual. The annual survey that produces the reassuring average is exactly the wrong instrument for finding the candidate quietly in trouble.
Koji for Education
Research & Editorial Team · June 13, 2026
Answer first: Doctoral supervision is hard to evaluate because it is a long, one-to-one, power-laden relationship in which true anonymity is nearly impossible and the headline satisfaction number is reassuringly high. Sector surveys like PRES report supervision at around 89% satisfaction, but that average masks a tail of candidates whose supervision is failing — the people evaluation most needs to reach. Evaluating supervision well requires confidential, formative, in-cycle feedback that probes the relationship without exposing the candidate, not a once-a-year summative survey that arrives too late to help anyone currently struggling.
The most important relationship in the university, barely evaluated
A doctoral candidate's relationship with their supervisor shapes years of their life, their mental health, their research, and their career. It is arguably the highest-stakes teaching relationship in higher education. And it is among the least systematically evaluated — because the methods that work for a 200-student lecture course fall apart when applied to a relationship of one.
What sector-level data exists is genuinely encouraging at the aggregate. Advance HE's Postgraduate Research Experience Survey (PRES) — the largest of its kind, drawing on more than 35,000 postgraduate researchers across 93 institutions in its 2025 cycle — found overall PGR satisfaction at 83%, the highest since 2011, with supervision the single best-performing area at around 89% satisfaction. The same data identified a candidate's sense of belonging as the strongest correlate of overall satisfaction, with roughly two-thirds of PGRs reporting that they belong.
Those are good numbers. They are also exactly why the average is dangerous.
Why a high average is the wrong thing to celebrate
When 89% are satisfied with supervision, attention naturally flows to the reassuring nine-in-ten. But evaluation exists to find and fix problems, and the problems live in the other 11% — a group that, across a large sector, represents thousands of doctoral candidates whose supervision is not working. For a PhD student, failing supervision is not a minor dissatisfaction; it can mean a stalled project, a lost career, or a serious mental-health crisis. Averaging is precisely the operation that makes this tail invisible, the structural problem we examine in why averaging Likert scores misleads and negativity and the way distributions hide minorities.
The European policy framework has long understood supervision as too important to leave to a satisfaction score. The EUA's Salzburg II Recommendations and the EU's Principles for Innovative Doctoral Training both frame supervision as a transparent, contractual relationship of shared responsibilities between candidate, supervisor, and institution — something to be actively assured, not assumed. Assuring it requires knowing when the relationship is breaking down. An annual aggregate cannot tell you that.
What makes supervision genuinely hard to evaluate
Four features of doctoral supervision defeat the standard course-evaluation toolkit:
- It is a relationship of one. With a single supervisor (or a small team), responses cannot be anonymised by volume. A candidate who criticises their supervisor knows that supervisor could identify them. The result is heavy self-censorship — the very anonymity-versus-confidentiality tension that determines whether feedback is honest.
- It is long and evolving. A PhD runs three to four years or more. A relationship that is excellent in year one can deteriorate by year three. A single end-point survey captures none of this trajectory; the candidate needs a channel that is open throughout, the formative logic we set out in formative vs summative evaluation.
- It is power-laden. The supervisor controls progression, funding references, examiner relationships, and future career doors. The cost of honest negative feedback is perceived — often correctly — as high. Any evaluation that does not credibly protect the candidate will simply harvest reassurance.
- It is high-variance and personal. Generic Likert items ("My supervisor provides helpful feedback: 1–5") cannot capture what actually goes wrong: mismatched expectations, neglect, over-control, interpersonal conflict, or a candidate who does not know what good supervision should even look like.
But supervision already scores 89% — why intervene?
This is the strongest objection, and it has two parts. First: supervision is the best-performing area in the sector's flagship survey, so why divert effort to it rather than to weaker areas like research culture or community? Second: supervision is so sensitive that formal evaluation risks damaging the very trust it depends on — better to leave it to informal pastoral channels.
Both deserve a serious answer. On the high score: evaluation is not a popularity contest, and a high mean is not evidence that no one is suffering — it is, statistically, perfectly compatible with a substantial tail in acute difficulty. The job of supervision evaluation is not to confirm that most candidates are fine (PRES already does that) but to find the ones who are not, early enough to act. A measure optimised for the aggregate is the wrong tool for that job. As we argue in triangulation, a single high number is necessary but nowhere near sufficient.
On sensitivity: the risk is real, and it is an argument for better method, not for no method. Crude, identifiable, punitive evaluation would indeed damage trust — which is exactly why formative, confidential, candidate-controlled feedback matters. The alternative to careful evaluation is not "trust preserved"; it is candidates in difficulty with no safe channel, discovered only at the point of withdrawal or formal complaint, when it is far too late. Informal pastoral care is valuable but unsystematic, and it disproportionately reaches the confident and well-networked while missing the isolated — the candidates least likely to have a sense of belonging in the first place.
There is also a triangulation point specific to doctoral education. Supervision quality is visible not only to the candidate but in the artefacts of the relationship — annual progress reviews, milestone panels, thesis-committee notes, completion times, and submission rates. Treating candidate feedback as one strand woven together with these institutional signals guards against both false alarms and false reassurance: a glowing self-report alongside repeatedly missed milestones tells a different story than either source alone. The aim is not to police supervisors but to give doctoral schools an early, multi-source picture of where a relationship needs support before it reaches crisis. Crucially, that picture should inform development and mediation, not punishment — the moment evaluation becomes a disciplinary weapon, the honest disclosure it depends on evaporates.
Where Koji fits
Koji for Education is designed for exactly the feedback that supervision needs and surveys cannot provide. Its AI-moderated conversational interviews create a confidential, judgement-free space where a candidate can describe what is actually happening in the relationship — and because the moderator is a standardised, bias-aware AI rather than a colleague or the supervisor's peer, the candidate is not performing for someone embedded in the same power structure. The interview can probe gently and follow up ("you mentioned feedback has slowed — can you say more?") in a way no fixed form and no awkward human interviewer reliably can.
Because Koji runs formative, mid-cycle collection rather than a single end-point survey, it can keep a channel open across the years of a doctorate, catching deterioration when something can still be done. Its automatic thematic analysis lets a doctoral school see patterns across a cohort — recurring issues of expectation-setting, feedback frequency, or isolation — and act at the structural level while protecting any individual candidate from exposure, navigating the anonymity and GDPR constraints that make this domain so delicate. Its programme- and institution-level reporting rolls these signals up for graduate schools and quality teams within GDPR/AVG-compliant, EU-appropriate handling, matching the assurance expectations that frameworks like Salzburg II set out. To be precise about the claim: Koji does not eliminate the power dynamic or guarantee disclosure — it reduces the barriers to honest feedback and surfaces trouble earlier than an annual aggregate ever could.
The same conversational engine powers the main Koji platform, where teams running longitudinal relationship research — onboarding journeys, long-term customer health — face the same need to track a relationship over time rather than snapshot a satisfaction score.
The 89% will always be reassuring. The question every doctoral school should ask is whether it can find the candidate who is part of the other 11% — before they disappear. See how Koji for Education evaluates the supervision relationship.