New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods10 min read

Evaluating Doctoral Supervision: What PRES and Supervision-Experience Measures Actually Capture

A research-grounded look at how the postgraduate research (PGR) supervision experience is measured, what the UK PRES and validated instruments like the QSDI capture, and where those measures fall short.

Koji Education Team

Product

In brief. Doctoral/postgraduate-research (PGR) supervision is evaluated through two complementary channels: large experience surveys such as Advance HE's Postgraduate Research Experience Survey (PRES), and validated relational instruments such as the Questionnaire on Supervisor-Doctoral Student Interaction (QSDI). These measure the perceived quality of a sustained one-to-one relationship — its functional, relational and developmental dimensions — rather than the delivery of a taught module. The best evidence shows relational quality, autonomy and belongingness genuinely predict satisfaction, progress and even mental health. But supervision-experience measures carry serious validity limits that taught-course evaluation does not: cohorts are tiny (so anonymity is fragile), the student is rating someone with direct power over their career, and cross-sectional self-report confounds supervision with project and personal factors.

Why doctoral supervision is a different measurement problem

Student evaluation of teaching (SET) has a mature, if contested, evidence base built on multidimensional models of what students can validly judge about a module delivered to a cohort. Supervision breaks almost every assumption behind that machinery. There is no lecture, no shared syllabus, and no large class within which one voice is anonymised. Instead there is a bespoke research project pursued over three to four years, mediated by a relationship with one or two supervisors. The "product" being evaluated is not a term's teaching but a long, asymmetric, developmental partnership.

This matters because the constructs worth measuring are different. Anne Lee's widely cited conceptual study — Lee, A. (2008), How are doctoral students supervised? Concepts of doctoral research supervision, Studies in Higher Education, 33(3), 267–281 — interviewed supervisors across disciplines and identified five concepts of supervision: functional (project management), enculturation (helping the student become a member of the disciplinary community), critical thinking (encouraging the student to question and analyse), emancipation (encouraging the student to question and develop themselves), and developing a quality relationship (in which the student is enthused, inspired and cared for). Lee's contribution is qualitative and conceptual rather than psychometric, but it frames the whole measurement task: an evaluation that only checks whether meetings happened and feedback arrived on time captures the functional layer and misses the enculturation, critical-thinking, emancipation and relational layers where doctoral education actually does its work.

What the research says

The relational instrument: the QSDI

The most influential validated instrument for supervision experience is the Questionnaire on Supervisor-Doctoral Student Interaction (QSDI). In Mainhard, T., van der Rijst, R., van Tartwijk, J., & Wubbels, T. (2009), A model for the supervisor–doctoral student relationship, Higher Education, 58(3), 359–373, the authors adapted the interpersonal-circumplex tradition from classroom research to the supervision dyad. The QSDI locates supervisory behaviour on two independent dimensions — influence (the dominance–submission axis) and proximity (the cooperation–opposition axis) — which combine into eight behavioural styles: leadership, helpful/friendly, understanding, giving the student freedom/responsibility, uncertain, dissatisfied, admonishing, and strict. Doctoral students rate 41 behavioural statements on a frequency scale. The original paper reported the QSDI as a reliable and valid instrument, and it has since been psychometrically re-examined in other national contexts.

The conceptual move here is important for anyone designing supervision feedback: the QSDI does not ask "how satisfied are you?" It asks students to describe behaviour on theoretically grounded dimensions, and it is explicitly framed as a tool to give an individual supervisor formative feedback on their interpersonal style toward a particular student. That is a very different design philosophy from a global satisfaction score.

The sector survey: PRES

At the system level, the dominant instrument in the UK and increasingly beyond is Advance HE's Postgraduate Research Experience Survey (PRES). It is the largest survey of its kind: in 2023, more than 37,000 postgraduate research students across 105 higher education institutions (including four in Australia) took part, and overall satisfaction sat at 79% (down slightly from 80% the previous year). PRES covers supervision alongside resources, research culture, progression and assessment, skills and professional development, and wellbeing, and offers institutions confidential benchmarking against the sector, breakable by discipline, mode of study, gender and domicile. Its supervision items ask, in essence, whether students feel their supervisors have the skills and subject knowledge to support them, give helpful and timely feedback, and are available when needed. PRES is a powerful macro-instrument — but by design it is aggregate, periodic and largely end-state; it is a sector thermometer, not a live relational diagnostic for an individual dyad.

What actually drives doctoral outcomes

The most useful corroborating evidence links these measured constructs to real outcomes. van Rooij, E., Fokkens-Bruinsma, M., & Jansen, E. (2021), Factors that influence PhD candidates' success: the importance of PhD project characteristics, Studies in Continuing Education, 43(1), 48–67, surveyed 839 PhD candidates at a Dutch university and modelled satisfaction, progress and quit intentions. Their headline finding is directly relevant to how we should weight supervision measures: a candidate's sense of autonomy (freedom in the project) and belongingness (the quality of the relationship and integration into the community) were the strongest predictors of satisfaction and the strongest buffers against intention to quit, while high perceived workload predicted lower satisfaction and higher quit intentions. More academic and personal support from the daily supervisor was associated with more satisfaction. In other words, the relational and developmental dimensions Lee described and the QSDI measures are not soft extras — they are the variables most tied to whether students thrive or leave.

A contrasting, sobering line of evidence raises the stakes further. Mavrogalou-Foti, Kambouri, & Çili (2024), The supervisory relationship as a predictor of mental health outcomes in doctoral students in the United Kingdom, Frontiers in Psychology, studied 141 UK doctoral students using the QSDI, an actual-versus-preferred relationship discrepancy scale, and the DASS-21. They found that the uncertain supervisory style — indecisive, ambiguous behaviour — was the strongest supervisory predictor of depression, anxiety and stress, and that the gap between a student's actual and preferred supervisory relationship significantly predicted worse mental-health outcomes. This reframes supervision evaluation as partly a duty-of-care instrument, not only a quality-assurance one, and it argues for measures sensitive enough to detect ambiguity and relational mismatch, not just gross dissatisfaction.

Why it matters for supervision evaluation in practice

Three practical implications follow. First, a single satisfaction number is nearly useless for supervision. The evidence says the actionable variance lives in specific relational and developmental behaviours — autonomy granted, availability, decisiveness, community integration — so instruments must decompose the experience the way the QSDI and PRES scales do, not collapse it. Second, timing changes the value of the data. Because the supervisory relationship is longitudinal and repairable, feedback gathered mid-candidature (formative) can change a live relationship, whereas an exit survey can only inform the next cohort. Third, the developmental and pastoral stakes are high enough that the instrument's job is partly early warning — surfacing the "uncertain" or mismatched relationships that predict distress before they become attrition.

Limitations and honest caveats

A critical doctoral reader will and should raise several objections, and none of them should be waved away.

Anonymity is structurally fragile. A supervisor may have only two or three students. Any free-text comment, and often any sufficiently specific rating, can be re-identified. This is not a hypothetical GDPR nicety; it is the single biggest barrier to honest supervision feedback, and it means small-n aggregation and careful data governance are prerequisites, not afterthoughts.

Power asymmetry contaminates the measurement. The person being evaluated typically controls references, authorship order, funding continuation, and the timing of progression and the viva. Rating them critically carries perceived career risk. This produces a systematic leniency/acquiescence bias that pushes scores upward and mutes exactly the signal you most want to catch. No instrument fully removes this; design can only mitigate it through anonymity, aggregation, and by separating formative feedback from anything that touches the supervisor's formal record.

Cross-sectional self-report confounds causes. The strongest studies here — van Rooij et al. (2021), Mavrogalou-Foti et al. (2024) — are correlational and mostly cross-sectional. A student who is struggling for personal or project reasons may rate their supervisor more harshly, and vice versa; supervision quality, project characteristics, workload and personal circumstances are entangled. Effect directions are plausible but not clean causal estimates, and single-site samples (one Dutch university; 141 UK students) limit generalisability across disciplines and systems.

Construct and response biases persist. Response rates to PGR surveys are often modest, raising non-response bias (dissatisfied or disengaged students may opt out — or over-participate). Halo effects, recency, and the tendency to rate a warm relationship as a competent one all threaten construct validity. And replication remains thinner than in the SET literature: instruments like the QSDI are validated but far less independently replicated across contexts than mainstream teaching-evaluation tools.

The honest summary: supervision-experience measures are informative and, when relationally specific, genuinely predictive — but they are best read as triangulated signals alongside progression data, completion rates and qualitative accounts, never as a single verdict on a supervisor.

How Koji incorporates this

Koji for Education is built around AI-moderated conversational interviews rather than one-shot Likert forms, and several of its mechanisms are designed to mitigate — not eliminate — the specific problems above.

  • Probing beyond the number. The AI-moderated interview does not stop at a scale response. When a PGR student rates, say, availability or feedback quality, the interviewer can follow up conversationally to surface the underlying behaviour — mirroring the QSDI's logic of describing behaviour on influence and proximity rather than collecting a global satisfaction figure. This is aimed squarely at the finding that actionable variance lives in specific relational behaviours, not in an overall score.
  • Structured question types where they add rigour. Studies can combine open_ended, scale, single_choice, multiple_choice, ranking and yes_no items — for example, ranking which forms of support matter most (echoing van Rooij et al.'s autonomy/belongingness/workload variables) while leaving room for open narrative on the relationship.
  • Automatic thematic analysis of open text. Free-text responses across a cohort are analysed thematically, so patterns (e.g. recurring "uncertain"/ambiguous-direction signals) can be detected at the aggregate level rather than requiring a human to read identifiable comments one by one — which is also relevant to anonymity.
  • Anonymity and aggregation for tiny cohorts. Because a supervisor may have only two or three students, reporting is designed to work at aggregate/cohort level and to avoid surfacing re-identifiable individual comments, directly targeting the structural anonymity problem that suppresses honest supervision feedback. (This is a mitigation, not a guarantee; genuinely small n always carries residual re-identification risk that governance must manage.)
  • Mid-cycle, formative collection. Feedback can be gathered mid-candidature rather than only at exit, so that a live, repairable relationship can actually be improved — the practical implication of supervision being longitudinal.
  • Quality scoring and bias-aware reporting. Response quality scoring and bias-aware reporting are intended to flag thin or acquiescent responses and to frame results with the leniency/power caveats a critical reader would insist on, rather than presenting inflated scores at face value.
  • Triangulation and closing the loop. Results are meant to be read across cohorts and alongside other evidence, and action-tracking supports closing the loop — recording what changed in response — so feedback informs practice rather than disappearing into an archive.

None of this overturns the underlying caveats: power asymmetry, confounding and small-n anonymity are real constraints on any supervision measure. Koji's aim is to make the instrument relationally specific, formative, and bias-aware enough to be worth acting on.

For teams outside higher education, the same engine powers Koji's core research platform at koji.so, where AI-moderated interviews are applied to product and customer research — the underlying method (conversational probing plus structured questions plus automatic thematic analysis) is identical; only the study design differs.

References

  1. Lee, A. (2008). How are doctoral students supervised? Concepts of doctoral research supervision. Studies in Higher Education, 33(3), 267–281. https://doi.org/10.1080/03075070802049202
  2. Mainhard, T., van der Rijst, R., van Tartwijk, J., & Wubbels, T. (2009). A model for the supervisor–doctoral student relationship. Higher Education, 58(3), 359–373. https://doi.org/10.1007/s10734-009-9199-8
  3. van Rooij, E., Fokkens-Bruinsma, M., & Jansen, E. (2021). Factors that influence PhD candidates' success: the importance of PhD project characteristics. Studies in Continuing Education, 43(1), 48–67. https://doi.org/10.1080/0158037X.2019.1652158
  4. Mavrogalou-Foti, A., Kambouri, M., & Çili, S. (2024). The supervisory relationship as a predictor of mental health outcomes in doctoral students in the United Kingdom. Frontiers in Psychology, 15, 1437819. https://doi.org/10.3389/fpsyg.2024.1437819
  5. Advance HE. (2023). Postgraduate Research Experience Survey (PRES) 2023. Advance HE. https://www.advance-he.ac.uk/knowledge-hub/postgraduate-research-experience-survey-2023

Related Resources