New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods11 min read

Evaluating Clinical Placements and Work-Based Learning: Why the End-of-Course Questionnaire Is the Wrong Instrument

Why standard end-of-course student evaluation questionnaires fail to capture placement quality, and what validated clinical/practice-learning-environment instruments (MCPI, CLES+T, UCEEM, PET, CLEI, DREEM) measure instead.

Koji Education Team

Product

In brief: The standard end-of-course student evaluation of teaching (SET) questionnaire is engineered around one lecturer delivering structured content in a classroom — a model that does not exist on a clinical placement, where learning is mediated by a supervisor, a workplace team, patient contact, and organisational culture. Decades of psychometric work show that placement quality is better captured by validated learning-environment instruments — the Manchester Clinical Placement Index (MCPI), the CLES+T scale, the UCEEM, the Placement Evaluation Tool (PET), the CLEI and the DREEM — which measure the environment as a multidimensional psychosocial construct rather than rating "the lecturer." If your professional programmes are evaluating placements with the same form you use for lectures, you are measuring the wrong thing.

Why the standard SET is the wrong instrument for placements

Student evaluation of teaching questionnaires are a mature, defensible technology for what they were built to do. As Marsh's programme of research established, effective SET instruments are multidimensional and reliable, capturing distinct factors such as instructor clarity, workload, and organisation. But every one of those factors presupposes a specific situation: a nameable instructor, a bounded module, structured content, and a classroom.

A clinical or practice placement violates all of these assumptions simultaneously. There may be no single "lecturer" — instead a named preceptor or clinical supervisor, a rotating team, a ward sister or clinic lead, and incidental teaching from many staff. The "content" is not a syllabus but the flow of real work with real patients or clients. The "classroom" is a workplace with its own culture, hierarchy, safety climate and tolerance for a learner who slows things down. Asking a student to rate "the clarity of the lecturer's explanations" or "how well the module was organised" in this setting produces answers that are, at best, noise and, at worst, actively misleading — because the items do not map onto the causal structure of how the student actually learned.

This is a construct validity problem, not a wording problem. You cannot fix a lecture-course SET for placements by editing a few items. The construct it measures — quality of taught instruction — is the wrong construct. The right construct is the learning environment.

What the research says: the learning environment as a measurable construct

Across medicine, nursing and the allied health professions, researchers have converged on treating the practice-learning environment as a multidimensional psychosocial construct: the relational, organisational and experiential conditions in which work-based learning happens. Six validated instruments illustrate both the shared logic and the useful variety.

Manchester Clinical Placement Index (MCPI). Dornan, Muijtjens, Graham, Scherpbier and Boshuizen (2012) developed a deliberately short instrument for medical undergraduates learning in hospital and community "communities of practice." The final MCPI has 8 items loading on two factors — Learning Environment (5 items: leadership, reception, supportiveness of people, organisation, facilities) and Training (3 items: instruction, observation, feedback) — rated on a 0–6 scale. It was validated on 754 completed questionnaires (84% response) across 168 placements (90 hospital, 78 community) among Year 3 Manchester students. Subscale reliabilities were high (Cronbach's α from 0.89 to 0.93), and the two factors explained 78% (hospital) and 82% (community) of variance. Crucially, the authors report generalisability analyses showing that roughly 7–11 student raters are needed to judge a placement reliably — a direct, quantified statement that placement quality is a property of a site, not of one respondent's mood.

CLES+T (Clinical Learning Environment, Supervision and Nurse Teacher). The most internationally adopted instrument in nursing. Saarikoski, Isoaho, Warne and Leino-Kilpi (2008) extended the earlier CLES by adding a nurse-teacher dimension, producing a 34-item scale with five subscales: pedagogical atmosphere on the ward (9 items), leadership style of the ward manager (4), premises of nursing care on the ward (4), the supervisory relationship (8), and the role of the nurse teacher (9). Notice what dominates: the single largest reliability contribution comes from the supervisory relationship, which in later validations reaches Cronbach's α as high as 0.98. The instrument has since been translated into more than 30 languages — evidence of both its usefulness and the community's insistence on re-validating it in each new context.

UCEEM (Undergraduate Clinical Education Environment Measure). Strand, Sjöborg, Stalmeijer, Wichmann-Hansen, Jakobsson and Edgren (2013) built a 25-item measure for undergraduate medical clinical education in Sweden, organised under two overarching dimensions — experiential learning and social participation — and four subscales: opportunities to learn in and through work & quality of supervision, preparedness for student entry, workplace interaction patterns & student inclusion, and equal treatment. Reported internal consistency ranged from roughly 0.79 to 0.91. The theoretical framing is significant: it treats learning as participation in a workplace, drawing on workplace-learning theory rather than classroom pedagogy.

Placement Evaluation Tool (PET). Cooper and colleagues (2020) took a feasibility-first approach for nursing. The PET has 19 Likert items plus one global satisfaction rating, resolving into two factors — Clinical Environment (8 items, α = 0.94) and Learning Support (11 items, α = 0.96) — with strong concurrent validity against CLES+T (r = 0.834). It was pilot-tested with 1,263 pre-registration nursing students across seven Australian institutions, and its median completion time was about 3.5 minutes. The PET is the pragmatic answer to a real QA constraint: an instrument nobody completes measures nothing.

CLEI (Clinical Learning Environment Inventory). Chan (2003) adapted classroom-climate methodology to the clinical setting, producing a 42-item, six-scale inventory (personalisation, student involvement, task orientation, innovation, satisfaction, individualisation) with a distinctive Actual versus Preferred form design. Validated with 108 pre-registration nursing students, it surfaced systematic gaps between the environment students experienced and the one they wanted — a diagnostic logic that echoes service-quality gap models.

DREEM (Dundee Ready Education Environment Measure). The broadest of the six, Roff, McAleer, Harden and colleagues (1997) built a 50-item measure of the overall educational environment with five domains (perceptions of learning, teachers, atmosphere, and academic and social self-perception). DREEM is not placement-specific, but it anchors the whole tradition: the recognition that "environment" is measurable, multidimensional, and diagnostic.

Read together, three findings are robust. First, the recurring dimensions — supervisory/preceptor relationship, pedagogical atmosphere/climate, unit leadership, opportunity to participate in real work, inclusion and equal treatment — appear again and again across instruments and disciplines. Second, none of these dimensions is captured by a lecture-course SET. Third, reliability is a property of the placement aggregated across several students, not of any single questionnaire.

Why it matters in practice

For an institution's quality assurance and accreditation function, this is not academic hair-splitting. Professional programmes — nursing, medicine, midwifery, allied health, teaching, social work — are accredited against standards that explicitly cover the practice learning environment. When a regulator or accrediting body asks how you assure placement quality, "we send students the standard module-evaluation survey" is a weak answer: the instrument is not aligned to the standard being assessed. As the medical-education accreditation literature (e.g. WFME-aligned frameworks) makes clear, evidence of environment quality needs to speak to supervision, workplace conditions and student participation — exactly the dimensions the placement instruments were built to measure.

There is also a management-signal difference. A lecture SET tells a programme director something about one academic's teaching. A placement instrument tells them something about a site: which wards, clinics or partner organisations offer a healthy learning environment and which are quietly harming students' development or wellbeing. That is site-level, longitudinal, comparative intelligence — the kind that supports decisions about where to place students, which supervisors need support, and which partnerships to expand or retire.

Limitations and honest caveats

None of this makes placement instruments a solved problem, and overselling them would be its own failure of rigour.

  • Self-report and response bias. Every instrument here is student self-report. It measures perceptions of the environment, which correlate with but are not identical to the environment itself. Halo effects, recency, and mood all intrude.
  • Power dynamics. This is the sharpest issue and it is worse than in lecture courses. A student evaluating a placement is frequently evaluating a named supervisor who may assess them, write references, or belong to a small professional community the student is trying to enter. The incentive to give inflated, socially desirable responses is structural, not incidental — and it directly threatens the validity of exactly the supervisory-relationship dimension that matters most.
  • Anonymity is hard in small cohorts. A ward that took two students this term cannot be given "anonymous" feedback without those two students being identifiable. Small-n placements require deliberate anonymity thresholds and aggregation, or students will (rightly) self-censor.
  • Context- and culture-specificity. These instruments are not freely portable. CLES+T's 30-plus translations exist because each new context re-tests factor structure; several transfer imperfectly. An instrument developed for Australian nursing or Swedish medicine cannot be assumed valid, unmodified, for a German or Dutch professional programme without local validation or at least cognitive pretesting.
  • Timing. End-of-placement data arrives too late to help the student who generated it. Purely summative collection systematically misses the students most at risk.

Honest practice means using these instruments as strong but fallible signals, triangulated with other evidence, not as verdicts.

How Koji incorporates this

Koji for Education is built for exactly the relational, environmental, small-cohort character of placement evaluation — the situation a fixed lecture-SET form handles worst. Its mechanisms are designed to mitigate the problems above, not to eliminate them.

  • AI-moderated conversational interviews. Because placement quality lives in relationships and workplace culture, a rigid checkbox form leaves the most important signal on the table. Koji's AI-moderated interviews probe adaptively — following up on a vague "my supervisor was fine" to understand what actually happened — which suits the narrative, situated nature of practice learning far better than a static questionnaire.
  • Structured question types to rebuild validated dimensions. Koji supports open_ended, scale, single_choice, multiple_choice, ranking and yes_no questions, so an institution can construct a placement-specific instrument that mirrors the validated dimensions above — supervisory relationship, pedagogical atmosphere, opportunity to participate, inclusion — rather than forcing placements through a lecture form. Scale items give you the comparable, aggregable numbers accreditation needs; open-ended items give you the "why."
  • Automatic thematic analysis of open text. Free-text placement comments are where safety and relational problems surface first. Koji analyses open responses thematically at scale, so recurring signals across a cohort are visible without a human hand-coding every transcript.
  • Anonymity handling for small placement cohorts. Koji is designed to support aggregation and anonymity thresholds so that feedback on a two-student ward is not trivially de-anonymised — directly addressing the power-dynamic and small-n risks that most threaten placement-evaluation validity.
  • Mid-placement / formative collection. Because Koji studies can run during a placement, a deteriorating supervisory relationship or an unsafe climate can surface while the student is still on site and something can still be done — not in a summative report filed after the harm is complete.
  • Triangulation and closing the loop. Results can be compared across sites and cohorts to build the site-level, longitudinal intelligence placements actually require, and action-tracking supports closing the loop so that a flagged site problem leads to a recorded response.

None of this removes self-report bias or guarantees candour under power imbalance; it is designed to reduce those distortions and to fit the instrument to the construct. The same AI-moderated interview engine underpins Koji's core research platform at koji.so, where it is applied to product and customer research — the education product is that engine pointed at the specific, high-stakes problem of evaluating how and where students learn through work.

References

  • Chan, D. S. (2003). Validation of the Clinical Learning Environment Inventory. Western Journal of Nursing Research, 25(5), 519–532. https://doi.org/10.1177/0193945903253161
  • Cooper, S., Cant, R., Waters, D., Luders, E., Henderson, A., Willetts, G., Tower, M., Reid-Searl, K., Ryan, C., & Hood, K. (2020). Measuring the quality of nursing clinical placements and the development of the Placement Evaluation Tool (PET) in a mixed methods co-design project. BMC Nursing, 19, 101. https://doi.org/10.1186/s12912-020-00491-1
  • Dornan, T., Muijtjens, A., Graham, J., Scherpbier, A., & Boshuizen, H. (2012). Manchester Clinical Placement Index (MCPI). Conditions for medical students' learning in hospital and community placements. Advances in Health Sciences Education, 17(5), 703–716. https://doi.org/10.1007/s10459-011-9344-x
  • Roff, S., McAleer, S., Harden, R. M., Al-Qahtani, M., Ahmed, A. U., Deza, H., Groenen, G., & Primparyon, P. (1997). Development and validation of the Dundee Ready Education Environment Measure (DREEM). Medical Teacher, 19(4), 295–299. https://doi.org/10.3109/01421599709034208
  • Saarikoski, M., Isoaho, H., Warne, T., & Leino-Kilpi, H. (2008). The nurse teacher in clinical practice: Developing the new sub-dimension to the Clinical Learning Environment and Supervision (CLES) scale. International Journal of Nursing Studies, 45(8), 1233–1237. https://doi.org/10.1016/j.ijnurstu.2007.07.009
  • Strand, P., Sjöborg, K., Stalmeijer, R., Wichmann-Hansen, G., Jakobsson, U., & Edgren, G. (2013). Development and psychometric evaluation of the Undergraduate Clinical Education Environment Measure (UCEEM). Medical Teacher, 35(12), 1014–1026. https://doi.org/10.3109/0142159X.2013.835389

Related Resources

Related articles

accreditation

Course Evaluation Evidence for WFME Medical-School Accreditation (Programme Evaluation, Area 7)

How to turn student course feedback into WFME-ready evidence: a buyer's guide mapping Area 7 Programme Evaluation (mechanisms, teacher and student feedback, cohort performance, stakeholder involvement) to concrete evaluation outputs.

research-methods

What Can Open-Text Student Comments Tell You That Likert Scores Cannot?

A research-grounded guide to open-ended student comments in course evaluation: what Alhija and Fresko found about who writes them and what they contain, how thematic analysis surfaces issues numbers miss (Stupans et al.), how rare abusive comments actually are (Tucker), and how to turn free text into reliable evidence.

research-methods

What Do Student Evaluations Actually Measure? Marsh, the SEEQ, and the Case for Multidimensional Feedback

Herbert Marsh spent decades showing that well-built student evaluations are multidimensional, reliable, and stable — and that a single global score throws away most of what they can tell you. What his work establishes, where critics push back, and how to design feedback that is actually usable.

research-methods

Are Students Customers? The SERVQUAL Gap Model, HEdPERF, and What Service-Quality Thinking Adds to Course Evaluation

The SERVQUAL gap model and its higher-education variant HEdPERF measure the distance between what students expect and what they perceive they received. We assess the evidence, the sharp limitations of treating students as customers, and how Koji uses expectation framing without collapsing learning into satisfaction.