Realist Evaluation for Course Feedback: 'What Works, for Whom, in What Circumstances'
Pawson and Tilley's realist evaluation replaces 'did it work?' with 'what works, for whom, in what circumstances?' using context-mechanism-outcome configurations. Here is the theory, the evidence, the caveats, and how to apply it to course evaluation.
Koji Education Team
Product
In brief
Realist evaluation, set out by Ray Pawson and Nick Tilley in Realistic Evaluation (1997), replaces the blunt question "did this course work?" with "what works, for whom, in what circumstances, and why?" Its central tool is the context-mechanism-outcome (CMO) configuration: an intervention (a new teaching method, a feedback change, a redesigned assessment) does not produce outcomes directly — it offers resources that trigger mechanisms (changes in students' reasoning and response), and whether those mechanisms fire depends on context. For course evaluation, this reframes an aggregate score into a set of testable hypotheses about which students benefited, under which conditions, and through what causal pathway.
What the research says
Pawson and Tilley developed realist evaluation as a "third way" between naive experimentalism (which asks only whether an intervention has a net effect) and pure constructivism (which resists causal claims altogether). Their philosophical starting point is scientific realism: programmes are theories incarnate. To evaluate one is to test the theory of how it is supposed to work. The signature output is the CMO statement — for example, "problem-based tutorials (intervention) improve deep engagement (outcome) for students who already have strong prior knowledge (context) because the open-ended tasks activate their sense of competence and autonomy (mechanism), but the same tutorials increase anxiety and surface learning for underprepared students in the same room."
The mechanism concept is doing the heavy lifting, and it has been sharpened considerably since 1997. Astbury and Leeuw (2010), in Unpacking Black Boxes: Mechanisms and Theory Building in Evaluation, argue that mechanisms are typically (a) hidden, (b) sensitive to context, and (c) about the reasoning and reactions of participants rather than the programme's moving parts. In education this is intuitive: the same lecture "causes" mastery in one student and disengagement in another because it interacts with different prior knowledge, motivation and identity. A mean rating averages these opposite reactions into a middling number that describes no one.
Realist evaluation has become mainstream in health-professions and higher education research. The RAMESES reporting standards (Wong and colleagues, 2013) formalised realist synthesis and evaluation for the health and education literatures, and recent applications continue to use CMO logic to explain professional learning. The approach is congenial to education for a specific empirical reason: the SET-and-learning literature repeatedly finds weak average relationships — Uttl, White and Gonzalez (2017) put shared variance near 1% — precisely the pattern you would expect if strong but opposite context-dependent effects are cancelling out in the aggregate. Realist evaluation says: stop averaging away the heterogeneity; model it.
Why it matters for course evaluation in practice
Conventional evaluation asks whether the mean moved. Realist evaluation asks a more useful set of questions that map cleanly onto quality-assurance decisions:
-
Segment before you conclude ("for whom"). A course rated 3.8 overall may be a 4.6 for continuing-track students and a 2.9 for those taking it as a required service course. The average conceals the actionable finding. Realist practice insists on disaggregating by the contexts that plausibly matter — prior attainment, discipline, mode, elective vs required.
-
Ask why, not just how much ("mechanism"). If a redesigned seminar lifted engagement, the reusable knowledge is the mechanism — was it psychological safety, cognitive challenge, relevance to assessment? That is what transfers to the next course; the score does not.
-
Build and test programme theory over time. Each cycle refines the CMO hypotheses. This turns evaluation from a recurring verdict into a cumulative, department-owned theory of what teaching works locally — the opposite of chasing a benchmark.
This is a natural partner to utilization-focused evaluation: CMO findings are inherently actionable because they name the condition and the lever, not just the result.
A fourth practical shift follows from the first three: report the configuration, not just the coefficient. A realist evaluation report is most useful when each finding is written as an explicit "for whom, in what context, through what mechanism" sentence that names the sub-group, the condition, and the presumed causal pathway. This forces the analyst to commit to a claim specific enough to be wrong — and therefore specific enough to be tested and acted on — rather than retreating to a defensible but useless average. It is also the format that lets a programme committee argue about the mechanism itself, which is where the pedagogically interesting disagreements live.
Limitations and honest caveats
- Causal claims from observational data. CMO configurations are hypotheses, not proven causal mechanisms. Without design safeguards, a realist story can rationalise almost any pattern after the fact. Realist evaluation is a theory-testing discipline; treating its outputs as confirmed causation overclaims.
- The mechanism concept is slippery. Astbury and Leeuw's whole paper exists because "mechanism" is used inconsistently. Teams frequently mislabel programme activities (the tutorial) as mechanisms when the mechanism is the student's reasoning (feeling competent). Sloppy specification produces sloppy conclusions.
- Sub-group analysis invites false positives. Slicing by many contexts multiplies comparisons; small cohorts yield unstable estimates. The multiple-comparison and small-sample cautions elsewhere in this knowledge base apply with full force — disaggregation must be pre-specified and read with appropriate uncertainty.
- Labour and expertise. Eliciting mechanisms requires probing qualitative data and analytic skill. It is heavier than tallying a Likert item, and demands evaluators comfortable with theory.
- Not a bias antidote. Realist evaluation explains heterogeneity; it does not by itself remove the gender, accent, or leniency biases documented in the SET literature. Those confounds can masquerade as "context" if not handled explicitly.
How Koji incorporates this
Realist evaluation needs two things that traditional survey tools supply poorly: the why behind a rating, and the ability to see for whom a course works. Koji is built around both.
- AI-moderated conversational interviews are, in effect, mechanism-elicitation engines. When a student rates a seminar highly, the interviewer asks what specifically changed for them and why — surfacing the reasoning (felt challenge, relevance, safety) that is the mechanism in a CMO configuration. A static questionnaire cannot do this; a dialogue can.
- Structured questions plus segmentation. Combining scale, single_choice and multiple_choice items (prior background, mode, motivation) with open-ended probing lets you construct the context side of CMO and disaggregate outcomes by sub-group — turning "3.8 overall" into "for whom, and why."
- Automatic thematic analysis clusters the mechanism language across a cohort, so recurring causal pathways become visible and auditable rather than anecdotal.
- Triangulation across cohorts and terms supports the cumulative theory-building that realist evaluation prizes: hypotheses generated one term can be probed the next.
Two honest guardrails, consistent with the method's own cautions. First, Koji is designed to help you generate and probe CMO hypotheses, not to certify causation — the platform surfaces patterns for expert judgement, it does not manufacture proof. Second, disaggregated reporting is offered with sample-size and uncertainty flags so small sub-groups are not over-interpreted. The same conversational engine underpins Koji's core research platform at koji.so, where "what works, for whom" is the everyday language of product and customer research.
A worked example: reading one score three ways
Consider a redesigned second-year quantitative-methods module that returns an unremarkable 3.7 overall. A conventional report files it as "satisfactory, no action." A realist reading refuses to stop there and disaggregates. Among students entering with strong prior mathematics, the module scores 4.5, and their open comments describe the new open-ended problem sets as energising: the mechanism is a reinforced sense of competence and autonomy, firing in the context of adequate preparation. Among students taking the module as a required service course with weaker preparation, it scores 2.8, and their comments describe the same problem sets as bewildering and anxiety-inducing: the identical resource triggers a different mechanism — threat and disengagement — because the context differs.
The CMO configurations that emerge are actionable in a way the 3.7 never was. They point not to "improve the module" in the abstract but to a specific structural fix: scaffold the open-ended tasks for underprepared students, or stream the service cohort differently, while preserving the challenge the prepared students value. Crucially, the average of 3.7 recommended nothing, because it described a course that helped and harmed in equal, cancelling measure. This is the everyday payoff of realist thinking: it converts an inert benchmark into a testable, improvement-oriented theory of who your course serves and how. It also disciplines the analyst — each configuration is a claim that can be checked next term, not a story that merely sounds plausible this one.
Related resources
- Utilization-focused evaluation: designing course feedback for use
- Multilevel models for nested course-evaluation data
- Measurement invariance and differential item functioning
- Small mean differences and confidence intervals
- Feeling of learning vs actual learning (Deslauriers)
- Is student evaluation of teaching valid? Spooren's state of the art
References
- Pawson, R. & Tilley, N. (1997). Realistic Evaluation. London: Sage.
- Pawson, R. & Tilley, N. (2004). Realist Evaluation. Paper for the British Cabinet Office. https://www.betterevaluation.org/methods-approaches/approaches/realist-evaluation
- Astbury, B. & Leeuw, F. L. (2010). Unpacking black boxes: Mechanisms and theory building in evaluation. American Journal of Evaluation, 31(3), 363-381. https://doi.org/10.1177/1098214010371972
- Wong, G., Greenhalgh, T., Westhorp, G., Buckingham, J. & Pawson, R. (2013). RAMESES publication standards: realist syntheses. BMC Medicine, 11, 21. https://doi.org/10.1186/1741-7015-11-21
- Uttl, B., White, C. A. & Gonzalez, D. W. (2017). Meta-analysis of faculty's teaching effectiveness: Student evaluation of teaching ratings and student learning are not related. Studies in Educational Evaluation, 54, 22-42. https://doi.org/10.1016/j.stueduc.2016.08.007
Related articles
Measurement Invariance: Can You Compare Course-Evaluation Scores Across Groups at All?
Comparing average ratings across departments, languages or online vs paper assumes the questionnaire means the same thing to every group. Measurement invariance is the test that assumption usually fails — and almost nobody runs it.
The 4.2 vs 4.4 Trap: Why Small Differences in Course-Evaluation Means Are Usually Noise
Faculty and administrators routinely read meaning into tiny gaps between course-evaluation means. Boysen (2015) and Boysen et al. (2014) show this happens even when confidence intervals say the difference is nothing — and that warnings barely help. How to report uncertainty honestly.
Students Are Nested in Courses: Why Multilevel Models Beat Raw Averages for Course Evaluation
Course-evaluation data has a nested structure — students within sections within instructors — and a flat average ignores it. This guide explains how multilevel (hierarchical linear) models partition variance, why ignoring clustering understates uncertainty, and what it means for fair reporting.
Designing Course Evaluation for Use: The Utilization-Focused Approach
The biggest failure of course evaluation is not bad data — it is data nobody acts on. Patton's Utilization-Focused Evaluation and the empirical research on evaluation use (Johnson et al. 2009) show how to design feedback for action from the start.