New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Methodology10 min read

Stop Asking "Did the Course Work?" Ask "What Worked, for Whom, in What Circumstances?"

A department average answers the wrong question. Realist evaluation reframes course evaluation around mechanisms and context — why a teaching change fired for some students and misfired for others — and turns a satisfaction score into a testable explanation you can act on.

Koji Education Team

Product ·

The short answer: Most course evaluation asks a yes/no question — did students like it? did the change work? — and answers it with an average. Realist evaluation, developed by Ray Pawson and Nick Tilley in Realistic Evaluation (Sage, 1997), argues that the average hides the only thing worth knowing: a course or teaching intervention does not "work" or "not work" in general — it triggers particular mechanisms in particular contexts to produce particular outcomes, and different for different students. Realist evaluation replaces the verdict with a context–mechanism–outcome (CMO) configuration: in this context, that mechanism fired for these students, generating those outcomes. For course evaluation that reframing is the difference between a number you defend and an explanation you can act on.

The trouble with "does it work?"

A programme team introduces group projects to build teamwork. The end-of-module evaluation comes back at 3.9 — down from 4.2. Verdict: the change did not work; revert. But the average is a blend of two hidden stories. For students who had done group work before and had confident peers, the projects consolidated skills and rated well. For students in dysfunctional groups, or those carrying most of the load, the experience was miserable and rated badly. The mean of 3.9 describes no actual student and recommends the wrong action. This is the ecological fallacy meeting the evaluation report: a single figure standing in for heterogeneous mechanisms.

Pawson and Tilley's target was exactly this. Evaluation science, they argued, had inherited a "successionist" model of causation from clinical trials — did the treatment, on average, beat the control? — that suits pills better than it suits social programmes, where the same intervention lands differently depending on who receives it and where. Their alternative is generative: interventions work by offering resources or reasoning that people take up (or do not), and whether they take them up depends on context. The evaluator's job is to surface those mechanisms, not just to score the outcome.

Context, mechanism, outcome

The CMO configuration is the working unit of realist evaluation, and it is more useful than it first sounds:

  • Mechanism — the reasoning or resource an intervention triggers. Group projects do not "cause" teamwork; they may trigger mutual accountability, or free-riding resentment, or skill modelling from a stronger peer. These are the mechanisms.
  • Context — the conditions under which a mechanism fires or fails: prior experience, group composition, workload elsewhere, disciplinary norms, whether assessment rewards the process or only the product.
  • Outcome — what actually results, for whom. Not a single mean but a patterned set of outcomes across contexts.

A CMO statement reads: For students with weak prior teamwork experience (C), placing them in mixed-ability groups with individual accountability (M: peer modelling plus reduced free-riding) improved both skill and rating (O); for students already competent, the same design (M: resentment at carrying peers) depressed ratings (O). That is a testable, actionable theory. It tells you not to abandon group projects but to fix group composition and accountability — a decision the 3.9 could never have supported.

The approach has matured into a recognised standard: the RAMESES II quality and reporting guidelines (Wong and colleagues, 2017) now govern realist evaluation in health and education research, and the method is explicitly method-neutral — you can build and test CMO configurations with quantitative data, qualitative data, or both.

How this differs from the frameworks you already use

Realist evaluation is not a rebranding of models we have covered before; it sits at a different layer. Logic models and theory of change map what should happen from input to outcome, but typically as a single intended pathway; realist evaluation insists there are multiple pathways contingent on context, and treats "for whom does this pathway hold?" as the central question. The CIPP model organises what to evaluate (context, input, process, product); realist evaluation specifies a theory of causation underneath. And where Kirkpatrick's four levels climb from reaction to results, realist evaluation asks why the climb succeeds for some learners and stalls for others. It is complementary, not competing — a causal lens you lay over the framework you already run.

"But isn't this just an untestable story-telling exercise?" — the counterargument

The strongest objection to realist evaluation is that CMO configurations can degenerate into post-hoc narratives: given any result, a clever evaluator can invent a mechanism-and-context story that "explains" it, with no way to be wrong. This criticism has teeth and deserves a real answer.

The realist reply is that CMO configurations are hypotheses stated in advance and tested against patterned data, not stories told afterwards. If you theorise that mixed-ability grouping improves outcomes for students with weak prior experience via peer modelling, that is a falsifiable prediction: you should see the outcome gain concentrated in that subgroup and absent elsewhere, and you should be able to find the mechanism in what students actually say. If the pattern does not appear, the configuration is wrong and you revise it. Realist evaluation done properly is iterative theory-testing — closer to the quasi-experimental logic of ruling out rival explanations than to free interpretation. The RAMESES standards exist precisely to hold practitioners to that discipline.

Two honest limitations remain. First, realist evaluation is more demanding than computing a mean — it needs subgroup-level data and genuine qualitative depth, which many evaluation systems cannot produce. Second, it resists the thing institutions most want: a single comparable number for league tables and personnel files. That tension with the dual-purpose problem is real, and realist evaluation is firmly an improvement instrument, not an accountability one. Used to rank instructors, it would be misused.

Where Koji fits

Realist evaluation needs two things a Likert form cannot give it: the mechanism (why did this work or fail?) and the context (for whom, under what conditions?). Koji for Education is built to capture both. Its AI-moderated conversational interviews do not stop at a rating — they probe why, surfacing the reasoning that constitutes a mechanism in Pawson and Tilley's sense. Structured question types capture the contextual variables — prior experience, group role, mode of study — that let you segment outcomes by subgroup rather than averaging across them. And automatic thematic analysis reads the open-text at scale to reveal which mechanisms recur in which contexts, turning hundreds of conversations into candidate CMO configurations a teaching team can actually test.

That is the practical unlock: realist evaluation has been theoretically compelling since 1997 but operationally heavy, because harvesting mechanisms and contexts by hand does not scale. AI-moderated interviewing plus thematic analysis makes the method feasible for a whole programme, under GDPR/AVG-compliant, EU-appropriate data handling. The same conversational engine powers mechanism-and-context discovery in general user research on the main Koji platform — because "what worked, for whom, and why?" is the right question in any evaluation, not only in a classroom.

Building your first CMO configuration

Starting is less daunting than the theory suggests. Take one teaching change you have already made and one outcome you already track, then work backwards through three questions. What was the mechanism? — not what you did, but the reasoning or resource it was meant to trigger in students (did the flipped-classroom design offer more practice, more peer discussion, or more perceived risk of looking unprepared?). In what context does that mechanism fire? — for which students, with what prior preparation, under what assessment incentives. What outcome pattern would that predict? — and does the data, segmented by subgroup rather than averaged, actually show it? Write the result as a single sentence in the "in this context, that mechanism, for these students, produced this outcome" form. Your first configuration will be wrong in places; that is the method working. Refining it against the next cohort's evidence is how a course evaluation stops being a scoreboard and becomes an accumulating theory of why your teaching works when it works.

The takeaway

The department average is not a small version of the truth; it is a different kind of statement, and usually the wrong one. Realist evaluation gives course evaluation a better question — what worked, for whom, in what circumstances, and through what mechanism? — and a structure for answering it that you can test and act on. A satisfaction score tells you a change scored 3.9. A CMO configuration tells you what to change about it. Only one of those is worth collecting.