Contribution Analysis: Making Credible Causal Claims from Course Evaluation
You changed a course and outcomes improved — but did the teaching cause it? Contribution analysis is a theory-based method for making defensible causal claims when a randomised trial is impossible. Here is how it applies to course evaluation.
Koji Education Team
Product
In brief. When a programme team changes its teaching and student outcomes improve, the honest question is whether the teaching caused the improvement or whether cohort differences, grade inflation, other courses or the labour market did. Randomised trials are usually impossible in curriculum change, so contribution analysis (Mayne) offers a disciplined alternative: build an explicit theory of change, surface rival explanations, gather evidence that tests each causal link, and end with a confidence-rated contribution claim rather than a naive before/after story. It reframes the goal from "proving impact" to "assessing how confident we can reasonably be that our teaching made a difference."
The attribution problem QA offices quietly ignore
Closing the feedback loop is meant to end with evidence that an action worked. A department redesigns assessment, response scores rise the following year, and the annual report declares success. But a rise in scores or outcomes after a change is not evidence the change produced it. The cohort may have been stronger on entry. Grading may have drifted. A parallel employability initiative may have run at the same time. The graduate labour market may have improved. Regression to the mean alone can manufacture an apparent gain after a bad year.
The gold standard for causal claims — a randomised controlled trial — is rarely available for curriculum change: you cannot ethically or practically randomise students to "good teaching" and "poor teaching," and even quasi-experimental designs are often infeasible at programme level. So evaluators are left needing to make a causal claim without an experiment. Contribution analysis is built precisely for that situation.
What the research says
Contribution analysis was developed by John Mayne and matured over a decade of use in public-sector and international evaluation. In Contribution Analysis: Coming of Age? (Evaluation, 2012, 18(3), 270–280), Mayne sets out the philosophical core: most interventions are neither necessary nor sufficient on their own to produce an outcome; they are contributory causes — one part of a package of causes that, working together, produce the result. A course is exactly this kind of cause. It does not single-handedly produce a graduate's employability; it contributes alongside prior ability, other modules, work experience and the labour market. Mayne's insight is that this is not a reason to give up on causal claims, but a reason to make them carefully and in proportion.
The method proceeds through a recognisable sequence of steps: (1) set out the attribution problem precisely; (2) develop the postulated theory of change, including the assumptions and risks to it — crucially, the plausible rival explanations; (3) gather the existing evidence on each link in that theory; (4) assemble the contribution story and honestly assess its weaknesses and the strength of the alternative explanations; (5) seek out additional evidence to fill the gaps and test the rivals; and (6) revise and strengthen the story. The output is not a p-value but a reasoned, evidence-backed claim of the form: "It is reasonable to conclude that the intervention made a meaningful contribution to the outcome, given that its theory of change is supported, its assumptions held, and the main alternative explanations have been examined and found wanting."
Mayne's earlier foundational work, Addressing attribution through contribution analysis: Using performance measures sensibly (Canadian Journal of Program Evaluation, 2001, 16(1), 1–24), introduced the approach as a way to use ordinary monitoring data to support — rather than merely assert — causal claims.
The method has since been sharpened by combining it with process tracing. Befani and Mayne (2014), in Process Tracing and Contribution Analysis: A Combined Approach to Generative Causal Inference for Impact Evaluation (IDS Bulletin, 45(6), 17–36), show how process-tracing "tests" (does a piece of evidence, if found, sharply raise or lower our confidence?) can be applied to the links in a contribution story. Their reframing is the memorable one: the combined approach shifts the focus of impact evaluation "from assessing impact to assessing confidence about impact." Notably, they tested the approach on the evaluation of a teaching programme's contribution to improving school performance — an education example, not a stretch to our domain.
Why it matters for course evaluation in practice
Contribution analysis gives programme-level quality assurance a rigorous backbone for the "did it work?" question:
-
It forces a theory of change before you claim success. Instead of "we changed assessment and scores went up," you articulate why the change should improve learning — the mechanism, the assumptions, the intermediate steps. That theory is then something evidence can support or undermine.
-
It makes rival explanations part of the method, not an afterthought. A credible contribution story explicitly lists and tests the alternatives — stronger cohort, grade inflation, concurrent initiatives, labour-market shifts — rather than hoping a reviewer will not raise them. This is exactly the intellectual honesty a critical accreditation panel rewards.
-
It fits lagging, confounded outcomes. Employability and long-run learning gains arrive late and are influenced by everything. Contribution analysis is designed for causes that are contributory rather than decisive, so it is well suited to claims about graduate outcomes where a clean experiment is impossible.
-
It produces calibrated, defensible claims. The end product states how confident you are and why, which is far more durable under scrutiny than an unqualified "our intervention improved outcomes."
Limitations and honest caveats
- It is not a substitute for a true counterfactual. Contribution analysis strengthens causal reasoning but does not deliver the clean effect size of a randomised experiment. Where an RCT or a strong quasi-experimental design (interrupted time series, difference-in-differences) is feasible, it provides more decisive evidence; contribution analysis is the tool for when it is not.
- It is labour-intensive and judgement-heavy. Building and testing a theory of change, gathering evidence for each link, and adjudicating rival explanations takes time and skill. It is disproportionate for a single course's routine feedback and belongs at programme or intervention level.
- Confirmation bias is a live risk. Because the evaluator often has a stake in the intervention, there is a temptation to build a story that flatters it and to under-test the rivals. The method's credibility depends entirely on treating alternative explanations seriously — ideally with independent challenge.
- Evidence quality caps confidence. A contribution story assembled from weak monitoring data yields a weak claim. Garbage links produce a garbage conclusion, however elegant the theory of change.
- "Confidence" is not certainty. Befani and Mayne's reframing is a strength but also a limit: the output is a defensible degree of belief, not proof. Stakeholders who want a single causal number may find that unsatisfying, and it must be communicated carefully.
How Koji incorporates this
Koji does not replace an evaluator's judgement, but it supplies the evidence base a contribution story needs — and it is built to gather the kind of mechanism-level evidence that theory-of-change testing requires.
- Evidence for each link, not just the endpoint. Because Koji's AI-moderated interviews probe why and how — asking students what specifically changed in their learning experience after a redesign — they generate evidence about the intermediate steps of a theory of change, not merely a satisfaction number at the end. That is exactly what testing a causal link requires.
- Surfacing rival explanations from the students themselves. Conversational follow-ups can elicit the alternative drivers students attribute their experience to (a concurrent module, a placement, prior preparation), giving evaluators empirical material to weigh the rival explanations contribution analysis demands.
- Triangulation across cohorts and time. Koji's reporting supports comparison across cohorts and mid-cycle collection, so a contribution story can draw on trend evidence and on before/after change described in students' own words — strengthening or weakening specific links.
- Thematic analysis that maps to theory-of-change steps. Automatic thematic analysis of open text can be organised around the postulated mechanisms, turning hundreds of comments into structured evidence for or against each assumption rather than an undifferentiated word cloud.
- Honest, confidence-oriented reporting. Koji is designed to present findings with their uncertainty rather than as false precision, aligning with the "assessing confidence about impact" stance and with closing-the-loop practice that asks whether an action actually worked.
The same interview engine underpins Koji's core research platform at koji.so, where product teams face the identical challenge of claiming a feature contributed to an outcome amid many confounding influences.
FAQ
How is contribution analysis different from just showing before-and-after scores? A before/after comparison shows correlation with a change, not causation. Contribution analysis requires an explicit theory of change, gathers evidence for each causal link, and actively tests rival explanations — ending with a confidence-rated claim rather than an unqualified assertion of impact.
When should I use it instead of a quasi-experiment? Use a quasi-experimental design (interrupted time series, difference-in-differences) when a credible comparison or counterfactual is available, because it gives more decisive evidence. Use contribution analysis when randomisation and strong comparisons are infeasible, which is common for programme-level curriculum change.
What is a "contributory cause"? A cause that is neither necessary nor sufficient on its own but is a genuine part of the package of causes that together produce the outcome. A course typically contributes to graduate outcomes alongside prior ability, other modules and the labour market — so its effect is contributory, not decisive.
Isn't this just telling a persuasive story? Only if done badly. The method's credibility rests on explicitly listing and testing rival explanations and on the quality of evidence for each link. Process-tracing tests (Befani & Mayne) add rigour by specifying what evidence would raise or lower confidence in each claim.
At what level should we apply it — course or programme? Programme or intervention level. It is labour-intensive and judgement-heavy, so it is disproportionate for a single course's routine feedback but well suited to evaluating a substantial teaching change or a programme-wide outcome claim.
Related Resources
- Realist evaluation: what works, for whom, in what circumstances
- Developmental evaluation for course innovation under complexity
- Goal-free evaluation and unintended side effects
- Interrupted time series: did a teaching change actually move outcomes?
- Closing the feedback loop: the evidence
- Utilization-focused evaluation: designing course feedback for use
References
- Mayne, J. (2001). Addressing attribution through contribution analysis: Using performance measures sensibly. The Canadian Journal of Program Evaluation, 16(1), 1–24. https://evaluationcanada.ca/system/files/cjpe-entries/16-1-001.pdf
- Mayne, J. (2012). Contribution analysis: Coming of age? Evaluation, 18(3), 270–280. https://doi.org/10.1177/1356389012451663
- Befani, B., & Mayne, J. (2014). Process tracing and contribution analysis: A combined approach to generative causal inference for impact evaluation. IDS Bulletin, 45(6), 17–36. https://doi.org/10.1111/1759-5436.12110
Related articles
Did the Teaching Change Actually Work? Interrupted Time Series for Course-Evaluation Trends
Comparing this year''s evaluation mean to last year''s cannot tell you whether a curriculum redesign worked. Interrupted time series with segmented regression can — here is how to apply it, and where it breaks.
Does Closing the Feedback Loop Actually Matter? The Evidence on Acting on Student Evaluations
Universities are good at collecting student feedback and bad at acting on it visibly. The research — Watson (2003), Leckey & Neill (2001), Shah et al. (2017) — shows that failing to close the loop drives the scepticism and declining response rates that quietly destroy your evaluation data.
Designing Course Evaluation for Use: The Utilization-Focused Approach
The biggest failure of course evaluation is not bad data — it is data nobody acts on. Patton's Utilization-Focused Evaluation and the empirical research on evaluation use (Johnson et al. 2009) show how to design feedback for action from the start.
Developmental Evaluation (Patton): Evaluating Course Innovation Under Uncertainty
Michael Quinn Patton''s developmental evaluation supports innovation in complex, fast-changing conditions by feeding rapid, real-time evidence back into ongoing design rather than judging a fixed programme against fixed outcomes. We explain the model, its evidence base, its limits, and how it maps onto continuous course-improvement cycles.