New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Methodology7 min read

Logic Models and Theory of Change: Making the Hidden Assumptions in Course Evaluation Explicit

Most course evaluation measures outputs and outcomes without ever stating why anyone expected them to be connected. A logic model — and the theory of change behind it — forces those assumptions into the open, where they can be tested.

Koji Education Team

Product ·

Bottom line up front: A logic model is a simple chain — inputs → activities → outputs → outcomes → impact — and a theory of change is the explanation of why you believe each arrow holds. Course evaluation that skips this step measures whatever is convenient and then guesses at causation. Building the logic model first tells you what to measure and, more importantly, exposes the untested assumptions on which your programme quietly depends. The catch: a logic model is a hypothesis, not proof, and treating it as proof is the most common way it misleads.

The problem: evaluation without an articulated theory

Ask a programme team why their redesigned module should improve graduate readiness and you will often get a confident answer that, written down, turns out to be a series of unexamined leaps: we added group projects (activity) → students collaborate more (output) → they develop teamwork skills (outcome) → employers value them more (impact). Each arrow is an assumption. Some are well-evidenced; some are wishful. Standard course evaluation never names them, so when the satisfaction scores come back fine but graduate outcomes do not move, no one can say which link in the chain broke.

This is the gap that program-evaluation theorists set out to close. Carol Weiss popularised the term "theory of change" in the mid-1990s, arguing that complex initiatives fail to demonstrate impact largely because their underlying assumptions are "poorly articulated" — evaluators measure the start and the end and ignore the mechanism in between. The W.K. Kellogg Foundation's widely used Logic Model Development Guide (2004) operationalised this: a logic model is "a picture of how your organisation does its work — the theory and assumptions underlying the program," linking activities to outcomes through the principles that connect them. As Kellogg put it, the theory of change is "what puts the logic in a logic model."

The two artefacts, and how they differ

The terms are often conflated; the distinction is useful.

  • A logic model is the what: a structured, usually linear map of inputs, activities, outputs, short- and medium-term outcomes, and long-term impact. It is a planning and communication tool.
  • A theory of change is the why and how: the causal reasoning, the assumptions, and the contextual conditions that have to hold for each step to lead to the next. Funnell and Rogers (Purposeful Program Theory, 2011) emphasise that a good theory of change makes the assumptions and the necessary conditions explicit and, ideally, testable.

For evaluation, the theory of change is the more important of the two, because it tells you which assumptions to put under empirical pressure. The logic model tells you what to collect; the theory of change tells you what would falsify your beliefs.

Why this matters for course and programme evaluation specifically

Three concrete payoffs for a quality-assurance or institutional-research team:

  1. It defines your indicators before you collect data. Instead of fielding a generic satisfaction survey and reverse-engineering a story, you derive each evaluation item from a specific link in the chain. If the theory says "structured peer feedback → improved revision behaviour → better final work," you measure revision behaviour, not just satisfaction with peer feedback.

  2. It locates failure. When outcomes disappoint, a logic model lets you ask which arrow broke? Did the activity not happen as designed (an implementation failure, in CIPP terms a Process problem)? Or did it happen but not produce the expected outcome (a theory failure)? These demand completely different responses, and you cannot tell them apart without the model.

  3. It surfaces the assumptions you would rather not examine. The "students will transfer this skill to the workplace" arrow is usually the weakest and least evidenced link in any employability claim — and a logic model puts it on the page where someone can ask for evidence. Our analysis of why course evaluation cannot directly measure employability is essentially an argument about one very long, very assumption-laden arrow.

"But isn't this just a linear oversimplification of messy reality?" — the counterargument

The strongest critique of logic models is real and worth stating plainly.

Education is non-linear, and logic models pretend it is not. Learning is recursive, context-dependent and shaped by factors far outside the curriculum. A tidy left-to-right arrow diagram can impose false order on a genuinely complex system and create an illusion of control. Critics in the complexity-evaluation tradition (and Funnell and Rogers themselves) warn against treating the model as reality rather than as a simplified hypothesis about reality.

The honest response is twofold. First, the value of a logic model is not its accuracy but its explicitness — a wrong assumption you have written down can be tested and corrected; an unexamined one cannot. Second, theories of change can and should incorporate feedback loops, enabling conditions and alternative pathways rather than a single arrow; the linearity is a limitation of lazy logic models, not of the method. A logic model is a starting hypothesis to be revised as evidence arrives — never a finished proof of impact.

Where Koji fits

A theory of change tells you what evidence would test each assumption. The practical obstacle is that the most revealing evidence — about mechanisms, about whether students actually changed their behaviour — is qualitative and has historically been impossible to collect at scale. That is the gap Koji is built for.

Once a programme team has mapped its logic model, Koji for Education lets them instrument the specific links. Its six structured question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) let you attach a tailored measure to each step rather than a one-size-fits-all survey. Its AI-moderated conversational interviews probe mechanism — not just whether an outcome occurred but whether it occurred for the reason your theory predicted, which is exactly the evidence that distinguishes a theory failure from an implementation failure. Automatic thematic analysis then aggregates those open-text mechanisms into structured patterns across a cohort, and closing-the-loop action tracking records which assumptions the evidence revised. The same conversational engine drives mechanism-level customer research on the main Koji platform, where "did the feature work for the reason we assumed?" is the identical question.

The precise claim: Koji helps you test the arrows in your theory of change by capturing mechanism-level evidence at scale. It does not validate your theory for you, and a flawed theory instrumented well is still a flawed theory. The model is your hypothesis; Koji is how you put it under pressure.

Implementation failure versus theory failure: the distinction that pays

The single most valuable thing a logic model gives an evaluator is a clean way to separate two failure modes that look identical in the outcome data but demand opposite responses.

Suppose a redesigned module aimed to improve students' analytical writing, and the final assessment shows no improvement. Without a logic model, the team argues in circles. With one, they ask a precise sequence: Did the planned activities actually happen as designed? If the structured writing workshops were cancelled, under-attended or delivered differently, the theory was never tested — this is implementation failure, and the fix is delivery, not redesign. If the workshops ran exactly as planned and writing still did not improve, the assumption that "workshops → better writing" is wrong — this is theory failure, and the fix is to revise the theory.

Conflating the two is how programmes thrash: teams scrap sound designs that were never properly delivered, or doggedly re-deliver designs whose underlying logic is broken. The mechanism-level evidence needed to tell them apart — did the activity happen, and did it produce the predicted intermediate change? — is exactly the evidence a number-only survey cannot provide, which is why theory-driven evaluation and qualitative depth belong together.

The takeaway

Logic models and theories of change do not make evaluation more complicated — they make its hidden assumptions visible. The discipline of writing "we believe X leads to Y because Z" before collecting a single data point is what separates evaluation that can learn from evaluation that merely scores. Build the chain first; then measure the arrows you are least sure about; then revise. That loop is the whole craft.

Want to instrument the assumptions behind your programme, not just its satisfaction score? See how Koji for Education captures mechanism-level evidence at scale.