New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
best-practices10 min read

Developmental Evaluation (Patton): Evaluating Course Innovation Under Uncertainty

Michael Quinn Patton''s developmental evaluation supports innovation in complex, fast-changing conditions by feeding rapid, real-time evidence back into ongoing design rather than judging a fixed programme against fixed outcomes. We explain the model, its evidence base, its limits, and how it maps onto continuous course-improvement cycles.

Koji Education Team

Product

In brief

Developmental evaluation supports innovation under complexity by feeding rapid, real-time evidence back into a programme while it is still being designed, rather than judging a finished programme against fixed outcomes. Developed by Michael Quinn Patton (2011), it is purpose-built for situations where "how to get there is not yet known," stakeholders disagree, and conditions change faster than a traditional summative study can keep up with. For higher education, it is the natural evaluation logic for a genuinely new course, a redesigned curriculum, or a fast-moving intervention (for example, teaching with generative AI) where the point is not to grade this year''s version but to help the design evolve. It trades the clean, defensible verdict of summative evaluation for speed, adaptability, and usefulness — and demands unusual discipline to avoid becoming mere unstructured tinkering.

What the model says

Patton set out the approach in Developmental Evaluation: Applying Complexity Concepts to Enhance Innovation and Use (2011, Guilford Press). It grows out of his earlier utilization-focused evaluation but breaks from a core assumption of most evaluation designs: that there is a stable, well-defined "program" whose fixed outcomes can be measured. In complex, emergent conditions, Patton argues, that assumption fails — the intervention is still mutating, causal pathways are non-linear, and the "model" you would evaluate against does not yet exist.

Patton defines complex situations by three features: high uncertainty about how to achieve desired results, disagreement among key stakeholders about what to do, and many interacting factors in a dynamic environment that defeat prediction and static models. In such conditions, developmental evaluation positions the evaluator inside the innovation team, providing timely, rigorous feedback that informs the next design decision. The evaluator''s role shifts from independent judge delivering a verdict to embedded sense-maker helping the team learn its way forward — while still bringing evaluative thinking, data discipline, and a check against self-serving interpretation.

Patton distinguishes developmental evaluation from formative evaluation, and this distinction matters. Formative evaluation improves a model on its way to a fixed, summative-ready form — it assumes there is a stable model being perfected. Developmental evaluation supports ongoing development where the model itself keeps changing and may never stabilise, because the environment keeps moving. He identifies several use situations: ongoing development of an initiative, adapting effective principles to new contexts, developing rapid responses in crisis, and pre-formative development of a new idea toward something that could later be evaluated more conventionally.

The approach draws explicitly on complexity science — emergence, non-linearity, adaptive systems — and on systems thinking. It has been elaborated with practical exemplars: Patton, McKegg and Wehipeihana (2016) collected worked cases in Developmental Evaluation Exemplars, and Gamble''s A Developmental Evaluation Primer (2008, J. W. McConnell Family Foundation) gives an accessible practitioner account. These sources also register the recurring critique that the approach can blur the line between evaluation and management consulting if rigour is not deliberately protected.

Why it matters for course evaluation in practice

Standard course evaluation is implicitly summative: administer a fixed instrument at the end of term, compare the mean to a benchmark, file the result. That logic works for a stable, repeatedly-offered module. It is actively unhelpful for the situations universities increasingly face: a brand-new programme in its first delivery, a curriculum being redesigned around competences or work-integrated learning, or teaching that must adapt within weeks to a fast-moving technology such as generative AI. Waiting until the end of term to discover that the design did not work — when the cohort has already moved on and the next redesign is already underway — wastes the most valuable feedback.

Developmental evaluation reframes the task. For an innovating course, the question is not "was this course good?" but "what is this course teaching us about how to build the next iteration, and how fast can we feed that back?" In practice this means rapid, iterative collection during delivery; tight coupling between evidence and the design team''s decisions; and comfort with an evolving, rather than fixed, set of questions. It aligns with the quality-culture conception the European University Association contrasts with mere quality-assurance: quality as continuous, owned, developmental improvement rather than periodic external compliance. It complements — does not replace — the summative annual survey the institution still owes its accreditors.

A concrete higher-education case makes the logic tangible. Consider a department introducing generative-AI tools into an assessment for the first time. There is no stable model to evaluate against: the assessment design, the guidance to students, and the marking approach are all being invented in real time, and last month''s decisions are already obsolete. A summative end-of-term survey would deliver its verdict long after the decisions that mattered were made. A developmental approach instead collects short bursts of student and staff feedback after each major milestone — the first AI-assisted task, the first marking round — and feeds each finding straight into the next design choice. It is worth distinguishing this from action research, with which it overlaps: developmental evaluation keeps an explicit evaluative stance and data discipline, whereas action research centres practitioner-led cycles of change. The two are cousins rather than twins, and Patton is careful to preserve the evaluator''s independent, evidence-testing role even while embedded in the team.

Limitations and honest caveats

It is not a summative verdict, and must not be sold as one. Developmental evaluation is designed to help a programme develop, not to prove it worked. Using it where a defensible, comparable judgement is required — accreditation evidence, personnel decisions, resource allocation between courses — is a category error. Institutions need both logics and must not substitute one for the other.

Rigour is fragile. Embedding the evaluator in the innovation team creates a real risk of capture: the evaluator becomes an advocate, confirmation bias creeps in, and "rapid feedback" degrades into cheerleading. The exemplar literature is explicit that developmental evaluation demands more methodological self-discipline than summative work, not less — triangulation, disconfirming evidence, and transparent reasoning — precisely because its independence is structurally weaker.

Weak generalisation and comparability. Findings are tightly bound to a specific, evolving context. They rarely support claims that generalise across courses or over time, and they resist benchmarking by construction.

Demanding of people and data infrastructure. Real-time, iterative evaluation requires fast data collection and analysis and a team culture that can act on feedback between iterations. Without that infrastructure the approach stalls; a slow feedback loop defeats the entire purpose.

Evidence base is largely case-based. The support for developmental evaluation is primarily theoretical and drawn from documented exemplars rather than controlled comparative trials showing it outperforms alternatives. That is appropriate to its subject matter but should be stated plainly: it is a well-theorised, practitioner-validated framework, not an empirically "proven-superior" method.

How Koji incorporates this

Koji for Education is well suited to the data-and-feedback engine that developmental evaluation requires, provided institutions keep its developmental use distinct from its summative use. We frame these mechanisms as designed to support rapid, rigorous development.

Rapid, mid-cycle, iterative collection. Developmental evaluation lives or dies on feedback speed. Koji supports mid-cycle and formative collection, so an innovating course team can gather student evidence during delivery and feed it into the next iteration rather than waiting for an end-of-term summative survey. Adaptive AI-moderated conversational interviews surface emergent issues — the concern nobody anticipated — which is exactly what a design team navigating uncertainty needs.

Evolving questions within a rigorous frame. Because developmental evaluation''s questions change as the innovation evolves, Koji''s flexible structured question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) and reconfigurable briefs let a team adjust what it asks between iterations while automatic thematic analysis and quality scoring keep the analysis disciplined — a partial guard against the "capture and cheerleading" risk by anchoring interpretation in the actual data.

Triangulation and disconfirmation. To counter confirmation bias, Koji reports distributions and divergent themes rather than a single agreeable narrative, and lets teams triangulate conversational evidence against comparable numeric items — supporting the hunt for disconfirming evidence that keeps developmental work honest.

Closing the loop across iterations. Action tracking helps a team record what each round of feedback changed and whether the change worked, building the iterative learning trail that distinguishes genuine developmental evaluation from undocumented tinkering.

Because Koji also produces standardised, comparable evidence, an institution can run a developmental cycle on an innovating course while still generating the summative record its accreditors expect — keeping the two logics separate but in one system. Koji''s core research platform at koji.so applies the same rapid, iterative, AI-moderated interview engine to product discovery and innovation research, the commercial analogue of developmental evaluation.

Related resources

References

  • Patton, M. Q. (2011). Developmental Evaluation: Applying Complexity Concepts to Enhance Innovation and Use. New York: Guilford Press.
  • Patton, M. Q., McKegg, K., & Wehipeihana, N. (Eds.). (2016). Developmental Evaluation Exemplars: Principles in Practice. New York: Guilford Press.
  • Gamble, J. A. A. (2008). A Developmental Evaluation Primer. Montreal: J. W. McConnell Family Foundation.
  • Patton, M. Q. (1994). Developmental evaluation. Evaluation Practice, 15(3), 311–319. https://doi.org/10.1177/109821409401500312
  • Preskill, H., & Beer, T. (2012). Evaluating Social Innovation. FSG / Center for Evaluation Innovation.

Related articles

best-practices

Does Closing the Feedback Loop Actually Matter? The Evidence on Acting on Student Evaluations

Universities are good at collecting student feedback and bad at acting on it visibly. The research — Watson (2003), Leckey & Neill (2001), Shah et al. (2017) — shows that failing to close the loop drives the scepticism and declining response rates that quietly destroy your evaluation data.

best-practices

Do Student Evaluations Actually Improve Teaching? The Feedback-Intervention Evidence

Student ratings handed back as a bare number barely change teaching (d ≈ 0.2). Paired with structured consultation, the same data produces moderate, durable improvement (d ≈ 0.6–0.7). What the intervention meta-analyses tell quality teams about closing the loop.

best-practices

Students Are Willing to Evaluate — They Just Doubt Anyone Listens: The Spencer & Schmelkin Evidence

Spencer and Schmelkin (2002) surveyed students about how they view course evaluations and found a clear pattern: students are generally willing to participate but have little confidence their feedback is actually used. That belief, not apathy, is the lever behind response rates and answer quality.

best-practices

Designing Course Evaluation for Use: The Utilization-Focused Approach

The biggest failure of course evaluation is not bad data — it is data nobody acts on. Patton's Utilization-Focused Evaluation and the empirical research on evaluation use (Johnson et al. 2009) show how to design feedback for action from the start.