Responsive Evaluation (Stake): Letting Stakeholder Issues Drive Course Evaluation
Robert Stake''s responsive evaluation begins not with pre-set objectives but with the issues and concerns that stakeholders actually raise, using emergent, largely qualitative design. We explain the model, its evidence base, its limits, and how it maps onto conversational course evaluation.
Koji Education Team
Product
In brief
Responsive evaluation, developed by Robert Stake, starts from the concerns and issues that stakeholders actually raise rather than from a fixed list of pre-specified objectives, and it lets the evaluation design emerge as those issues come into focus. Where a conventional objectives-based evaluation asks "did the course meet its stated learning outcomes?", a responsive evaluation asks "what do students, teachers, and programme leaders find important, puzzling, or troubling about this course, and how can we render that faithfully to those who must act on it?" It privileges naturalistic, largely qualitative inquiry, continuous dialogue with stakeholders, and communication in forms audiences can actually use. It is a strong fit for formative, improvement-oriented course evaluation — and a poor fit when the sole requirement is a standardised, comparable number.
What the model says
Stake introduced the approach in Program Evaluation, Particularly Responsive Evaluation (1975), extending his earlier "countenance" model. His well-known contrast is between preordinate evaluation — which fixes objectives, instruments, and hypotheses in advance — and responsive evaluation, which "orients more directly to program activities than to program intents; responds to audience requirements for information; and refers to the different value perspectives present" in reporting success and failure. In Stake''s account, the evaluator does not arrive with a questionnaire and a scoring rubric; the evaluator arrives, observes the program, talks to the people involved, and lets the salient issues — the points of contention, worry, or curiosity that stakeholders themselves name — become the organising spine of the study.
Several features follow from this starting point. The design is emergent: what gets measured is adjusted as the evaluator learns which issues matter, rather than frozen at the outset. The methods are predominantly qualitative and naturalistic — observation, conversation, document review, narrative portrayal — though Stake did not exclude quantitative evidence where it served an issue. The evaluator attends to multiple, divergent value perspectives rather than assuming a single criterion of merit, and reports in audience-appropriate forms (vignettes, portrayals, thematic accounts) designed to be understood and used, not merely archived.
Responsive evaluation belongs to a family of stakeholder-oriented and constructivist approaches. Guba and Lincoln (1989) built directly on it in Fourth Generation Evaluation, organising inquiry around stakeholders'' claims, concerns, and issues and treating evaluation findings as negotiated constructions rather than objective readings. Tineke Abma has developed the practical and political dimensions of the approach: in The Practice and Politics of Responsive Evaluation (2006, American Journal of Evaluation) she shows how creating dialogue among stakeholders with unequal power is the central craft — and the central difficulty — of doing responsive evaluation well. The approach has been applied in higher and professional education; for example, responsive-evaluation methods have been used to study medical-education programmes where standardised outcome measures failed to capture what teachers and trainees actually experienced.
Why it matters for course evaluation in practice
Most institutional course evaluation is thoroughly preordinate: a fixed Likert instrument, administered identically to every course, generating comparable means. That design answers accountability questions ("is this course above or below the faculty mean?") but is weak at the improvement questions programme directors actually care about ("why are final-year students disengaging from the capstone, and what would they change?"). A preordinate form can only surface issues its authors anticipated; it is structurally blind to the concern nobody thought to ask about.
Responsive evaluation is the corrective. In practice it means: before or alongside the standard survey, let students and staff name the issues; build the evaluation around those; and report back in a form that provokes action rather than filing. It is especially valuable for new or redesigned modules, problem courses where the numbers are bad but the cause is unclear, and programme-level review, where the relevant questions are emergent and contested. It also directly serves the "closing the loop" obligation in the European Standards and Guidelines: an evaluation organised around stakeholders'' own issues produces findings stakeholders recognise as theirs, which is a precondition for them acting on the results.
Critically, responsive and preordinate evaluation are complements, not rivals. The comparable annual number has a legitimate accountability role; responsive inquiry is what you reach for when you need to understand rather than merely rank.
A useful way to operationalise responsive evaluation without abandoning the annual survey is to sequence the two. Begin a redesign or a problem-course review with a short, open, issue-eliciting phase — a handful of conversational interviews or an open prompt asking students and teaching staff what they find most important, confusing, or frustrating about the course. Let the recurring issues from that phase define the questions the fuller evaluation then pursues, and design the reporting around those issues so that the people who raised them recognise their concerns in the output. This inverts the usual order, in which the instrument is fixed first and stakeholder concerns are squeezed into whatever boxes happen to exist. It also reframes the perennial anxiety about response rates: a responsive study does not need a representative sample of a fixed population to be valid, because its warrant is the depth and authenticity of the issues surfaced, not the precision of a mean. That makes it robust in exactly the small, engaged, or hard-to-survey settings — capstones, laboratories, doctoral cohorts — where conventional response-rate thresholds are hardest to meet.
Limitations and honest caveats
A rigorous reader should not romanticise the approach.
Comparability and standardisation are sacrificed by design. Emergent, issue-driven inquiry does not produce numbers you can benchmark across 400 courses or trend cleanly year over year. For the accountability uses that dominate institutional evaluation — and for accreditation panels that expect systematic, comparable evidence — responsive evaluation cannot be the whole system.
Resource intensity. Observation, dialogue, and narrative portrayal are labour-intensive relative to auto-scored surveys. Applied naïvely to every course every term, responsive evaluation does not scale, which is precisely why it has remained a specialist method rather than routine practice.
Subjectivity, credibility, and the politics of voice. Because the evaluator selects issues and portrays them, responsive evaluation is vulnerable to charges of bias and of privileging the loudest or most powerful voices. Abma''s work is candid that stakeholders enter the dialogue with unequal power, and that a poorly facilitated process can amplify rather than correct that inequality. Establishing trustworthiness (credibility, dependability, confirmability) requires deliberate methodological discipline — member checking, triangulation, transparent audit trails — that is easy to skip.
Weak on generalisation. Findings are local and contextual by design. That is a feature for improving this course but a limitation when leaders want defensible claims about teaching quality across a faculty.
How Koji incorporates this
Koji for Education is not a pure responsive-evaluation instrument — it must also serve the standardised, comparable needs of institutional QA — but several of its mechanisms operationalise responsive principles at a scale Stake''s labour-intensive method never could. We frame these as designed to support responsive inquiry, not to replace the evaluator''s judgement.
Issue-led conversation instead of fixed items only. Koji''s AI-moderated conversational interviews let a student raise and elaborate the issues they find salient, not only respond to pre-set scales. Adaptive, neutral follow-up probes pursue an emerging concern — the responsive move of letting the issue, not the instrument, lead — while the interview stays bounded and auditable.
Multiple value perspectives, rendered faithfully. Automatic thematic analysis of open text surfaces divergent perspectives across a cohort rather than collapsing them into a mean, echoing Stake''s insistence on representing multiple, sometimes conflicting, values. Structured question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) let a team triangulate a naturalistic account against comparable numbers, so responsive and preordinate evidence coexist in one cycle.
Emergent design within a scalable frame. Mid-cycle and formative collection lets issues identified early reshape what is asked later — an emergent design applied across many courses at once. Closing-the-loop action tracking helps ensure that stakeholder-named issues are reported back in a usable form and followed to resolution, addressing the "so what happened?" gap responsive evaluation cares about.
Guarding against the politics-of-voice risk. Because Abma''s caution about unequal power is real, Koji''s anonymity and bias-aware reporting are designed to give quieter or lower-status voices weight they might lose in a facilitated meeting, and to report distributions rather than a single dominant narrative.
Koji''s core research platform at koji.so applies the same issue-led, AI-moderated interview engine to product and customer research, where responsive, stakeholder-driven inquiry is equally valuable.
Related resources
- /docs/realist-evaluation-cmo-what-works-for-whom-course-evaluation
- /docs/utilization-focused-evaluation-designing-course-feedback-for-use
- /docs/closing-the-feedback-loop-course-evaluation-evidence
- /docs/students-willing-but-doubt-feedback-is-used-spencer-schmelkin
- /docs/framework-analysis-open-text-course-feedback
- /docs/mid-semester-feedback-consultation-meta-analysis
References
- Stake, R. E. (1975). Program Evaluation, Particularly Responsive Evaluation (Occasional Paper No. 5). Kalamazoo, MI: Evaluation Center, Western Michigan University.
- Stake, R. E. (2004). Standards-Based and Responsive Evaluation. Thousand Oaks, CA: SAGE. https://doi.org/10.4135/9781412985932
- Guba, E. G., & Lincoln, Y. S. (1989). Fourth Generation Evaluation. Newbury Park, CA: SAGE.
- Abma, T. A. (2006). The practice and politics of responsive evaluation. American Journal of Evaluation, 27(1), 31–43. https://doi.org/10.1177/1098214005283189
- Abma, T. A. (2005). Responsive evaluation: Its meaning and special contribution to health promotion. Evaluation and Program Planning, 28(3), 279–289. https://doi.org/10.1016/j.evalprogplan.2005.04.003
Related articles
The Framework Method for Open-Text Course Feedback: A Structured Alternative to Thematic Coding
When a quality committee — not a lone qualitative researcher — has to make sense of thousands of student comments, the Framework Method offers a transparent, auditable matrix-based approach. What Gale et al. (2013) actually proposed, and how to use it.
Does Closing the Feedback Loop Actually Matter? The Evidence on Acting on Student Evaluations
Universities are good at collecting student feedback and bad at acting on it visibly. The research — Watson (2003), Leckey & Neill (2001), Shah et al. (2017) — shows that failing to close the loop drives the scepticism and declining response rates that quietly destroy your evaluation data.
Students Are Willing to Evaluate — They Just Doubt Anyone Listens: The Spencer & Schmelkin Evidence
Spencer and Schmelkin (2002) surveyed students about how they view course evaluations and found a clear pattern: students are generally willing to participate but have little confidence their feedback is actually used. That belief, not apathy, is the lever behind response rates and answer quality.
Designing Course Evaluation for Use: The Utilization-Focused Approach
The biggest failure of course evaluation is not bad data — it is data nobody acts on. Patton's Utilization-Focused Evaluation and the empirical research on evaluation use (Johnson et al. 2009) show how to design feedback for action from the start.