New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
best-practices9 min read

Designing Course Evaluation for Use: The Utilization-Focused Approach

The biggest failure of course evaluation is not bad data — it is data nobody acts on. Patton's Utilization-Focused Evaluation and the empirical research on evaluation use (Johnson et al. 2009) show how to design feedback for action from the start.

Koji Education Team

Product

In brief: The most common failure in course evaluation is not measurement error — it is non-use: surveys are run, scores are filed, and nothing changes, which in turn teaches students that responding is pointless. Utilization-Focused Evaluation (UFE), developed by Michael Quinn Patton, reframes the goal: an evaluation should be judged by whether its findings are actually used, and use is engineered from the start by identifying primary intended users and their intended uses before a single question is written. The empirical research on evaluation use (Johnson et al., 2009) supports the core mechanism — stakeholder involvement is among the strongest predictors of whether evaluation findings get used. For European quality assurance, UFE turns "closing the loop" from an afterthought into the design principle.

What the research says

Decades of evaluation scholarship document a stubborn paradox: organisations invest heavily in collecting evaluation data, then fail to use it. Patton's Utilization-Focused Evaluation is the most influential response. Its premise is deceptively simple — evaluations should be judged by their use — and its method follows directly: identify the primary intended users (the specific people who can and will act) and their intended uses at the outset, and let those drive every subsequent decision about design, instruments, analysis and reporting. Patton's argument, grounded in research on evaluation, is that "intended users are more likely to use evaluations if they understand and feel ownership of the evaluation process and findings," and they are more likely to feel ownership if they have been actively involved in shaping it.

This is not merely an assertion of preference. Johnson, Greenseid, Toal, King, Lawrenz & Volkov (2009), "Research on Evaluation Use: A Review of the Empirical Literature From 1986 to 2005," in the American Journal of Evaluation, systematically reviewed 41 empirical studies of evaluation use that met quality standards. Building on Cousins & Leithwood's (1986) foundational framework, they confirmed that use is shaped by factors such as evaluation quality, credibility, relevance, communication and timeliness — and they added stakeholder involvement as a distinct category and evaluator competence as a key characteristic. The headline for practitioners: whether people are involved in and feel ownership of an evaluation is one of the most consistent correlates of whether its findings get acted upon.

This evidence base matters for course evaluation specifically because the SET literature documents the symptom. Spencer & Schmelkin (2002), in Assessment & Evaluation in Higher Education, found that students are generally willing to evaluate teaching but doubt that their feedback is meaningfully used — a perception that depresses both response rates and candour. When students sense the loop is open, they disengage; when they see change, they invest. UFE provides the design logic that turns that dynamic around.

European quality frameworks already encode the expectation of use without prescribing the method. The Standards and Guidelines for Quality Assurance in the European Higher Education Area (ESG, 2015) require institutions to collect, analyse and act on information about their programmes (notably under ongoing monitoring and information-management standards). UFE is, in effect, an operational answer to the ESG question "how do we make sure feedback leads to action?"

Why it matters for course evaluation in practice

Adopting a utilization-focused mindset changes concrete decisions across the evaluation cycle.

  1. Start with the decision, not the questionnaire. Before drafting items, name the primary intended users (the module leader, the programme committee, the QA officer) and the decisions they will make (revise assessment? change sequencing? invest in teaching support?). Every question then earns its place by informing a real decision. This single discipline eliminates the bloated, generic forms that produce data nobody owns.

  2. Design for timeliness. Findings used for decisions must arrive before the decision. An end-of-semester survey whose results land after next term is timetabled is, in UFE terms, designed to fail. This is the strongest argument for mid-cycle / formative collection, where there is still time to act for the very cohort that gave the feedback.

  3. Engineer ownership to raise response and candour. Involving students and teachers in what is asked — and visibly reporting what changed — builds the ownership that the use literature identifies as decisive. "You said, we did" is not PR; it is the mechanism by which the next round of data becomes both more plentiful and more honest.

  4. Match the report to the user. A 40-page statistical appendix is unused by a busy programme director. UFE insists the format of findings be shaped to how the intended user will actually consume and act on them — short, prioritised, decision-linked.

The net effect is to move course evaluation from a compliance ritual (we surveyed, therefore we complied) to a change instrument (we surveyed, therefore something improved).

Limitations and honest caveats

Utilization-focused evaluation is a powerful design philosophy, but it has real tensions a thoughtful QA professional should weigh.

  • The evidence base is associational. The research on evaluation use (Johnson et al., 2009; Cousins & Leithwood, 1986) is largely observational; stakeholder involvement correlates with use, but the studies vary in design and rarely permit clean causal claims. Involvement is a well-supported lever, not a guaranteed cause.
  • Stakeholder involvement can bias the evaluation. Engaging primary intended users improves ownership, but it can also narrow questions to what insiders want to hear and crowd out inconvenient findings. UFE must be balanced with independence and methodological integrity, or "use" degrades into self-justification.
  • Whose use counts? Privileging primary intended users (often staff and managers) risks under-weighting students, whose experience the evaluation is ostensibly about. A purely managerial UFE can quietly sideline the student voice it depends on.
  • Resource intensity. Genuine involvement, tailored reporting and timely turnaround cost staff time. Under-resourced QA units may struggle to do UFE well, and a half-implemented version (consultation theatre without real influence) can be worse than none.
  • Use is hard to measure. Judging an evaluation "by its use" presupposes you can observe use; in practice, influence is diffuse, delayed and entangled with other drivers of change, making the success criterion itself slippery.

The balanced view: UFE rightly centres the neglected question of action, and the evidence supports involvement and timeliness as levers — but it is a design discipline to be combined with methodological rigour and a deliberate protection of the student voice, not a licence to let stakeholders curate their own conclusions.

How Koji incorporates this

Koji is built around the conviction that evaluation exists to drive action — which is the heart of utilization-focused evaluation. Several capabilities operationalise it, framed as support for use rather than a guarantee of it.

  • Decision-linked study design. Koji encourages building evaluations around the questions a programme actually needs answered, using structured types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) so each item maps to a decision rather than padding a generic form.
  • Fast, readable synthesis for the intended user. Koji's automatic thematic analysis and report generation turn raw responses into prioritised, plain-language themes and recommendations — the report-shaped-to-the-user that UFE demands, instead of an unread statistical dump.
  • Mid-cycle and formative collection supports the timeliness principle directly: feedback can be gathered and acted on while a module is still running, so the cohort that speaks can benefit from the change.
  • Closing-the-loop / action tracking. Koji supports recording and communicating what changed in response to feedback, which is the visible "you said, we did" mechanism that builds the ownership the use literature shows to be decisive — and which lifts future response rates and candour.
  • AI-moderated interviews that surface the actionable specifics. Because use depends on findings being concrete enough to act on, Koji's conversational interviewer probes vague ratings into specific, decision-ready evidence ("the second assessment clashed with the group project deadline") that a committee can actually move on.

Koji frames all of this as designed to increase the likelihood of use — the data still has to meet a real decision, and the institution still has to act. Koji's core research platform at koji.so applies the same use-first philosophy to product and customer research, where a study that does not change a decision is a study that was not worth running.

Related resources

References

Related articles

accreditation

Turning Student Feedback into ESG / ENQA Accreditation Evidence

A buyer's guide mapping the ESG 2015 internal quality assurance standards to concrete, accreditation-ready evidence you can generate from student feedback — and how AI-moderated evaluation closes the loop.

analysis-reporting

Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees

The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.

best-practices

Does Closing the Feedback Loop Actually Matter? The Evidence on Acting on Student Evaluations

Universities are good at collecting student feedback and bad at acting on it visibly. The research — Watson (2003), Leckey & Neill (2001), Shah et al. (2017) — shows that failing to close the loop drives the scepticism and declining response rates that quietly destroy your evaluation data.

best-practices

Students Are Willing to Evaluate — They Just Doubt Anyone Listens: The Spencer & Schmelkin Evidence

Spencer and Schmelkin (2002) surveyed students about how they view course evaluations and found a clear pattern: students are generally willing to participate but have little confidence their feedback is actually used. That belief, not apathy, is the lever behind response rates and answer quality.