Is Your Course-Evaluation System Any Good? Meta-Evaluation with the Program Evaluation Standards
Universities evaluate courses relentlessly but rarely evaluate the evaluation itself. The Program Evaluation Standards (JCSEE) offer a validated framework — utility, feasibility, propriety, accuracy and accountability — for auditing whether your own course-evaluation system is useful, fair and sound.
Koji Education Team
Product
In brief
Institutions pour enormous effort into evaluating courses and almost none into evaluating the evaluation. Yet a course-evaluation system is itself an evaluation, and it can be judged by the same criteria as any other. The Program Evaluation Standards (Yarbrough, Shulha, Hopson & Caruthers, 2011), maintained by the Joint Committee on Standards for Educational Evaluation (JCSEE), define five attributes of a sound evaluation — utility, feasibility, propriety, accuracy, and evaluation accountability — across 30 specific standards. Running your own course-evaluation process against them is a meta-evaluation (Stufflebeam, 2001): a structured audit that asks not "what did students say?" but "is our way of asking, analysing and using their answers useful, practical, ethical, accurate, and answerable for itself?" Most systems fail predictably on utility (data collected but not used) and propriety (anonymity and fairness), and those failures — not the survey items — are usually why course evaluation loses credibility.
What the research says
Meta-evaluation — the evaluation of an evaluation — is a long-standing principle in the field. Michael Scriven coined the term, and Daniel Stufflebeam formalised it as a professional obligation in "The metaevaluation imperative," arguing that evaluations should themselves be held to explicit quality standards rather than trusted on faith (Stufflebeam, 2001). The instrument the field converged on is the JCSEE Program Evaluation Standards, now in their third edition and accredited by the American National Standards Institute (Yarbrough et al., 2011). Developed over six years with input from more than 400 stakeholders, the 30 standards are grouped into five attributes:
- Utility (8 standards) — the extent to which stakeholders find the evaluation's processes and findings valuable and actually use them. Covers stakeholder identification, relevant information, meaningful reporting, and timely dissemination that leads to use.
- Feasibility (4 standards) — whether the evaluation is practical, efficient, politically viable and resource-conscious. An evaluation that is too burdensome to run properly fails here.
- Propriety (7 standards) — what is proper, fair, legal, right and just: responsive to participants, transparent, protective of rights and confidentiality, and free of conflicts of interest.
- Accuracy (8 standards) — whether the findings are technically sound: valid and reliable information, sound designs and analyses, justified conclusions, and clear communication of uncertainty.
- Evaluation accountability (3 standards) — whether the evaluation documents itself, is internally and externally meta-evaluated, and is answerable for its own quality.
The framework's power for course evaluation is that it reorders the usual questions. Institutions obsess over accuracy — scale points, reliability coefficients, bias — while quietly failing utility and propriety. The Standards insist all five attributes matter, and that a technically accurate evaluation nobody uses (utility failure) or one that exposes students to identification (propriety failure) is a bad evaluation regardless of its psychometrics. This aligns with the utilization-focused tradition (Patton), which holds that an evaluation's worth is decided by its use, and with Linse's (2017) guidance that responsible reporting is an ethical, not merely technical, matter.
Why it matters for course evaluation in practice
A meta-evaluation turns a vague sense that "our evaluations are not landing" into a specific, fixable diagnosis. Mapping a typical university system onto the five attributes exposes recurring failure points.
Utility is where most systems die. Data is collected diligently and then does nothing: results reach instructors weeks after the cohort has left, students never learn whether their feedback changed anything, and reports are tables of means no committee can act on. The well-documented finding that students believe no one listens is, in Standards terms, a utility failure — and it is the direct cause of falling response rates. Fixing utility (timely, meaningful reporting and visible closing-the-loop) does more for a course-evaluation system than any psychometric tweak.
Propriety is where systems create risk. Small-class feedback that can re-identify a student, open-text routes for abusive comments, and using the same instrument for both formative development and high-stakes personnel decisions are all propriety violations. GDPR and duty-of-care obligations live here. A meta-evaluation forces these into view before a data-protection or wellbeing incident does.
Feasibility failures cause survey fatigue. Over-surveyed students and administrators drowning in low-value questionnaires are a feasibility problem: the system consumes more goodwill and resource than it returns. The Standards legitimise doing less, better.
Accuracy without the other four is hollow. A beautifully validated instrument delivering unused, unfair, or impractical evaluation is still a failed evaluation. The Standards stop QA teams from over-investing in accuracy while the real leaks are elsewhere.
Accountability makes the system defensible. For accreditation, being able to show that your evaluation process is itself documented and periodically meta-evaluated is powerful evidence of a mature quality culture — exactly what external reviewers look for.
Limitations and honest caveats
The Standards are a framework, not a formula, and should be applied with the same scepticism they demand.
They are consensus standards, not empirical laws. The 30 standards codify professional judgement, not experimental findings. Reasonable evaluators weight them differently, and the attributes can conflict — maximising utility (fast, actionable reporting) can strain accuracy (careful analysis takes time), and propriety (protecting confidentiality) can limit the granularity utility wants. The Standards offer no algorithm for resolving these trade-offs; they require deliberation.
North American provenance. The JCSEE is a North American body, and the Standards' examples and legal framing reflect that. European institutions operating under ESG/ENQA expectations and GDPR must map the propriety standards onto their own regulatory reality rather than adopt them verbatim.
Meta-evaluation costs resources. Auditing your evaluation against 30 standards is itself an evaluation with feasibility constraints. A full external meta-evaluation is not proportionate for every institution every year; a lighter internal self-audit against the five attributes usually is.
Standards do not fix; they diagnose. Identifying a utility failure does not tell you how to make reporting more usable — that requires design work. The framework is a lens, not a solution.
Risk of box-ticking. Like any standards set, the Program Evaluation Standards can be reduced to a compliance checklist that satisfies the letter while missing the point. The utility and accountability standards are precisely the ones most easily gamed.
How Koji incorporates this
Koji for Education is, in effect, designed against the Program Evaluation Standards — the platform's architecture targets the attributes where traditional course-evaluation systems most reliably fail.
- Utility. The Standards' core demand is that evaluation leads to use. Koji's closing-the-loop action tracking records what an institution commits to and re-measures it next cycle, and automatic thematic analysis turns raw open text into decision-ready themes rather than un-actionable tables — directly targeting the utility failures (unused data, unmeaningful reports) that sink most systems.
- Propriety. Koji is designed to mitigate re-identification risk in small cohorts and supports anonymity and GDPR-aligned handling, addressing the fairness and rights-protection standards that expose institutions to real legal and duty-of-care risk.
- Accuracy. Koji's quality scoring flags low-effort or careless responses, its bias-aware reporting contextualises scores rather than presenting bare means, and its structured question types (scale, single_choice, multiple_choice, ranking, yes_no, open_ended) support sound instrument design — feeding the accuracy standards without letting them crowd out the rest.
- Feasibility. By replacing multiple overlapping surveys with AI-moderated conversational interviews that adapt to each respondent, Koji is designed to extract more signal from fewer, shorter interactions — easing the survey-fatigue and resource-burden problems the feasibility standards target.
- Accountability. Because Koji retains the study configuration, questions, and analysis trail, the evaluation documents itself — making a course-evaluation system that can be meta-evaluated and shown to accreditors, which is exactly what the accountability attribute requires.
Koji does not claim to satisfy every standard automatically — meeting them remains an institutional responsibility, and the platform is designed to support that work rather than replace the judgement it requires. Teams applying the same use-focused discipline to product and customer research run it on Koji's core platform at koji.so.
A lightweight self-audit
A full external meta-evaluation is rarely proportionate, but a short internal self-audit against the five attributes almost always is. For utility, ask: do results reach instructors while they can still act, and do students ever learn what changed? For feasibility, ask: are we surveying more than we can meaningfully use? For propriety, ask: could any respondent be re-identified in a small cohort, and is the same instrument doing double duty for formative development and high-stakes promotion? For accuracy, ask: do our reports communicate uncertainty, or present bare means as if they were precise? For evaluation accountability, ask: is the process itself documented well enough that an external reviewer could inspect and judge it? Five honest answers usually locate the real problem faster than another round of item-wording debate — and the exercise itself becomes accreditation evidence of a reflective, self-correcting quality culture.
Related resources
- Designing Course Evaluation for Use: The Utilization-Focused Approach
- Does Closing the Feedback Loop Actually Matter?
- Interpreting and Reporting Student Ratings Responsibly
- Can a Student Be Re-Identified From Their Course Feedback?
- Validity Is About the Use, Not the Instrument: Kane's Framework
- Do Student Evaluations Actually Improve Teaching?
References
- Yarbrough, D. B., Shulha, L. M., Hopson, R. K., & Caruthers, F. A. (2011). The program evaluation standards: A guide for evaluators and evaluation users (3rd ed.). Joint Committee on Standards for Educational Evaluation. Sage.
- Stufflebeam, D. L. (2001). The metaevaluation imperative. American Journal of Evaluation, 22(2), 183–209. https://doi.org/10.1177/109821400102200204
- Scriven, M. (1991). Evaluation thesaurus (4th ed.). Sage.
- Linse, A. R. (2017). Interpreting and using student ratings data: Guidance for faculty serving as administrators and on evaluation committees. Studies in Educational Evaluation, 54, 94–106. https://doi.org/10.1016/j.stueduc.2016.12.004
Related articles
Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees
The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.
Does Closing the Feedback Loop Actually Matter? The Evidence on Acting on Student Evaluations
Universities are good at collecting student feedback and bad at acting on it visibly. The research — Watson (2003), Leckey & Neill (2001), Shah et al. (2017) — shows that failing to close the loop drives the scepticism and declining response rates that quietly destroy your evaluation data.
Do Student Evaluations Actually Improve Teaching? The Feedback-Intervention Evidence
Student ratings handed back as a bare number barely change teaching (d ≈ 0.2). Paired with structured consultation, the same data produces moderate, durable improvement (d ≈ 0.6–0.7). What the intervention meta-analyses tell quality teams about closing the loop.
Designing Course Evaluation for Use: The Utilization-Focused Approach
The biggest failure of course evaluation is not bad data — it is data nobody acts on. Patton's Utilization-Focused Evaluation and the empirical research on evaluation use (Johnson et al. 2009) show how to design feedback for action from the start.