New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Sector trends9 min read

Who Evaluates Your Course-Evaluation System? The Case for Meta-Evaluation

Universities evaluate every course and every teacher, then never turn the same scrutiny on the evaluation system itself. Meta-evaluation — judging your feedback process against recognised standards for utility, feasibility, propriety and accuracy — is the missing quality loop.

Koji Education Team

Product ·

Bottom line: Your institution subjects every module and every lecturer to systematic evaluation — and almost never applies the same discipline to the evaluation system itself. That system is a programme, it consumes real resources, and it can do real harm when it is wrong. Meta-evaluation — the formal evaluation of an evaluation — is the missing quality loop. There is an internationally recognised yardstick for it, and running your student-feedback process against that yardstick once a year is one of the highest-leverage things a quality-assurance office can do.

The blind spot at the centre of quality assurance

Higher education has industrialised course evaluation. Every semester, tens of thousands of surveys go out, dashboards fill with means, and committees make decisions — about curriculum, about probation, sometimes about promotion — on the strength of the numbers. The one thing that is almost never evaluated is the evaluation apparatus that produces those numbers.

This is a striking omission for a sector that prizes reflexivity. We would never accept a research instrument whose own validity had never been examined. Yet the student-evaluation-of-teaching (SET) system — arguably the most consequential measurement instrument most universities operate — is typically inherited, tweaked ad hoc, and left unexamined for years. Nobody owns the question "is our evaluation process itself any good, judged against a defensible standard?"

The discipline of programme evaluation has a name for the answer: meta-evaluation, the evaluation of an evaluation. Michael Scriven coined the term, and Daniel Stufflebeam developed it into a professional practice: applying the same rigour to an evaluation that the evaluation applies to its object. A course-evaluation system is exactly the kind of standing programme meta-evaluation was built for.

There is already a standard — you do not have to invent one

The reason meta-evaluation is practical rather than academic is that the yardstick already exists and is freely available. The Joint Committee on Standards for Educational Evaluation (JCSEE) Program Evaluation Standards (3rd edition, Yarbrough, Shulha, Hopson & Caruthers, 2011) set out 30 standards grouped into five categories:

  • Utility (8 standards) — will the evaluation serve the information needs of its users? Are results actually used?
  • Feasibility (4 standards) — is the process realistic, prudent, and cost-effective in context?
  • Propriety (7 standards) — is it legal, ethical, and fair to everyone involved, including the staff being evaluated?
  • Accuracy (8 standards) — does it produce technically sound, valid, reliable information?
  • Evaluation Accountability (3 standards) — added in the 2011 edition — is the evaluation itself documented and open to review?

You do not have to adopt all 30 verbatim. The value is in the audit posture: taking your SET system and asking, category by category, where it holds up and where it fails. Most systems fail in predictable places.

What a meta-evaluation typically exposes

Run an honest audit and the same weaknesses recur across institutions.

Utility failures are the most common — and the most damning. The JCSEE utility standards ask whether findings are actually used. Yet the "action gap" is the sector's open secret: feedback is collected, aggregated, and then not visibly acted upon. If a system scores well on accuracy but nobody closes the loop, it fails the category that matters most to the students who spent their time responding. A meta-evaluation forces this into the open by asking for evidence of use, not just evidence of collection.

Propriety failures hide in the dual-purpose problem. The propriety standards demand fairness to those evaluated. Using the same end-of-term instrument for formative improvement and for high-stakes personnel decisions — tenure, promotion, contract renewal for precarious staff — is a propriety problem the accuracy numbers will never reveal. It can be technically accurate and still improper.

Accuracy failures are the ones institutions think they have covered but usually have not. Reporting a raw mean for a class of 12 without an uncertainty interval; comparing departments on incommensurable scales; ignoring the well-documented biases in the underlying ratings — each is an accuracy-standard failure. The evidence that SET scores carry systematic bias related to instructor gender and other characteristics is strong enough that ignoring it is itself an accuracy problem.

Feasibility failures show up as survey fatigue. A system that technically works but exhausts students into non-response has traded accuracy for coverage without anyone deciding to.

The point of the framework is that these are not scattered gripes. They are named, categorised deficiencies against a published professional standard — which makes them actionable and reportable to a governing body in a way that "some faculty are unhappy" never is.

The European regulatory tailwind

For European institutions there is an added reason this is timely. The Standards and Guidelines for Quality Assurance in the European Higher Education Area (ESG), maintained under ENQA, already push institutions toward closing feedback loops and demonstrating that quality processes are themselves subject to review. Meta-evaluation is the concrete mechanism that satisfies that expectation: it is how you show an external quality-assurance agency, or a national body such as NVAO, that your feedback system is not just running but is itself held to a standard and periodically reviewed. Meta-evaluation turns a vague ESG aspiration ("continuous improvement of quality assurance") into a documented annual practice.

But isn't this just bureaucracy evaluating bureaucracy?

The strongest objection is regress: if you meta-evaluate the evaluation, must you then meta-meta-evaluate the meta-evaluation, and so on forever? And doesn't adding an audit layer just create more forms for exhausted staff?

Three responses. First, the regress terminates in practice because the JCSEE standards are self-applying — the Evaluation Accountability category explicitly asks the meta-evaluation to document itself, so the loop closes at one level, not infinitely. Second, meta-evaluation is lightweight relative to what it governs: a system that drives personnel decisions and consumes thousands of student-hours per year deserves at least an annual half-day audit against a checklist. The cost asymmetry is overwhelming. Third — and this is the real answer — the alternative to structured meta-evaluation is not no judgement of the system. It is ad hoc, political judgement: the loudest complaint in a committee meeting drives the next tweak. A standards-based audit replaces politics with evidence. That is less bureaucracy, not more, because it settles arguments that otherwise recur every year.

A fair critic will also note that the JCSEE standards are American in origin. True — but they are explicitly principles, not prescriptions, and they map cleanly onto the ESG's own emphasis on utility, fairness, and soundness. The categories travel.

Where Koji fits: a system built to pass its own audit

Meta-evaluation is method-neutral — you can and should audit any instrument, including a paper form. But some systems are easier to defend against the standards than others, and this is where the design of the underlying tool matters.

  • Utility. Koji for Education is built around closing the loop: automatic thematic analysis of open-text feedback and action-tracking turn findings into documented follow-up, which is exactly the evidence the utility standards demand.
  • Accuracy. Instead of a single number vulnerable to halo and central-tendency effects, Koji uses AI-moderated conversational interviews with six structured question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) and quality scoring, and it reports at programme and institution level with appropriate context — reducing the accuracy failures a raw-mean dashboard bakes in. The AI moderation is standardised, removing human-moderator inconsistency.
  • Propriety. Formative, mid-cycle collection supports separating improvement-focused feedback from summative judgement, easing the dual-purpose propriety problem. Data handling is GDPR/AVG-compliant and EU-appropriate.
  • Accountability. Because the process is standardised and its analysis is transparent and reproducible, the system can document itself — satisfying the newest JCSEE category.

The same conversational engine powers the main Koji platform for organisations running general user and customer research, where the case for auditing your own feedback process is no different.

Koji does not exempt you from meta-evaluation — it makes passing one substantially easier.


Audit the instrument that audits everything else. Explore Koji for Education and see how a standardised, loop-closing evaluation system stands up to the JCSEE Program Evaluation Standards.