New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Sector trends9 min read

Responsible Metrics for Teaching: What the Leiden Manifesto and the Metric Tide Mean for Course Evaluation

Research assessment spent a decade building a governance answer to a misused number. Teaching evaluation has the same disease and never took the cure. The responsible-metrics canon transfers to course evaluation almost line for line.

Koji Education Team

Product ยท July 13, 2026

The short version

The research world spent a decade building a governance answer to a metric that was being misused: the journal impact factor and the h-index, deployed as if a single number could stand in for the quality of a scholar. That answer has a name, responsible metrics, and a canon: the Leiden Manifesto (2015), the UK's Metric Tide review (2015), the San Francisco Declaration on Research Assessment (DORA), and now the Coalition for Advancing Research Assessment (CoARA). Teaching evaluation has exactly the same disease and has never taken the cure. The course-evaluation average is the impact factor of the classroom: a crude aggregate, presented with false precision, used to make decisions it was never validated to support. The principles built for research metrics transfer to teaching almost line for line, and adopting them is the most credible reform available to a quality office.

The parallel is not loose

Consider what the responsible-metrics movement was reacting to. A single indicator, easy to compute and easy to rank, was being read as a proxy for a rich, multidimensional thing (research quality). It was compared across fields that had no business being compared. It was reported to more decimal places than its reliability justified. And it changed behaviour: people optimised for the metric rather than the underlying good.

Now read that paragraph as a description of student evaluation of teaching. A single mean, easy to compute and easy to rank, read as a proxy for teaching quality; compared across disciplines with different rating cultures (why you cannot compare a 4.1 in engineering to a 4.4 in history); reported to two decimals it does not deserve; and, as we have written, prone to Goodhart's law the moment it becomes a target. The diagnosis is identical. The difference is that research assessment produced a shared governance framework in response, and teaching evaluation mostly did not.

The Leiden Manifesto was published in Nature in 2015 by Diana Hicks, Paul Wouters and colleagues at Leiden University's Centre for Science and Technology Studies, offering ten principles so that "researchers can hold evaluators to account, and evaluators can hold their indicators to account." The Metric Tide review, the same year, distilled the same spirit into five dimensions of responsible metrics: robustness, humility, transparency, diversity and reflexivity. Neither was written with classrooms in mind. Both fit them.

The principles, translated

The most load-bearing principles need almost no translation to apply to course evaluation:

Quantitative evaluation should support, not supplant, qualitative expert judgement. For SET this is the whole ballgame. A rating average is an input to a human reading of the teaching, not a verdict that replaces it. Where evaluations feed tenure and promotion, this principle alone would rule out mechanical cut-offs.

Humility: avoid misplaced concreteness and false precision. A 3.9 is not meaningfully different from a 4.1 in a class of 20, yet institutions routinely act as if it were. Responsible metrics demands you report uncertainty, not just a point estimate, and resist the false flags a dashboard throws up by chance.

Account for variation by field. Rating cultures differ by discipline, class size, and course type exactly as citation cultures differ by field. Benchmarking that ignores this manufactures spurious winners and losers.

Let those evaluated verify the data and analysis. Faculty should be able to see and contest how their scores were produced and contextualised, a transparency most legacy systems never offered.

Scrutinise indicators regularly and recognise systemic effects. Reflexivity means auditing your own evaluation system, the case we have made for meta-evaluation: who evaluates the evaluation?

None of these is exotic. Together they describe a course-evaluation regime that would be almost unrecognisable to an institution currently ranking staff on raw means.

"But teaching isn't research, the analogy breaks down"

The strongest objection: research metrics and teaching evaluation differ in kind. Students are not peers; they rate their experience, not the scholarly quality of the work, and their ratings carry documented biases that citation counts do not. Doesn't that make the borrowing superficial?

Partly, and the differences matter. But they cut in favour of responsible metrics, not against it. If SET data is more contaminated by bias and construct-irrelevant variance than citation data, then the humility, context-sensitivity and support-not-supplant principles are more necessary, not less. The Leiden logic does not assume the indicator is good; it assumes the indicator is partial and builds governance around that partiality. That is precisely the situation teaching evaluation is in. It is also why a bare average serving two masters at once, improvement and accountability, is so fragile, a tension we unpack in the dual-purpose problem.

A second, fairer objection: principles without enforcement are decoration. An institution can adopt a responsible-metrics statement and carry on ranking staff by mean the next morning. True. Responsible metrics is a culture change, and the research-assessment movement's own mixed record shows how easily a declaration becomes a poster. The answer is not to abandon the principles but to wire them into the tooling and the workflow, so that the responsible reading is the default one, not an act of individual virtue.

What a responsible-metrics statement for teaching would say

The research world did not stop at principles; institutions turned them into commitments. When the Coalition for Advancing Research Assessment (CoARA) launched in 2022, signatories agreed to concrete pledges, chief among them that peer review and expert judgement lead, with quantitative indicators in a supporting role. A teaching-evaluation equivalent is overdue, and it is not hard to draft. A credible statement would commit an institution to a short, enforceable set of rules: never reduce a personnel decision to a rating threshold; always report the distribution and an uncertainty range alongside any mean; never compare raw scores across disciplines or class sizes without adjustment; give every member of staff the right to see and annotate the context of their own results; require qualitative evidence to sit beside any quantitative flag; and review the evaluation instrument itself on a fixed cycle.

What makes such a statement more than a poster is that each clause is checkable. "We do not use cut-offs" is either true of your promotion process or it is not. "We report uncertainty" is visible on the dashboard or it is absent. The failure mode of responsible metrics, in research and in teaching alike, is the gap between the declaration and the Tuesday-morning committee meeting, and the only remedy is to write the principles into the workflow and the tooling so that violating them takes deliberate effort. A statement an institution cannot be held to is worse than none, because it launders an unreformed practice in the language of reform.

Where Koji fits

This is the practical difference an AI-native platform can make: it can turn the responsible-metrics reading into the path of least resistance.

Koji for Education operationalises several of these principles directly. Its AI-moderated conversational interviews and automatic thematic analysis restore the qualitative substance the Leiden Manifesto insists must sit alongside any number, so a committee reads reasons, not just a mean. Its reporting can present distributions and uncertainty rather than a single false-precision average, honouring the humility principle we have argued for in how the mean hides the spread. Its standardised, bias-aware moderation narrows the construct-irrelevant variance that makes naive comparison so hazardous. And its closing-the-loop action tracking builds in reflexivity: the system records what was done with feedback, which is the beginning of scrutinising the indicator itself. The same interview engine powers the general Koji platform for teams doing broader stakeholder research under the same principles.

Koji does not turn a contested metric into an uncontested one. What it does is make it far harder to commit the specific sins, false precision, decontextualised ranking, number-without-narrative, that responsible metrics exists to prevent.

The limit worth stating

Responsible metrics is a governance stance, not a validity guarantee. Adopting the Leiden principles will not make a biased evaluation unbiased; it will make you honest about the bias and disciplined about how the data is used. That is a smaller claim than "we fixed course evaluation" and a much more defensible one. For a quality office trying to win back the trust of a sceptical faculty, honesty about the limits of the number may be worth more than any number.