New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Sector trends9 min read

Should Universities Publish Course Evaluation Results Publicly?

Public league tables of teaching scores feel like accountability. But publishing a noisy, bias-laden number turns a weak instrument into a high-stakes incentive to game — the opposite of quality.

Koji Education Team

Product ·

Bottom line up front: Publishing raw student-evaluation scores publicly — as searchable per-instructor or per-course league tables — is hard to defend on the evidence. Student ratings are too noisy and too biased to bear that weight, and making them public predictably amplifies both bias and gaming. Transparency is a genuine value in quality assurance, but the right object of transparency is what the institution did with the feedback, not the raw mean. There is a defensible middle path between secrecy and a public leaderboard.

The case for publishing — and why it is seductive

The argument for openness is intuitive. Students pay rising fees and deserve information to choose modules and programmes. Sunlight is a disinfectant: if teaching quality is measured, why hide it? Several systems already lean this way, from third-party platforms like RateMyProfessors to institutional dashboards that expose course scores. The instinct is democratic and, on its face, pro-student.

The problem is not the value of transparency. It is what happens when you attach public stakes to a measurement that cannot carry them.

The instrument is too weak to publish

Start with validity. The strongest evidence on whether student ratings track actual learning comes from multisection studies, and the most rigorous recent synthesis — Uttl, White and Gonzalez's 2017 meta-analysis, "Meta-analysis of faculty's teaching effectiveness: Student evaluation of teaching ratings and student learning are not related" — found that once prior ability is accounted for, the correlation between ratings and learning is essentially zero. Their reanalysis showed that earlier optimistic conclusions (notably Cohen's 1981 meta-analysis) were inflated by a small-sample bias that gave tiny studies the same weight as large ones.

Layer on the documented biases. Statisticians Philip Stark and Richard Freishtat, in Berkeley's widely read "An Evaluation of Course Evaluations", concluded that averages of student ratings have an "air of objectivity" but are "not a reliable measure of teaching effectiveness," partly because they correlate with characteristics that have nothing to do with teaching. Controlled experiments such as Boring, Ottoboni and Stark (2016) found gender bias in ratings even when teaching was held constant. An instrument with near-zero validity and demonstrable demographic bias is a poor candidate for a public scoreboard.

Publishing turns a measurement into a target

The deeper danger is incentive distortion — Goodhart's Law in an academic key: when a measure becomes a target, it ceases to be a good measure. Wolfgang Stroebe's analysis, "Student Evaluations of Teaching Encourage Poor Teaching and Contribute to Grade Inflation", lays out the mechanism: because lenient grading and lighter workloads reliably raise ratings, tying consequences to ratings pressures instructors toward exactly those behaviours. Make the scores public and you sharpen that pressure: the rational move for an anxious academic is to teach to the survey — easier content, higher grades, more entertainment — rather than to learning.

Public ratings also magnify self-selection. Third-party platforms illustrate the failure mode: voluntary online ratings draw disproportionately from students at the extremes, and easiness and perceived quality tend to move together, so the "scoreboard" rewards leniency rather than rigour. Publishing institutional scores invites the same dynamic, just with the university's logo on it.

But isn't hiding the data paternalistic?

Here is the strongest counterargument, stated fairly: refusing to publish looks like a guild protecting itself. Students generate the data; surely they are entitled to see it, and withholding it concentrates power with administrators who already use these scores behind closed doors in tenure and promotion decisions.

This objection has real force, and it should not be waved away. But it conflates two different things: transparency of the raw metric and transparency of the process. The honest response to "we use flawed scores in high-stakes decisions" is to stop doing that — as a growing scholarly consensus urges — not to broadcast the flawed scores more widely. A metric too unreliable to decide a promotion is also too unreliable to publish as consumer guidance. The genuinely student-serving form of openness is to publish what changed: the issues students raised, and the actions the department took in response. That is accountability students can actually use, and it is exactly what ESG-style "closing the loop" expects.

A defensible middle path

Between the locked filing cabinet and the public leaderboard, there is a more honest design — one that treats transparency as a property of the institution's conduct rather than of an individual's score. The goal is to let students see that their feedback has weight without converting a weak instrument into a tool that punishes the wrong people:

  • Publish actions, not averages. Share "you said, we did" summaries at programme level — the changes made in response to feedback — rather than per-instructor means.
  • Report with uncertainty. Where numbers are shared internally, show distributions and margins of error, never a single decimal, so a 4.1 is not mistaken for "better" than a 3.9.
  • Aggregate to protect against bias and small samples. Programme- and cohort-level reporting is both more reliable and less exposed to the demographic biases that distort individual scores.
  • Separate formative from summative. Feedback meant to help an instructor improve should not double as public marketing copy; mixing the two corrupts both.

The chilling effect on rigour — and on diverse faculty

Public scoreboards do not distribute their harms evenly. Because student ratings carry documented demographic bias, making them public amplifies that bias into a reputational signal. A 2024 study of UK academics, "Understanding the impact of biased student evaluations", documents how women, faculty of colour and international staff — who already tend to receive lower ratings for reasons unrelated to teaching — bear the brunt when those ratings acquire stakes. Publish the numbers and you publish the bias, attaching it to the names of exactly the staff a sector committed to inclusion says it wants to retain.

There is a second-order effect on the curriculum itself. When scores are visible and consequential, the rational response is risk aversion: drop the demanding text, soften the marking, trim the challenging assessment that a few students will punish in the survey. Over time a public-rating regime selects against rigour and in favour of palatability — not because any individual academic is cynical, but because the incentive gradient points that way for everyone at once. The students who lose most are future cohorts, who inherit a quietly easier degree.

None of this means institutions should be opaque. It means the object of transparency matters. Publishing a department's response to feedback — the issues raised, the changes made, the things explicitly decided against and why — gives students real, usable accountability without handing a biased instrument the power to shape hiring and teaching. Other sectors learned this the hard way: crude public metrics, from hospital league tables to school rankings, reliably produce gaming and demoralisation when the measure is weak and the stakes are high. Teaching evaluation, with its near-zero correlation to learning, is a textbook case of a metric that should not be turned into a target.

Where Koji fits

Moving from "publish the number" to "publish the response" requires evaluation data rich enough to act on — and that is where an AI-native approach helps. Koji for Education is built around AI-moderated conversational interviews that probe beyond a rating to why students felt as they did, with automatic thematic analysis turning thousands of open-text responses into the specific, recurring issues a department can actually address. Its closing-the-loop action tracking and programme- and institution-level reporting are designed for exactly the "you said, we did" transparency that serves students, while standardised, bias-aware AI moderation reduces the human inconsistency that makes raw scores so unstable. Koji does not claim to eliminate bias — no system can — but it shifts the centre of gravity from a publishable number toward documented, actionable change. The same conversational interview engine underpins general research on the main Koji platform, so institutions running broader stakeholder studies stay on one method.

Make accountability mean action, not a leaderboard. Explore Koji for Education to turn student feedback into changes worth publishing.