New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Sector trends10 min read

Does Course Evaluation Quietly Shape What You Are Willing to Teach? The Chilling-Effect Problem

High-stakes course evaluation does not just measure teaching — it shapes it. When ratings drive promotion and reputation, rational academics drift toward what raises scores and away from productive difficulty. The evidence for a curricular chilling effect, the strongest objection to it, and how to keep student voice without the distortion.

Koji Education Team

Product ·

Bottom line up front: When a single comparable evaluation score is tied to promotion, contract renewal, and public league tables, it stops being a measurement and becomes an incentive. And academics respond to incentives. A growing body of evidence suggests that high-stakes student evaluation quietly nudges teaching toward what reliably lifts scores — leniency, lighter workloads, smoothness and charisma — and away from what can depress them in the short term: demanding content, productive struggle, unfamiliar pedagogy, intellectually uncomfortable material. This is a chilling effect on academic freedom that operates without anyone ever being told what to teach. It is distinct from the well-documented precarity of contingent faculty; it touches tenured professors too, because the pressure is built into the instrument, not just the contract.

How a measurement becomes a leash

The mechanism is not conspiracy; it is Goodhart''s law. As soon as a measure becomes a target, people optimise for the measure — see our piece on Goodhart''s and Campbell''s laws in course evaluation. If your renewal, your teaching prize, your standing in the department, and increasingly your visibility on a public ratings site all hinge on a number between 1 and 5, the rational move is to manage that number. The trouble is that the easiest ways to lift it often have nothing to do with — or actively work against — student learning.

The best-documented lever is grading. In a careful synthesis of the evidence, Wolfgang Stroebe (2016, Perspectives on Psychological Science) showed that students evaluate courses more positively the more leniently they are graded, and work less in such courses — meaning that good teaching evaluations can reward what he bluntly calls bad teaching, if "bad teaching" means courses where students do not learn much. In a follow-up analysis (Stroebe, 2020, Basic and Applied Social Psychology), he argued that student evaluations actively encourage poor teaching and contribute to grade inflation. This sits alongside the long-run pattern researchers have tracked for decades: average grades drifting upward while the time students invest in study drifts down. The American Association of University Professors reached a similarly stark verdict in its 2018 review, Student Evaluations of Teaching Are Not Valid.

The rigour penalty

The chilling effect bites hardest on exactly the practices that the learning-sciences evidence most strongly endorses. Active learning — retrieval practice, problem-solving, productive struggle — reliably improves outcomes, yet it often feels harder and less fluent to students than a polished lecture, and that feeling shows up as lower ratings. We covered this directly in the active-learning penalty: when students are asked to work, their feeling of learning can drop even as their actual learning rises. An instructor who reads their evaluations as a verdict learns a perverse lesson — that the way to better scores is to make the course feel easier, smoother, and less demanding.

Multiply that lesson across a career and a department and you get a slow, invisible narrowing. Difficult set texts get dropped. Genuinely challenging assessment gets softened. Risky, unfamiliar teaching methods get abandoned after one bruising round of feedback. No committee ever banned the hard version of the course; the incentive structure simply made teaching it irrational. That is what a chilling effect looks like — not censorship, but pre-emptive self-editing.

The strongest counterargument, taken seriously

Here is the objection that deserves a fair hearing: Isn''t this overblown? Plenty of students respect demanding teachers, and abolishing student feedback would be far worse than keeping it.

Both halves are right, and an honest analysis has to concede them. Stroebe himself notes that the fear of student retaliation may be exaggerated, and that many students who encounter rigorous, authoritative teaching are glad of it afterwards. Student voice is not the enemy: students experience teaching from the inside, and ignoring their perspective would be both unjust and unwise. The problem is not that we ask students what they think. The problem is what we ask and what we do with the answer.

A single summative satisfaction score, stripped of context, tied to high-stakes decisions, and published in rankings, is almost perfectly designed to produce a chilling effect. The same students, asked different questions in a lower-stakes, developmental setting, can tell you something genuinely useful — including whether a course''s difficulty was the productive kind that built understanding or the unproductive kind that just bred confusion. A number cannot distinguish those two. That distinction is the whole game.

Keeping the voice, losing the chill

If the chilling effect is driven by high-stakes, comparable, score-based evaluation, the antidote is to change those properties without silencing students:

  • Separate formative from summative. Mid-cycle, developmental feedback that an instructor owns is far less likely to distort behaviour than an end-of-term score that a promotion committee owns.
  • Stop reducing teaching to one comparable number. Much of the pressure comes from the comparability of a single Likert mean. Evidence that explains why students responded as they did resists being stacked into a league table.
  • Use evaluation as one source among several, never as the sole basis for tenure and promotion decisions — which is also the surest way to rebuild the faculty trust that high-stakes scoring has eroded.

Where Koji fits

Koji for Education is built to gather student voice in a way that informs teaching without weaponising it.

  • Conversational interviews that distinguish productive from unproductive difficulty. Koji''s AI moderator can follow up on "the course was hard" with "did the difficulty help you learn, or just frustrate you?" — surfacing the very distinction a satisfaction score erases, so that rigour is not punished blindly.
  • Formative, mid-cycle collection lets instructors hear and act on feedback while the course is live, in a developmental register, rather than receiving a high-stakes verdict after the fact.
  • Automatic thematic analysis separates "demanding but worthwhile" from "demanding and pointless," giving programme leaders signal that rewards good challenging teaching instead of penalising it.
  • Closing-the-loop action tracking turns feedback into visible change, shifting the culture from scoring teachers to improving courses.
  • Standardized, bias-aware AI moderation and GDPR/AVG-aligned data handling keep the process consistent and trustworthy across an institution.

The same conversational engine powers the general-purpose koji.so platform that teams use for user and customer research; the education edition adds the pedagogy-aware question design and quality-assurance reporting universities need.

The European quality-assurance angle

European quality assurance is, in principle, on the side of the solution. The Standards and Guidelines for Quality Assurance in the European Higher Education Area (the ESG) put student-centred learning at the heart of good practice (ESG 1.3) and frame student feedback as evidence for improving programmes, not as a stick for ranking individuals. Read that way, the ESG point away from the chilling effect: feedback is meant to feed continuous enhancement, with the loop closed back to students and staff.

The danger is in implementation. When institutions operationalise "student-centred" as a single satisfaction metric, benchmark it across staff, and feed it into workload and progression decisions, they recreate exactly the high-stakes scoreboard the ESG never asked for — and they import the chilling effect under a reassuring label. National agencies and review panels increasingly want to see that feedback leads to action, which is healthy; but "evidence of acting on feedback" should mean documented enhancement, not a ranked list of teachers. The distinction between using evaluation to improve a course and using it to judge a person is the line between a quality culture and a compliance culture — and it is the line that decides whether your evaluation system protects academic freedom or quietly erodes it.

The takeaway

Academic freedom is rarely lost in a single dramatic act. It erodes at the margin, one dropped reading and one softened assignment at a time, as conscientious teachers quietly optimise for a number that was never meant to govern the curriculum. The answer is not to stop listening to students. It is to stop turning their voice into a high-stakes scoreboard — and to start treating it as what it actually is: rich, contextual evidence about how a course is experienced and how it could be better.

Want student feedback that improves teaching without chilling it? See how Koji for Education works.