New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Sector trends8 min read

When the Reviewer Holds Your Contract: Course Evaluations and the Precarity of Contingent Faculty

Most teaching in European and North American universities is now done by staff on fixed-term or hourly contracts. When their renewal hinges on a satisfaction score, the evaluation stops measuring teaching and starts measuring who dares to challenge students. This is a structural problem, not a personal failing.

Koji Education Team

Product ·

Bottom line: A growing majority of university teaching is delivered by contingent staff—adjuncts, fixed-term lecturers, hourly tutors—whose contracts can turn on student-evaluation scores. When a satisfaction rating carries direct employment consequences for someone with no job security, it stops being a measure of teaching quality and becomes a measure of how much an instructor dares to challenge, fail, or push students. The predictable results are grade inflation, risk-averse pedagogy, and a chilling effect that falls hardest on those least able to absorb a single bad review. The remedy is not to abolish student feedback—it is to stop using thin, high-stakes ratings as a disciplinary instrument against the most precarious teachers, and to collect richer, developmental evidence instead.

Who actually teaches now

The faculty model that course evaluations were designed around—tenured professors with secure positions—is no longer the norm. In the United States, contingent appointments grew to roughly three-quarters of instructional positions; the AAUP's background data on contingent faculty documents how non-tenure-track teaching became the majority. Europe is no less precarious. In Germany, reporting compiled by sector unions indicates that the great majority of academic staff below professorial rank work on fixed-term contracts—one widely-cited figure puts it at over 90% of non-professorial academics under 45—and the UK's University and College Union precarity work has tracked a comparable casualisation of teaching. The OECD has flagged the precarity of academic careers as a systemic concern across member states.

This matters for evaluation because the stakes are not symmetric. A tenured professor with a weak evaluation has a conversation with a head of department. A fixed-term lecturer with a weak evaluation may not be rehired. Same number; radically different consequence.

What high stakes do to the data

When the cost of a poor rating is your livelihood, the rational response is to manage the rating. The mechanism is well-documented. Wolfgang Stroebe's analysis, Student Evaluations of Teaching Encourage Poor Teaching and Contribute to Grade Inflation (2020), lays out how the incentive structure rewards leniency: easier grading and lighter workloads buy higher satisfaction scores, and instructors in precarious positions have the strongest reason to respond. Qualitative work such as Pressure to Please: Adjunct Faculty Experiences with Grade Inflation reports adjuncts describing exactly this calculus—softening assessment to protect the evaluations that protect their contracts.

The consequences compound:

  • Grade inflation. When challenge is punished, challenge declines. (See our coverage of the grading-leniency mechanism.)
  • Risk-averse teaching. Contingent instructors are discouraged from demanding pedagogy—the kind that produces the active-learning "penalty" of lower short-term satisfaction but better learning.
  • Amplified bias. Documented biases in raw ratings—including the gender bias the evidence supports—do more damage when the rated person is disposable. A biased score against a tenured professor is an annoyance; against an adjunct it can be a termination.
  • Vulnerability to single complaints. With little formal review or mentoring, contingent staff can be let go over one or two negative comments, giving disproportionate weight to outliers in already-noisy data.

Why this is a quality-assurance problem, not just a fairness one

It would be easy to file this under labour ethics and move on. That would be a mistake, because the distortion contaminates the evidence base quality assurance depends on. If a measurable fraction of your teaching staff is grading to the rating rather than to the standard, then the evaluation scores, the grade distributions, and the inferences you draw about programme quality are all jointly corrupted. You are not measuring teaching; you are measuring an equilibrium between students who reward leniency and teachers who cannot afford not to supply it. A quality system that produces this incentive is undermining the very standards it exists to protect.

"But students deserve a voice—and some contingent teaching genuinely is weak"

This is the strongest objection, and it is correct on both counts. Students absolutely have a legitimate interest in the quality of their teaching, and the answer is not to shield contingent staff from all feedback—that would patronise students and protect genuine underperformance. Some contingent teaching is weak, and pretending otherwise helps no one.

But the objection conflates two different uses of feedback. There is formative feedback—developmental, low-stakes, intended to help an instructor improve—and summative feedback used to make employment decisions. The problem is not that contingent staff are evaluated; it is that a single thin satisfaction score is asked to do both jobs at once, and the high-stakes use poisons the developmental one. The defensible position is to keep listening to students intently while refusing to let a noisy, bias-prone, easily-gamed number be the deciding evidence for whether a precarious person keeps their job. Feedback for improvement should be rich and frequent; feedback for judgement should be triangulated with peer review and teaching artefacts and read with full awareness of its limits. (The same logic applies to tenure and promotion.)

The European policy backdrop

This is not only an institutional choice; it sits inside a regulatory context that increasingly recognises precarity as a quality risk. The OECD has explicitly examined how to reduce the precarity of academic research careers, and national systems are under pressure to reform fixed-term employment in higher education—debates over EU-level limits on successive fixed-term academic contracts have been live for years precisely because the status quo is seen as unsustainable. Meanwhile the European quality-assurance framework, the Standards and Guidelines for the European Higher Education Area (ESG), expects institutions to assure themselves of the competence and the conditions of their teaching staff. An evaluation regime that quietly pressures the most precarious teachers into grade inflation is not a neutral measurement instrument sitting outside that framework—it is a mechanism actively working against the staff-quality standards the framework asks institutions to uphold. Quality officers therefore have a defensible, standards-aligned reason to scrutinise how evaluation scores are used for contingent staff, not merely a compassionate one. Getting this right is part of doing quality assurance properly, not an exception to it.

How Koji helps decouple feedback from punishment

Koji cannot fix academic labour markets, and it does not claim to. What it can do is make it practical to separate developmental feedback from high-stakes judgement, so that listening to students stops being a threat to the people doing most of the teaching:

  • Formative, mid-cycle collection gives contingent instructors actionable feedback while a course is running, when they can still act on it—shifting evaluation from a verdict toward a development tool.
  • AI-moderated conversational interviews surface why something is or is not working, replacing a bare satisfaction number with specific, improvable detail that a renewal decision could never fairly rest on alone.
  • Automatic thematic analysis and quality scoring distinguish a genuine, recurring teaching issue from one or two outlier complaints—protecting precarious staff from being dismissed on statistical noise.
  • Standardised, bias-aware moderation applies the same well-formed process to every instructor, reducing the inconsistency that lets bias do extra damage to those without job security.
  • Programme- and institution-level reporting lets quality teams see patterns across a programme rather than pillorying individuals, supporting the triangulated, structural view that fairness and good measurement both require.

The same conversational engine underpins koji.so for organisations running internal feedback and 360-style research, where tying a single satisfaction metric to someone's job security produces the same defensive, low-trust behaviour it does in the lecture hall. Good feedback systems separate growth from judgement; bad ones collapse the two.

The takeaway

When the reviewer effectively holds the contract, a course evaluation is no longer a neutral measurement—it is a power relationship with a number attached. The majority of university teaching now happens under exactly those conditions, which means the distortion is not marginal; it is central to what your evaluation data actually represents. The evidence-led response is to keep students' voices loud and constant, but to stop asking a thin, gameable, bias-prone score to decide who teaches next term. Listen more. Punish less. Measure properly.

Want to give your contingent and permanent staff developmental feedback that is fair to act on? See Koji for Education.