Should Student Evaluations Decide Tenure and Promotion? What the Evidence Says
The averaged student evaluation score is the most consequential number in many academic careers — and one of the least defensible for high-stakes use. We review the evidence on validity and bias, and explain what a fairer, multi-source approach looks like.
Koji for Education
Research & Editorial Team · June 6, 2026
Bottom line: Using the institutional average of student evaluation of teaching (SET) scores as the primary, comparative measure for tenure, promotion, and renewal decisions is not supported by the evidence. SET scores carry useful formative signal, but as a high-stakes, norm-referenced ranking tool they are confounded by factors unrelated to teaching quality and only weakly related — if at all — to how much students actually learn. The defensible position is not to abolish student feedback, but to stop treating a single decimal-point average as a verdict.
The number that decides careers
In most European and North American universities, a faculty member''s teaching is summarised, for personnel purposes, as an average: a 4.2 out of 5, a 3.8, an 82%. That number is then compared — to a departmental norm, to a threshold, to a colleague up for the same promotion. It is fast, quantitative, and feels objective. The problem is that the speed and the apparent objectivity are exactly what make it dangerous when the stakes are someone''s tenure case.
The methodological case against high-stakes use rests on two distinct findings that are often conflated but should be kept separate: bias (the score measures things other than teaching) and validity (the score doesn''t track learning).
Finding 1: SET averages are confounded by factors instructors don''t control
The most cited single study here is Boring, Ottoboni and Stark (2016), Student Evaluations of Teaching (Mostly) Do Not Measure Teaching Effectiveness, published in ScienceOpen Research. Analysing roughly 23,000 evaluations from a French university alongside a controlled US online experiment, the authors found SET scores correlated with the instructor''s perceived gender and with students'' grade expectations rather than with measured learning. In the online experiment, students rated an instructor they believed to be male more highly than the same instructor presenting as female — a difference in perception, not in teaching.
This is consistent with a broad bias literature. Our own deep-dives summarise the evidence on gender bias in student evaluations, and on how a single averaged Likert figure hides more than it reveals. The point for personnel committees is narrow but decisive: if a score moves with gender, accent, attractiveness, discipline, class size, and grade expectations, then comparing two instructors'' averages does not compare their teaching unless all of those are held equal — which they never are.
Finding 2: SET scores barely track student learning
The validity question is even more damaging for high-stakes use. Uttl, White and Gonzalez (2017), in a meta-analysis published in Studies in Educational Evaluation (vol. 54, pp. 22–42), re-analysed the classic multisection studies — the design where students are randomly distributed across sections of the same course and learning is measured by a common final exam. Once the analysis accounted for small-sample effects, they found that SET ratings explain at most about 1% of the variance in student learning, and in the larger, better-powered studies essentially none. Earlier reviews reporting moderate SET–learning correlations, they argue, were inflated by small samples and publication bias.
In plain terms: students do not reliably learn more from instructors they rate highly. A measure that is meant to certify teaching effectiveness for a career decision should, at minimum, correlate with the thing it certifies. This one largely does not.
This is not a fringe position
Scholarly and professional bodies have moved on this. The American Sociological Association''s Statement on Student Evaluations of Teaching (2019) warns against using SET as the primary measure of teaching effectiveness and recommends multiple sources of evidence. The American Association of University Professors has published the case that student evaluations, used this way, are not valid. In a widely noted 2018 Canadian arbitration, an arbitrator ruled that Ryerson University could not rely on SET averages to measure teaching effectiveness for tenure and promotion. The direction of travel among methodologists is clear and converging.
But doesn''t this just let bad teaching off the hook?
This is the strongest counterargument, and it deserves a direct answer. No — and the distinction matters. Rejecting the averaged-score-as-verdict model is not the same as rejecting student input. Students are uniquely positioned to report on things they genuinely observe: whether the course was organised, whether feedback arrived in time to be useful, whether they felt able to ask questions, whether assessment matched what was taught. That is real, decision-relevant evidence.
What students are not well positioned to do is provide a calibrated, bias-free, single-number ranking of instructors that can be compared across genders, disciplines, and class sizes for a promotion file. The error is not listening to students; it is laundering rich, qualitative, context-bound feedback into one decimal and then treating that decimal as a measurement instrument it was never validated to be.
A second objection: "If not SET averages, then what — peer observation is biased too, and learning outcomes are hard to measure." Correct. No single source is sufficient. That is precisely the argument for triangulation rather than for any one replacement metric — a theme we develop in our piece on using multiple sources of evidence to evaluate teaching.
What a defensible model looks like
The methodological consensus points toward a few concrete principles:
- Separate formative from summative. Use student feedback heavily to help instructors improve, and only cautiously — alongside other evidence — in summative personnel decisions.
- Stop ranking on tiny differences. A 4.1 versus a 4.3 average from two classes of thirty students is statistical noise, not a quality gap. Report distributions and uncertainty, not league tables.
- Read the comments, don''t just average the scales. The actionable signal in evaluations is overwhelmingly in the open text — what students describe, not the number they circle.
- Triangulate. Combine student voice with peer review, self-reflection, teaching portfolios, and, where possible, evidence of learning.
The European angle: ESG, ENQA and the audit trail
For European institutions there is an additional, often-overlooked dimension. The Standards and Guidelines for Quality Assurance in the European Higher Education Area (ESG 2015) and national bodies such as the NVAO in the Netherlands and Flanders expect programmes to act on student feedback and to evidence a closed loop — not merely to collect a number and file it. A high-stakes personnel process built on confounded averages sits awkwardly inside that framework: it uses student data for a purpose the data does not support, while frequently failing to demonstrate the formative, improvement-oriented use that quality reviewers actually want to see.
This creates a quiet opportunity. The same shift that makes evaluation fairer to individual academics — moving from a comparative average to triangulated, qualitative, improvement-focused evidence — also makes an institution''s quality case stronger at accreditation. Treating student feedback primarily as formative intelligence, documenting the actions it drives, and reserving summative judgement for triangulated evidence is both the methodologically defensible position and the one that aligns with European QA expectations. Committees that cling to the SET league table risk the worst of both worlds: legally and methodologically exposed on personnel decisions, and thin on the closing-the-loop evidence that ESG-aligned reviews increasingly demand.
Where Koji fits
Koji for Education is built around the parts of this that legacy SET tools handle poorly. Instead of reducing a course to a Likert average, Koji runs AI-moderated conversational interviews that probe why a student answers as they do — turning "3 out of 5 on clarity" into a specific, attributable account of what was unclear and when. Its automatic thematic analysis surfaces patterns across hundreds of open-text responses, so committees and teaching-and-learning centres can act on substance rather than on a mean. Standardised, bias-aware AI moderation asks every student comparable, neutral questions — removing the human-moderator inconsistency that adds yet another confound. And because Koji supports mid-cycle, formative collection and closing-the-loop action tracking, it is designed for the use SET actually has good evidence for — helping teaching improve — rather than the high-stakes ranking it does not.
Koji does not claim to eliminate bias; no instrument can. It is designed to reduce reliance on a single confounded number and to surface the qualitative evidence that a fair evaluation process needs. Many of the same teams also run general user and customer research on the main Koji platform, which shares the same AI interview engine.
If your institution is rethinking how teaching evidence feeds promotion and review, explore Koji for Education — and, just as importantly, read the primary sources above. The credibility of any evaluation reform depends on getting the methodology right first.