New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
best-practices11 min read

Should Student Evaluations Decide Tenure? The Ryerson Arbitration and the Limits of High-Stakes SET

A landmark 2018 Canadian arbitration ruled that student evaluations of teaching should not be used to measure teaching effectiveness for promotion and tenure. This guide explains the decision, the evidence behind it, and what it means for governing the high-stakes use of course-evaluation data.

Koji Education Team

Product

Answer first: A landmark 2018 arbitration, Ryerson University v. Ryerson Faculty Association, ruled that student evaluations of teaching (SETs) should not be used to measure teaching effectiveness for promotion and tenure, accepting expert evidence that the scores are biased and unreliable for that purpose. The defensible position is not "abolish course evaluations" but "do not let a single, biased, low-reliability number drive irreversible personnel decisions." Use SETs formatively and as one triangulated input among several, with explicit safeguards against bias and false precision.

A ruling that changed the conversation

In 2018, arbitrator William Kaplan released a decision in a long-running dispute between Ryerson University (now Toronto Metropolitan University) and its faculty association. The grievance, which had simmered since the early 2000s, concerned whether the university could rely on student evaluation scores to assess teaching effectiveness in tenure and promotion. Kaplan's ruling was striking in its bluntness: he ordered that the parties amend their collective agreement so that SET results are not used to measure teaching effectiveness for promotion or tenure, that the institution stop relying on averages of student evaluations, that a numerical rating system be reconsidered, and that those sitting on personnel committees be educated about the systemic biases in student evaluations.

What makes the decision important for quality-assurance professionals is not the legal mechanics — it is an Ontario labour arbitration, not binding worldwide — but the evidentiary record. The arbitrator accepted, as essentially uncontested expert testimony, that SET scores are distorted by a long list of factors unrelated to teaching quality.

What the research says

The expert evidence in the Ryerson case rests on a now-substantial methodological literature, three threads of which are decisive.

SETs carry measurable bias. Boring, Ottoboni & Stark (2016), analysing a five-year natural experiment at a French university (≈23,000 evaluations) plus a US randomised dataset, concluded that SETs are more sensitive to students' gender bias and grade expectations than to teaching effectiveness, and that the bias can be large enough to make a more effective instructor receive a lower score than a less effective one. This is the heart of the validity problem: if a score moves with the instructor's gender or the grades students expect, it is not a clean measure of teaching.

SETs are too unreliable to rank people finely. Esarey & Valdes (2020) showed by simulation that even under fairly generous assumptions about how well SETs correlate with true teaching quality, using them to rank or threshold instructors produces frequent misclassification — many instructors are placed in the wrong tier purely by chance. That work underpins the related finding, covered elsewhere on this site, that you cannot fairly rank instructors by their evaluation scores and that even fair evaluations misclassify.

Statisticians warn against the way scores are reported. Stark & Freishtat (2014), in "An Evaluation of Course Evaluations," argued that averaging an ordinal scale, comparing means to one or two decimal places, and benchmarking against department norms all create false precision — points expanded in our summary of the Berkeley statisticians' conclusions. Their recommendation was to de-emphasise averages, report distributions, and never treat small numerical differences as meaningful.

Taken together, this literature supports a narrow but firm conclusion: SETs are weak instruments for high-stakes, summative, comparative judgements about individuals, even though they remain useful for formative feedback and programme-level quality signals. The Ryerson decision is best read as institutionalising that distinction.

Why it matters for course evaluation in practice

Separate formative from summative use. The single most important governance move is to stop using one instrument for two incompatible jobs. Formative use — helping an instructor improve — tolerates noise and bias far better than summative use deciding tenure. Evidence that feedback can improve teaching is about the formative channel, not a licence for high-stakes ranking.

Triangulate; never rely on a single source. Sound personnel evaluation of teaching draws on multiple methods — peer observation, teaching portfolios, self-evaluation, and student feedback — because each has different blind spots. Relying on SET alone is a textbook case of common-method dependence.

Ban the league table of averages. Following Ryerson and Stark–Freishtat, report distributions and response context, not a ranked list of decimals, and train committees to read them. A 4.2 and a 4.4 are not different.

Educate the readers, not just the respondents. The arbitrator's remedy of educating evaluators about bias is as important as anything done to the instrument. Decision-makers who know about gender, accent, and grading-leniency effects interpret scores more cautiously.

Designing a defensible teaching-evaluation policy

Institutions that take the Ryerson record seriously without discarding student voice tend to converge on a small set of policy commitments. They write into procedure that student feedback is one source among several and never the sole or decisive basis for a tenure or promotion judgement. They replace ranked decimal averages in personnel files with response distributions, response rates, and contextual flags, so a committee reads a pattern rather than a spurious point estimate. They require that anyone interpreting the data for a personnel decision has been briefed on the documented biases — gender, accent, discipline, grading leniency, class size — so caution is built into the reading rather than left to chance. They reserve the richest, most diagnostic feedback for formative use by the instructor, where noise and bias do far less damage. And they audit their own evaluation process periodically for adverse impact, treating the instrument as something to be validated and governed rather than trusted by default. None of this silences students; it disciplines how their voice is used.

Limitations and honest caveats

Several caveats keep this from becoming an anti-evaluation polemic. The Ryerson ruling is a single arbitration in one jurisdiction; it sets a persuasive precedent and reflects strong evidence, but it is not a universal legal standard, and other institutions and courts have reached more permissive conclusions. The bias literature, while robust, is not unanimous in magnitude: some studies find smaller or context-dependent gender effects, and a few find none in particular settings, so "biased and unreliable at worst" should not be flattened into "always severely biased." Critics of the abolitionist reading note that no measure of teaching is unbiased — peer observation and portfolios have their own halo and gaming problems — so the goal is a better-balanced portfolio, not the elimination of student voice. Finally, removing students from summative evaluation entirely risks silencing a legitimate and informative stakeholder; the defensible reform is disciplined, triangulated, bias-aware use, not abolition. Honest practice acknowledges both that SETs are too weak to decide careers alone and that they carry real, irreplaceable information about the student experience.

How Koji incorporates this

Koji for Education is designed for exactly the formative-first, triangulated, bias-aware model the Ryerson record points toward — and deliberately not for producing a single decimal to rank a career.

  • Built for formative depth, not a summative number. Koji's AI-moderated conversational interviews probe why behind a rating, surfacing actionable, improvement-oriented feedback rather than a lone scalar. That is the use case the evidence actually supports.
  • Triangulation by design. Koji captures structured items, rich open-text, and thematic patterns across cohorts, and is built to sit alongside peer review and self-evaluation rather than to substitute for them — reducing reliance on a single biased source.
  • Bias-aware, distribution-first reporting. Following Stark & Freishtat, Koji's reporting favours full response distributions, response-context flags, and theme-level analysis over headline averages and decimal-place league tables, and is designed to flag thin samples whose means cannot bear weight.
  • Surfacing rather than laundering bias. Because thematic analysis can highlight patterns of gendered language or abusive comments, Koji helps evaluators see potential bias rather than burying it inside an average — supporting the arbitrator's remedy of educating decision-makers.

These are designed-to-mitigate mechanisms; Koji does not claim to make SET data fit for unsupported high-stakes ranking, because the evidence says no instrument can. Its core research platform at koji.so applies the same triangulated, qualitative-plus-quantitative philosophy to product and customer research, where over-reliance on a single satisfaction metric carries analogous risks.

Related Resources

References

Related articles

analysis-reporting

Can You Fairly Rank Instructors by Their Course-Evaluation Scores? What Esarey & Valdes (2020) Proved

Even if course evaluations were unbiased, reliable, and valid, ranking instructors by their scores would still misclassify many good teachers. A walk through the Esarey & Valdes (2020) simulation and what it means for how you report and use evaluation data.

analysis-reporting

Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees

The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.

analysis-reporting

An Evaluation of Course Evaluations: What the Berkeley Statisticians Concluded About Reporting Student Ratings

Two Berkeley statisticians argued that averaging student-rating scores to rank instructors for promotion and tenure should be abandoned for both statistical and substantive reasons. What Stark & Freishtat (2014) actually recommend — and how to report ratings responsibly instead.

analysis-reporting

Even Fair Evaluations Can Rank Unfairly: Why Course-Evaluation Scores Misclassify Instructors

A simulation-based finding that even unbiased, reliable, and valid course-evaluation scores produce high misclassification rates when used to rank or compare instructors — and what it means for tenure, promotion, and merit decisions.