Do Instructors Actually Use Student Ratings? What Beran & Rokosh Found About Faculty Acceptance
Student evaluations only improve teaching if instructors act on them. Beran and Rokosh''s survey of 357 university instructors found faculty accept ratings as an accountability tool but consider them only marginally useful for actually improving their teaching — and resent public release. Here is the evidence and what it means for closing the loop.
Koji Education Team
Product
In brief
Course evaluations only change teaching if the instructors who receive them believe the data are useful and act on them — and the evidence says many do not. In a survey of 357 instructors at a major Canadian university, Beran and Rokosh (2009) found that faculty broadly accept student ratings as a legitimate accountability instrument and as useful to administrators for summative decisions, but rate them only marginally useful for improving their own instruction. Instructors were also notably unhappy about results being released publicly. The practical lesson is that a course-evaluation system''s impact is gated not by the psychometrics of the instrument but by the consequential stage: whether teachers find the feedback specific, credible, and actionable enough to use. A system that produces statistically immaculate numbers nobody acts on has failed at its only job that matters.
What the research says
Beran and Rokosh (2009), writing in Instructional Science, surveyed 357 instructors at a large Canadian research university where student evaluation is mandatory in every course each term. They asked two questions: what are instructors'' attitudes toward student ratings, and how useful do instructors find ratings for instructional improvement?
The findings were nuanced rather than simply hostile. Instructors agreed that student ratings are an acceptable way to assess institutional accountability and that ratings are useful to administrators making summative decisions — promotion, tenure, merit. In other words, faculty did not reject the legitimacy of being evaluated. But on the question that matters most for quality improvement — do ratings help me teach better? — instructors rated them only marginally valuable for enhancing instruction. And they reported being disappointed with the university''s policy of disseminating results publicly, a finding that anticipates the chilling and gaming effects public ranking can produce.
This pattern is corroborated across the faculty-perceptions literature. Studies of faculty attitudes (e.g., the 2019 Social Psychology of Education work linking attitudes toward evaluations to instructors'' self-image as teachers) show that acceptance is conditional and identity-laden: teachers who see ratings as a fair reflection of their effort engage with them, while those who see them as a popularity contest disengage. Reviews of the validity literature (Spooren, Brockx, & Mortelmans, 2013, Review of Educational Research) reach a compatible conclusion from the measurement side — that the use to which ratings are put, not the instrument alone, determines whether the exercise is defensible. And Linse (2017, Studies in Educational Evaluation) documents that the people interpreting ratings — committees and instructors alike — frequently lack the statistical literacy to read them responsibly, which further erodes instructors'' trust that the numbers mean what administrators claim.
The deeper mechanism connecting these results is the feedback-intervention literature. Decades of research (and meta-analyses such as Kluger & DeNisi, 1996, Psychological Bulletin) show that feedback only improves performance under specific conditions: when it is task-focused rather than ego-focused, specific rather than global, and paired with guidance on what to change. A bare numeric profile — "your clarity was 4.1, the faculty mean is 4.3" — fails every one of those conditions. It is global, ego-threatening, and silent on what to do differently. Seen through this lens, instructors'' verdict that ratings are "marginally useful for improvement" is not anti-intellectual resistance; it is an accurate read of feedback that is structurally unsuited to driving change.
Why it matters for course evaluation in practice
The entire justification for spending student and staff time on evaluations is that the data will improve teaching. Beran and Rokosh''s result punctures the assumption that this happens automatically. The implications for a QA office are concrete:
- The bottleneck is consequential, not psychometric. You can perfect your scale, your response rate, and your analysis and still produce zero teaching improvement if instructors do not engage with the output. Investment that stops at "we collected and reported the data" stops one step short of the goal.
- Format determines whether feedback is actionable. A profile of Likert means tells an instructor that something is wrong but not what or why. Open-text comments are richer but, in raw form, arrive as an unsorted, emotionally charged pile that triggers the negativity bias — one cruel comment outweighing twenty constructive ones — so instructors avoid them.
- Public release backfires. Beran and Rokosh''s disaffection finding aligns with the broader concern that publishing individual scores incentivises defensive teaching, grade leniency, and gaming rather than reflection (see our note on tenure and high-stakes use).
- Closing the loop is the live variable. The evidence that ratings can improve teaching is strongest when feedback is paired with consultation (see mid-semester feedback and consultation) and when students see that their input led to change (see closing the feedback loop). Faculty acceptance and student willingness (students are willing but doubt anyone listens) are two halves of the same trust problem.
Limitations and honest caveats
The honest reading of this evidence requires several caveats:
- Single-institution, single-country, and dated. Beran and Rokosh surveyed one Canadian university in the late 2000s. Attitudes vary by national QA culture, discipline, and institutional reward structure; the direction of the finding (accountability accepted, improvement value doubted) replicates widely, but the magnitudes do not transfer mechanically to a 2026 European context.
- Self-report about usefulness is not a measure of effect. Instructors saying ratings are "marginally useful" is an attitude, not a demonstration that ratings fail to improve teaching. Some instructors may improve without crediting the ratings; others may overstate their engagement. Attitudes and behaviour diverge.
- Acceptance is conditional and confounded. Faculty who teach large required quantitative courses face structurally lower ratings (see discipline differences) and may discount ratings for self-serving reasons unrelated to the instrument''s actual quality. Disentangling legitimate methodological scepticism from motivated reasoning is hard.
- The improvement question is under-powered by design. Most evaluation systems never give instructors the consultation, comparison over time, or specific behavioural detail that the feedback-intervention literature says is required for change — so the low "usefulness" ratings may reflect a poorly designed process, not an inherent limit of student feedback.
In short: the finding is robust as a warning that collection without actionable feedback wastes effort, but it is not proof that student ratings cannot improve teaching. It is proof that the current form of most ratings rarely does.
How Koji incorporates this
Koji for Education is built around the consequential stage that Beran and Rokosh identify as the weak point — turning ratings into feedback instructors will actually use:
- Actionable over global. Instead of returning a bare numeric profile, Koji''s AI-moderated conversational interviews probe why a student rated as they did and what specifically would have helped, producing task-focused, behaviourally specific feedback — the form the feedback-intervention literature says drives improvement — rather than an ego-threatening number.
- Automatic thematic synthesis of open text. Koji clusters open-ended responses into themes and surfaces representative quotes, so an instructor receives "twelve students asked for worked examples before the problem sets" rather than an unsorted comment pile dominated by the one hostile outlier. This directly counters the negativity bias that makes raw comments feel unusable.
- Mid-cycle and formative collection. Koji supports running short formative check-ins during the term, when there is still time to act and when feedback is paired with the instructor''s own follow-up — the consultation-style design with the strongest improvement evidence.
- Closing-the-loop tracking. Koji can record what action an instructor or programme took in response to feedback and report it back to students, addressing the public-trust problem from the opposite direction to public ranking: students see their input mattered without exposing instructors to a defensive league table.
- Designed to support, not replace, judgement. Koji frames its outputs as decision support for reflective practice and programme review, explicitly cautioning against using a single number for high-stakes ranking — consistent with instructors'' documented distrust of summative misuse.
The same conversational interview engine underpins Koji''s core research platform at [koji.so], which applies the identical "probe beyond the score, synthesise the themes, act on them" workflow to product and customer research.
Related Resources
- Do Student Evaluations Actually Improve Teaching? The Feedback-Intervention Evidence
- Does Closing the Feedback Loop Actually Matter?
- Mid-Semester Feedback and the Power of Consultation
- Should Student Evaluations Decide Tenure? The Ryerson Arbitration
- Students Are Willing to Evaluate — They Just Doubt Anyone Listens
- Interpreting and Reporting Student Ratings Responsibly
References
- Beran, T. N., & Rokosh, J. L. (2009). Instructors'' perspectives on the utility of student ratings of instruction. Instructional Science, 37(2), 171–184. https://doi.org/10.1007/s11251-007-9045-2
- Kluger, A. N., & DeNisi, A. (1996). The effects of feedback interventions on performance: A historical review, a meta-analysis, and a preliminary feedback intervention theory. Psychological Bulletin, 119(2), 254–284. https://doi.org/10.1037/0033-2909.119.2.254
- Spooren, P., Brockx, B., & Mortelmans, D. (2013). On the Validity of Student Evaluation of Teaching: The State of the Art. Review of Educational Research, 83(4), 598–642. https://doi.org/10.3102/0034654313496870
- Linse, A. R. (2017). Interpreting and using student ratings data: Guidance for faculty serving as administrators and on evaluation committees. Studies in Educational Evaluation, 54, 94–106. https://doi.org/10.1016/j.stueduc.2016.12.004
Related articles
Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees
The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.
Does Closing the Feedback Loop Actually Matter? The Evidence on Acting on Student Evaluations
Universities are good at collecting student feedback and bad at acting on it visibly. The research — Watson (2003), Leckey & Neill (2001), Shah et al. (2017) — shows that failing to close the loop drives the scepticism and declining response rates that quietly destroy your evaluation data.
Students Are Willing to Evaluate — They Just Doubt Anyone Listens: The Spencer & Schmelkin Evidence
Spencer and Schmelkin (2002) surveyed students about how they view course evaluations and found a clear pattern: students are generally willing to participate but have little confidence their feedback is actually used. That belief, not apathy, is the lever behind response rates and answer quality.
Do Student Evaluations Actually Improve Teaching? The Feedback-Intervention Evidence
Student ratings handed back as a bare number barely change teaching (d ≈ 0.2). Paired with structured consultation, the same data produces moderate, durable improvement (d ≈ 0.6–0.7). What the intervention meta-analyses tell quality teams about closing the loop.