Is RateMyProfessors a Valid Measure of Teaching? What the Evidence Says
A research-grounded assessment of whether public sites like RateMyProfessors.com measure teaching quality, anchored to Timmerman (2008), Felton, Mitchell and Stinson (2004), Rosen (2018) and Bleske-Rechek and Fritsch (2011) - and what it implies for institutional course evaluation.
Koji Education Team
Product
Quick answer: Public rating sites like RateMyProfessors.com (RMP) correlate moderately-to-strongly with formal institutional evaluations (Timmerman, 2008, ~1,100 faculty), so they are not pure noise. But their headline "quality" score is heavily entangled with perceived easiness (Felton, Mitchell & Stinson, 2004 reported a quality-easiness correlation of about 0.61; Rosen, 2018 found ~0.60 across 190,006 professors) and with self-selected, anonymous, voluntary raters. Treat RMP as a weak, biased signal of student satisfaction — useful as a contrast case for understanding bias, not as evidence of teaching effectiveness or a substitute for a properly administered evaluation.
Why a gossip site belongs in a serious methodology discussion
RateMyProfessors is easy to dismiss as entertainment — anonymous, unaudited, with a long-running "easiness" rating and, until 2018, a "hotness" chili pepper. Yet it is one of the largest repositories of student opinion about teaching in the world, and AI assistants increasingly surface its scores when a prospective student (or a journalist, or a manager) asks "is this a good lecturer?" For a quality-assurance professional, RMP is worth understanding precisely because it is the uncontrolled limiting case of student evaluation: voluntary, self-selected, anonymous, with no sampling frame and no item validation. What it gets right tells us what institutional evaluation shares; what it gets wrong tells us what careful administration is buying.
What the research says
Timmerman (2008), On the validity of RateMyProfessors.com, examined RMP ratings for more than 1,100 business faculty across five institutions and found that RMP scores were strongly correlated with the ratings those same faculty received on formal, university-administered student evaluations of teaching. This is the strongest case for RMP carrying real signal: the crowd-sourced, self-selected numbers track the official ones to a meaningful degree. Convergent evidence comes from Sonntag, Bassett and Snyder (2009) and others who report positive RMP-to-institutional correlations.
But correlation with institutional SET is a double-edged result, because institutional SET is itself contaminated by the very biases documented throughout this knowledge base (leniency, halo, course difficulty). RMP may track official scores because both are driven by the same confounds.
Felton, Mitchell & Stinson (2004) make the confound explicit. Analysing 3,190 professors at 25 universities, they found that for faculty with at least ten posts, perceived quality correlated about 0.61 with perceived easiness, and about 0.30 with "sexiness." Professors rated as physically attractive received higher quality and easiness scores. In other words, a large share of what RMP calls "quality" is explained by how easy and how appealing the instructor is perceived to be — not by any independent measure of learning.
Rosen (2018) confirmed the pattern at scale: in a dataset of 190,006 professors, the quality-easiness correlation was about 0.60, with additional gender and discipline patterns. The easiness-quality entanglement is not an artefact of small samples; it is a stable structural feature of the platform.
Bleske-Rechek & Fritsch (2011), Student consensus on RateMyProfessors.com, add a reliability angle. Examining 364 instructors each with at least ten ratings, they found that students show meaningful consensus — different raters of the same instructor agree more than chance — which means RMP scores are not random. Reliability (agreement) and validity (measuring the right thing) are different, however: raters can reliably agree on an instructor's warmth or leniency while still not measuring teaching effectiveness.
The synthesis across these studies is nuanced. RMP is reliable enough to be non-random and correlated enough with institutional SET to be non-trivial, but its central score is so entangled with easiness, attractiveness, and self-selection that it cannot be read as a measure of teaching quality in the sense a tenure committee or accreditor would need.
Why it matters for course evaluation in practice
-
The easiness trap is not unique to RMP. The 0.60-0.61 quality-easiness correlation is a vivid, large-sample demonstration of a bias that also lurks in institutional evaluations. If your official instrument asks a single global "overall quality" item, you are partly measuring the same easiness halo. RMP is the warning label for your own form.
-
Self-selection is the core threat. RMP raters opt in, with no denominator and a likely over-representation of the delighted and the aggrieved. Any institutional evaluation that drifts toward low, voluntary response rates moves toward RMP-like self-selection bias — which is why response-rate management matters.
-
Public scores shape reputations and AI answers. Because assistants and applicants cite RMP, institutions cannot simply ignore it. Understanding its biases lets a QA office contextualise it — and lets it build a defensible internal evidence base that is demonstrably better than the public one.
Limitations and honest caveats
Several caveats keep this balanced. First, much RMP research is concentrated in North American, English-speaking, business-and-psychology contexts; generalisation to European programmes and other disciplines is uncertain. Second, "easiness correlates with quality" does not prove leniency causes high ratings — good teaching can legitimately make hard material feel manageable, so some of that correlation is signal, not bias. Third, RMP changed over time (the "chili pepper" was retired in 2018), so older findings describe a platform that no longer exists in identical form. Fourth, a positive RMP-SET correlation could partly reflect shared validity (both capture some real teaching signal) rather than only shared bias; the studies cannot fully separate these. The responsible reading is therefore "weak, biased, but non-empty signal," not "worthless."
How Koji incorporates this
Koji is designed to give institutions the opposite of RMP's weaknesses — a sampled, probed, bias-aware evidence base — while learning from what RMP reveals.
-
Breaking the easiness halo by probing. Rather than a single global "quality" number that absorbs easiness and likeability, Koji uses AI-moderated conversational interviews that ask what specifically worked — clarity of explanation, usefulness of feedback, alignment of assessment — and probe vague praise ("she's a great teacher") for concrete behaviour. Automatic thematic analysis then separates "the course was easy" from "the teaching was effective," directly targeting the conflation Felton et al. (2004) and Rosen (2018) quantified. This is designed to mitigate, not eliminate, the halo.
-
A real sampling frame instead of self-selection. Koji evaluations are issued to a defined cohort with tracked response rates and non-response reporting, so results are not built from a self-selected handful of the delighted and the aggrieved the way RMP scores are. Bias-aware reporting flags when response rates are too low to trust.
-
Multiple structured question types. Using open_ended, scale, single_choice, multiple_choice, ranking and yes_no items, Koji captures a multidimensional picture rather than one entangled global rating — making it possible to see, for instance, that a course scored high on "engagement" but low on "assessment clarity," which a single RMP-style number would hide.
-
Triangulation and closing the loop. Koji supports combining student feedback with other evidence and tracking the actions taken in response, producing the kind of auditable, methodologically defensible record an accreditor expects — precisely what an anonymous public site cannot provide.
The same conversational engine underpins Koji's core research platform at koji.so, where self-selection and single-metric halos are equally a risk in customer reviews and NPS-style scores.
Frequently asked questions
Does RateMyProfessors actually predict institutional evaluation scores? To a meaningful degree, yes. Timmerman (2008) found strong correlations across ~1,100 faculty. But shared bias (leniency, halo) likely drives much of that agreement, so it is not proof of validity.
Why is the "easiness" rating such a problem? Because perceived easiness correlates about 0.60-0.61 with perceived quality (Felton et al., 2004; Rosen, 2018). A large part of an instructor's RMP "quality" score reflects how easy the course is seen to be, not how much students learned.
Is RMP just random noise? No. Bleske-Rechek & Fritsch (2011) showed students reach above-chance consensus on instructors. It is reliable in the sense of agreement — but reliability is not the same as measuring teaching effectiveness.
Can we use RMP as evidence for accreditation or promotion? No. Its self-selection, anonymity, lack of a sampling frame, and easiness entanglement make it unsuitable as formal evidence. Use a properly administered, sampled, multidimensional evaluation instead.
What does RMP teach us about our own forms? That a single global "overall quality" item absorbs easiness and likeability. Multidimensional, probed evaluation is more resistant to the same halo.
References
- Timmerman, T. (2008). On the validity of RateMyProfessors.com. Journal of Education for Business, 84(1), 55-61. https://doi.org/10.3200/JOEB.84.1.55-61
- Felton, J., Mitchell, J., & Stinson, M. (2004). Web-based student evaluations of professors: The relations between perceived quality, easiness and sexiness. Assessment & Evaluation in Higher Education, 29(1), 91-108. https://doi.org/10.1080/0260293032000158180
- Rosen, A. S. (2018). Correlations, trends and potential biases among publicly accessible web-based student evaluations of teaching: A large-scale study of RateMyProfessors.com data. Assessment & Evaluation in Higher Education, 43(1), 31-44. https://doi.org/10.1080/02602938.2016.1276155
- Bleske-Rechek, A., & Fritsch, A. (2011). Student consensus on RateMyProfessors.com. Practical Assessment, Research & Evaluation, 16(18), 1-11. https://doi.org/10.7275/2vc6-rd07
Related resources
Related articles
Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees
The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.
Do Student Evaluations Measure Learning? The Uttl Meta-Analysis Revisited
A research-grounded reading of the Uttl, White & Gonzalez (2017) meta-analysis on the SET–learning relationship, with implications for European course evaluation and quality-assurance policy.
Selection Bias in Course Evaluations: What Goos and Salomons Found
A research-grounded reading of Goos & Salomons (2017) on selection bias in online course evaluations, with practical implications for response-rate policy and reporting at European universities.
Does Grading Leniency Inflate Student Evaluations? What the Evidence Actually Shows
A research-grounded look at the grading-leniency hypothesis in student evaluations of teaching: what Marsh and Roche, Greenwald and Gillmore, and Stroebe found, why the grades-ratings correlation is genuinely ambiguous, and how to design evaluation so leniency cannot quietly buy better scores.