Sixteen Myths About Student Ratings — and What 75 Years of Research Actually Says (Aleamoni 1999)
Aleamoni's 1999 review tested 16 common myths about student ratings against research from 1924 to 1998 and found most of them to be myths. This is a working guide to which beliefs about course evaluations the evidence supports, which it rejects, and where the honest answer is "it depends".
Koji Education Team
Product
The short answer
Lawrence Aleamoni's 1999 review, Student Rating Myths Versus Research Facts from 1924 to 1998, took 16 widely held beliefs about student evaluations of teaching (SET) — that students cannot make consistent judgements, that ratings are just popularity contests, that easy graders win, that the time of day or class size determines the score — and checked each against roughly 75 years of empirical research. His conclusion: most of these beliefs are, on the whole, myths. Student ratings are generally reliable and reasonably valid for some purposes, and many of the "obvious" biases that committees worry about are small or absent on average.
That headline needs two pieces of honesty around it. First, "myth, on average" does not mean "never a problem in your specific course" — several effects Aleamoni filed under myth (grading leniency, class size, discipline) are real but small, and a few biases discovered or sharpened after 1999 (gender, accent, the rating–learning disconnect) are better described as open problems than settled myths. Second, the review is now a generation old. Used carefully, though, Aleamoni's list is still one of the most useful sanity-checks a quality-assurance office can run against its own folklore.
What the research says
The structure of Aleamoni's review
Aleamoni, L. M. (1999), published in the Journal of Personnel Evaluation in Education (13(2), 153–166), organized the literature around 16 myth statements and, for each, summarized the weight of evidence. The myths span the full anxiety list of any teaching-evaluation committee:
- Students cannot make consistent or reliable judgements until years after graduation.
- Student ratings are merely popularity contests; warm, entertaining instructors win regardless of substance.
- Grading leniency buys high ratings — give easy A's and your scores rise.
- The time of day, the class size, whether the course is required or elective, and the discipline determine the rating more than teaching does.
- Students cannot be a meaningful source of information about teaching at all.
- Ratings are unreliable, invalid, and too unstable to be useful.
Against each, Aleamoni marshalled the accumulated reliability, validity and bias research. The recurring finding: ratings of the same instructor tend to be stable across raters and across time, they correlate with other indicators of effective teaching, and the feared contaminants explain far less variance than intuition assumes. Where a relationship exists at all (e.g., expected-grade and rating), it is typically modest and open to a non-biasing interpretation — students who learn more may both expect better grades and rate the teaching more highly, which is validity, not contamination.
Corroboration from the broader literature
Aleamoni did not stand alone. The IDEA Center's synthesis by Benton and Cashin (2012), Student Ratings of Teaching: A Summary of Research and Literature (IDEA Paper No. 50), reviewed the major reviews from the 1970s to 2010 and concluded that student ratings "tend to be statistically reliable, valid, and relatively free from bias … perhaps more so than any other data used for faculty evaluation." Marsh and Roche's work on the SEEQ, and Feldman's many syntheses, point the same way: well-constructed, multidimensional rating instruments behave like respectable measurement instruments, not noise.
Where the post-1999 evidence is harder on ratings
A research-aware reader must hold the contrary evidence in view, because the field did not stop in 1999:
- The rating–learning link is weak. The multisection literature once read as supportive (Cohen, 1981) has been re-estimated toward zero (Uttl, White & Gonzalez, 2017). "Ratings are reliable" and "ratings measure learning" are different claims; the first survives, the second largely does not.
- Demographic biases are real in places Aleamoni treated cautiously. Controlled and natural-experiment studies since 2014 (e.g., MacNell et al.; Boring) document gender effects, and other work documents accent and ethnicity effects, that are not comfortably dismissed as myths.
- Aleamoni's own filing of leniency, class size and discipline as "myths" is better read as "small real effects, not decisive biases" — a distinction that matters when scores are used at two-decimal precision for personnel decisions.
So the mature reading is: Aleamoni was right that the catastrophic folk-beliefs are overblown, and the post-1999 literature was right that "low bias on average" hides specific, patterned biases that can matter for individuals.
Why it matters for course evaluation in practice
Aleamoni's list is, in effect, a de-bunking script for the recurring arguments inside a teaching committee. Three practical uses:
1. Stop relitigating settled folklore. When a colleague insists "the 9 a.m. slot is why my scores are low" or "students just can't judge teaching," you can point to decades of evidence that these effects are small or absent. This protects good instructors from being dismissed on superstition and protects committees from chasing phantom confounds.
2. Reserve your scepticism for the effects that survive scrutiny. Spend your governance energy on the biases the modern literature does support — gender, accent, ethnicity, the rating–learning gap, and the misuse of tiny score differences — rather than on time-of-day or class-size adjustments that the evidence does not justify.
3. Build instruments that earn their reliability. Aleamoni's "ratings are reliable" verdict applies to well-designed instruments with enough items and respondents, not to a two-question afterthought. Reliability is a property you engineer, not a property ratings have automatically.
Limitations and honest caveats
- The review is a generation old. Some of the most consequential causal evidence on bias and on learning (Carrell & West 2010; MacNell 2015; Braga et al. 2014; the Uttl reanalyses) postdates it. Treat Aleamoni as a strong baseline, not the last word.
- "Myth on average" is not "myth for you." Aggregate near-zero bias can coexist with real bias for particular groups, courses, or score ranges. Averages hide individuals.
- Narrative review, not meta-analysis. Aleamoni summarized the weight of evidence rather than computing pooled effect sizes, so the strength of each verdict varies and is not always quantified.
- Validity is purpose-specific. "Ratings are valid" is shorthand. Valid for describing the student experience is well-supported; valid as a direct measure of learning is not. A myth-busting verdict should always carry the question "valid for what use?"
How Koji incorporates this
Koji is an AI-native course-evaluation platform, and Aleamoni's central lesson — that reliability and low bias are properties of a well-built instrument, not gifts — runs through its design.
- Engineering reliability, not assuming it. Aleamoni's "ratings are reliable" verdict holds for multi-item, adequately-sampled instruments. Koji supports structured, multidimensional question sets (
scale,single_choice,multiple_choice,ranking,yes_no,open_ended) and reports response counts and distributions so a programme can see whether a given course actually cleared the reliability bar — rather than acting on three responses to one question. - Targeting the biases that survived scrutiny. Because the live problems are gender, accent and ethnicity bias plus the misuse of small differences, Koji's reporting is designed to be bias-aware: it favours distributions and uncertainty over decimal-point league tables, the practice the misclassification literature warns against.
- Conversational probing to separate validity from contamination. When a rating co-varies with expected grade, is that bias or genuine learning? A bare number cannot say. Koji's AI-moderated conversational interviews follow up on a score in the student's own words and run automatic thematic analysis of the open text, helping a programme distinguish "easy grader" from "clear, well-organized teacher" — the exact ambiguity at the heart of the leniency myth.
- Formative, multi-cohort triangulation. Aleamoni's caution that no single indicator should decide a career is operationalized through mid-cycle/formative collection and cross-cohort comparison, so ratings sit alongside other evidence rather than standing alone.
As always, these are framed as mitigations. Koji does not claim to eliminate bias; it is designed to make it visible and to keep ratings tied to the inferences they support. Teams that also run product or customer research can apply the same AI-moderated engine through Koji's core platform at koji.so.
Related Resources
- Is Student Evaluation of Teaching Valid? What the State-of-the-Art Review Concludes
- Does Grading Leniency Inflate Student Evaluations? What the Evidence Actually Shows
- Gender Bias in Student Evaluations of Teaching: Evidence and Mitigation
- Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees
- Do Cookies, Treats, and Mood Bias Course Evaluations?
- Is RateMyProfessors a Valid Measure of Teaching? What the Evidence Says
References
- Aleamoni, L. M. (1999). Student rating myths versus research facts from 1924 to 1998. Journal of Personnel Evaluation in Education, 13(2), 153–166. https://doi.org/10.1023/A:1008168421283
- Benton, S. L., & Cashin, W. E. (2012). Student ratings of teaching: A summary of research and literature (IDEA Paper No. 50). The IDEA Center. https://www.ideaedu.org/idea_papers/student-ratings-of-teaching-a-summary-of-research-and-literature/
- Cohen, P. A. (1981). Student ratings of instruction and student achievement: A meta-analysis of multisection validity studies. Review of Educational Research, 51(3), 281–309. https://doi.org/10.3102/00346543051003281
- Uttl, B., White, C. A., & Gonzalez, D. W. (2017). Meta-analysis of faculty's teaching effectiveness: Student evaluation of teaching ratings and student learning are not related. Studies in Educational Evaluation, 54, 22–42. https://doi.org/10.1016/j.stueduc.2016.08.007
- Spooren, P., Brockx, B., & Mortelmans, D. (2013). On the validity of student evaluation of teaching: The state of the art. Review of Educational Research, 83(4), 598–642. https://doi.org/10.3102/0034654313496870
Related articles
Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees
The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.
Gender Bias in Student Evaluations of Teaching: Evidence and Mitigation
A research-grounded synthesis of gender bias in student evaluations of teaching (SET), anchored to MacNell, Driscoll & Hunt (2015) and corroborated by Boring (2017) and Mengel, Sauermann & Zölitz (2019), with practical mitigations for European quality-assurance teams.
Do Cookies, Treats, and Mood Bias Course Evaluations?
Two controlled studies show that giving students chocolate or cookies before an evaluation measurably raises teaching scores. What the affect heuristic means for the validity of course evaluations — and how to design around it.
Is Student Evaluation of Teaching Valid? What the State-of-the-Art Review Concludes
A close reading of Spooren, Brockx and Mortelmans'' 2013 Review of Educational Research synthesis of SET validity — what it actually concludes, the evidence behind it, and what it means for how universities use student ratings.