Do Older Professors Get Lower Course Evaluations? Age and Seniority as a Confound
Stonebraker and Stone found a small but robust negative effect of instructor age on student ratings, emerging after the mid-forties. What the evidence does and does not show, and how to stop age from contaminating comparisons.
Koji Education Team
Product
In brief: There is modest, replicated evidence that student ratings decline gently with instructor age. Stonebraker and Stone (2015), analysing a large cross-section of instructors on RateMyProfessors, found a small but statistically robust negative age effect that does not begin until the mid-forties, then holds roughly flat — and that all but vanishes for instructors students judge physically attractive. The effect is real but small, is entangled with appearance and cohort factors, and is not evidence that older academics teach worse. For quality assurance the lesson is practical: comparing instructors of very different ages on a single evaluation scale, or tracking an individual across decades without allowing for it, risks reading an age artefact as a change in teaching.
What the research says
The anchor study is Stonebraker and Stone (2015), Too Old to Teach? The Effect of Age on College and University Professors (Research in Higher Education, 56(8), 793–812). Using a large sample of instructors drawn from RateMyProfessors.com across a broad range of institutions and disciplines, the authors modelled overall student ratings against instructor age with controls for discipline, institution type and other characteristics. They found that age has a small negative impact on student ratings that is robust across genders, disciplines and institution types. Three features of the effect are important:
- It is delayed. Ratings do not decline through the thirties and early forties. The downturn only appears from roughly the mid-forties onward.
- It plateaus. Having emerged, the effect does not keep worsening — ratings for professors in their late sixties or seventies are not dramatically below those in their fifties. There is no cliff at traditional retirement age.
- It is small and offsettable. The quantitative magnitude is modest, and it can be offset by other factors — most strikingly, when the authors restricted the sample to instructors students rated as physically attractive ("hot"), the age effect disappeared entirely.
That last result is the interpretive key. It strongly suggests the "age effect" is not a decline in teaching but is bound up with appearance, self-presentation and student expectations — the same family of surface cues that drive attractiveness, thin-slice and warmth effects elsewhere in the SET literature. Older instructors are, on average, judged less on the appearance dimensions that inflate ratings, and that shows up as a small age penalty.
Corroborating context: the broader literature on attractiveness and "hotness" (e.g., work linking RateMyProfessors "hotness" flags to higher quality ratings) and on thin-slice judgements (students form durable impressions of an instructor within seconds of exposure) makes an appearance-mediated age effect entirely plausible. It is also consistent with the general finding that global student ratings absorb a good deal of non-instructional signal. Against this, institutional studies using administrative data rather than a self-selected review site tend to find age effects that are weaker still or inconsistent, which is itself informative about how much the RateMyProfessors context contributes.
Why it matters for course evaluation in practice
Cross-sectional comparison. If a department ranks instructors against one another on a single overall score, a 58-year-old and a 34-year-old are not being compared on a level field. A small systematic age penalty, layered on top of the other well-documented confounds (class size, course level, discipline), can be enough to move an instructor across an arbitrary threshold in a league table — the misclassification problem in another guise.
Longitudinal misreading. An instructor tracked across a twenty-year career may show a gentle drift downward in scores that reflects the age artefact rather than any real change in their teaching. Read naively, a QA office might "detect" a decline and intervene where nothing has actually changed — a cousin of regression-to-the-mean errors.
Interaction with other biases. Because the age effect appears to run through appearance, it does not act alone: it can compound gender and attractiveness biases, and it may fall differently on women, for whom appearance-based judgement is already more pronounced in the evaluation literature. Treating age as a clean, isolated variable understates how it entangles with the rest.
Limitations and honest caveats
- Source matters. RateMyProfessors is a self-selected, voluntary sample skewed toward students with strong opinions and toward certain disciplines. Effects estimated there may overstate what a census-style institutional evaluation would find. The finding should be read as "this is how age behaves on a public rating platform," not as a universal law of the classroom.
- Correlation, not decline in skill. The study measures ratings, not learning. There is no evidence here that older instructors teach less effectively; the most defensible reading is that surface cues correlated with age move the ratings, which is a statement about the instrument, not the teacher.
- Cohort and selection effects. Age is confounded with academic generation, seniority, the kinds of courses senior staff are assigned, and survivorship (who is still teaching at 65 is a selected group). A cross-sectional age coefficient cannot fully separate "getting older" from "belonging to an earlier cohort."
- Small magnitude. The effect is genuinely small. It matters for fairness at the margins and for how we interpret fine-grained comparisons, but it is not a licence to discount older instructors' scores wholesale, nor a large enough effect to explain big differences on its own.
How Koji incorporates this
Koji for Education cannot change how students perceive an instructor, but it is designed to stop age-linked surface cues from silently driving the evidence a QA office acts on.
- Low-inference, behaviour-anchored items. Rather than a single global "overall quality" score — the measure most vulnerable to appearance and age halos — Koji foregrounds specific teaching behaviours (Did explanations come in more than one form? Did feedback arrive in time to use? Were sessions well organised?). Concrete behavioural evidence is far less permeable to an age or attractiveness halo than an impressionistic star rating.
- AI-moderated interviews that ask for evidence. Koji's conversational moderator pushes students past first impressions toward specific examples, and automatic thematic analysis reports what students actually cite. This dilutes the thin-slice, appearance-driven component that the age effect appears to travel on.
- Within-instructor trend framing with the right baseline. Because Koji foregrounds an instructor's own trajectory and benchmarks against comparable courses, it helps distinguish a genuine change in teaching from a slow, artefactual drift — the longitudinal trap age creates.
- Bias-aware reporting. Koji's reporting is designed to flag when demographic or course confounds may be shaping a comparison, prompting evaluators to interpret rather than rank on a raw number — the responsible stance the SET-reliability literature demands.
Koji is deliberate about not overclaiming: these mechanisms are designed to mitigate age-linked artefacts, not to eliminate human perception. The aim is to move the decision-relevant evidence away from the surface cues age rides on and toward what the instructor demonstrably did. Institutions that also run wider staff or product research use the same AI-moderated engine on Koji's core platform at koji.so, where the same principle — probe for specifics, distrust the snap impression — applies.
What this means for a fair evaluation policy
Because the age effect is small, appearance-mediated and about ratings rather than teaching, the right response is not to "correct" scores with an age adjustment — that would embed a contested causal claim into official data. The safer policy choices are structural. Avoid ranking instructors on a single global score, which is the measure most exposed to appearance and age halos; report a profile of specific behaviours instead. Interpret long-run individual trends cautiously: a gentle downward drift over a career may be an artefact, not a decline, so pair the trend with other evidence before acting on it. Watch for compounding: because the age effect runs through appearance, it can amplify gender and attractiveness biases and may fall harder on women, so it should never be analysed as if it were a clean, isolated variable. And prefer census-style institutional instruments over public rating sites as the basis for any decision — the Stonebraker and Stone effect is estimated on RateMyProfessors, a self-selected population that likely exaggerates surface-cue effects relative to a well-administered internal evaluation. The overarching principle is proportionality: the effect is real enough to matter for fairness at the margins of fine-grained comparisons, but far too small and too entangled with confounds to justify treating any individual older instructor's scores as suspect, or to explain a substantial score gap on its own.
Related resources
- Beauty Bias in Course Evaluations: Are Attractive Instructors Rated as Better Teachers?
- Do First Impressions Decide Your Course Evaluation? The Thin-Slice Evidence
- Is RateMyProfessors a Valid Measure of Teaching? What the Evidence Says
- Gender Bias in Student Evaluations of Teaching: Evidence and Mitigation
- Are Student Ratings Just a Mood? The Longitudinal Stability Evidence
- The Halo Effect in Course Evaluations: When One Impression Colours Every Rating
References
- Stonebraker, R. J., & Stone, G. S. (2015). Too Old to Teach? The Effect of Age on College and University Professors. Research in Higher Education, 56(8), 793–812. https://doi.org/10.1007/s11162-015-9374-y
- Ambady, N., & Rosenthal, R. (1993). Half a Minute: Predicting Teacher Evaluations From Thin Slices of Nonverbal Behavior and Physical Attractiveness. Journal of Personality and Social Psychology, 64(3), 431–441. https://doi.org/10.1037/0022-3514.64.3.431
Related articles
Gender Bias in Student Evaluations of Teaching: Evidence and Mitigation
A research-grounded synthesis of gender bias in student evaluations of teaching (SET), anchored to MacNell, Driscoll & Hunt (2015) and corroborated by Boring (2017) and Mengel, Sauermann & Zölitz (2019), with practical mitigations for European quality-assurance teams.
Beauty Bias in Course Evaluations: Are Attractive Instructors Rated as Better Teachers?
What the research says about physical-attractiveness bias in student evaluations of teaching — from Hamermesh & Parker (2005) to German and laboratory replications — and how to design evaluation so appearance does not masquerade as teaching quality.
The Halo Effect in Course Evaluations: When One Impression Colours Every Rating
When students like an instructor, that single global impression bleeds into their ratings of unrelated specifics. From Nisbett & Wilson (1977) to Feistauer & Richter (2018) and Cannon & Cipriani (2021), here is how the halo effect distorts itemised course evaluations — and what to do about it.
Do First Impressions Decide Your Course Evaluation? The Thin-Slice Evidence
Ambady and Rosenthal showed that silent 30-second clips of an instructor predict end-of-term ratings. What thin-slice judgments mean for the validity of course evaluations and how to design feedback that probes substance, not first impressions.