Age Bias in Student Evaluations of Teaching: Does Getting Older Cost You Marks?
Instructor age shows up in student evaluation scores — but the effect is small, conditional, and tangled with gender. We read the evidence honestly and show what better feedback design can do about it.
Koji Education Team
Product · June 17, 2026
Short answer: Instructor age does appear in student evaluations of teaching (SET), but rarely as a clean, large effect. The most rigorous longitudinal evidence finds women's ratings fall as they move from their late twenties toward middle age while men's ratings rise over the same span. Age is less a standalone bias than a lens that amplifies other biases — gender chief among them. The practical implication is not "throw out the data," but "stop reducing a career's worth of teaching to one demographic-contaminated average."
What the evidence actually shows
The strongest single study comes from Jennifer Chatman and colleagues, published in Organizational Behavior and Human Decision Processes in 2022. They analysed thousands of student ratings of 126 business-school professors collected over fifteen years (2003–2017). The headline finding: women's teaching ratings declined significantly as they aged from young adulthood into middle age, reaching their lowest point around age 47, whereas men's ratings increased across the same age range. The decline for women persisted even after controlling for research productivity and parental status — which makes a pure "older faculty simply teach worse" explanation hard to sustain.
This pattern is consistent with earlier work. Stonebraker and Stone's 2015 study (Studies in Higher Education) found that instructor age negatively affected evaluations, but with a threshold: the effect did not really begin until instructors reached their mid-forties. Other studies report a "halo effect" favouring younger teachers, with students reporting greater perceived rapport with younger instructors. Systematic reviews of SET bias generally classify age as a real but modest non-instructional influence — smaller and less consistent than grading leniency or the gender effects that have been replicated many times.
There is a tempting shortcut here — pointing at RateMyProfessors patterns, where younger and "easier" professors often score higher. Treat those with caution: that platform is self-selected, unverified, and not the controlled institutional data quality assurance runs on. The defensible claim is narrower and more useful: age effects in SET are real, generally small, and overwhelmingly conditional on other characteristics.
Why "age" is almost never a main effect
The most important thing the evidence teaches is that asking "does age bias exist?" is the wrong question. Age interacts. A 50-year-old man and a 50-year-old woman teaching the same syllabus to comparable cohorts do not face the same evaluative climate. The Chatman data show the lines literally crossing: the same decade of life that lifts men's ratings depresses women's. That is a textbook interaction effect, and it has a mechanism — students appear to penalise middle-aged women for perceived deficits in warmth (a communal trait) while rewarding middle-aged men for perceived competence and authority (agentic traits). The bias is not "old," it is "old and female read against a gendered template."
This is why a department that benchmarks instructors against each other on a single average number — and then reads age into the gaps — is building a measurement error into a personnel decision. It is also why this topic sits next to our work on the statistics of small classes and why averaging Likert scores misleads: the average is exactly the place where these interactions disappear from view.
But critics argue the effect is too small to matter
This is the strongest counterargument and it deserves a fair hearing. Three versions of it:
-
The effect sizes are small. True. Age typically explains a small share of variance in ratings, and a few hundredths of a point on a five-point scale is not, on its own, an injustice. But small average effects become consequential precisely at the margins where SET is used as a high-stakes tiebreaker — promotion, contract renewal, "teaching excellence" awards. A bias too small to see in the aggregate can still decide a close case.
-
It could be confounding, not bias. Also fair. Older faculty are disproportionately assigned large, required, quantitative service courses — and those course types score lower for reasons unrelated to the instructor. Some of the raw age–rating correlation is course allocation, not student prejudice. The honest response is that good studies (like Chatman's) control for several confounds and the gendered pattern survives; but yes, naive age comparisons are partly confounded, which is an argument for better instrumentation, not against caring.
-
SET still contains signal. Correct, and we agree. None of this means student feedback is worthless. It means the number is a noisy, partly-contaminated proxy, and that the richer signal — what students actually struggled with, what helped them learn — is sitting in the open-text comments and the unasked follow-up questions, not in the demographic-tinted average.
What to do about it
You cannot make a student forget how old their lecturer looks. You can change what you ask, how you ask it, and what you do with the answer.
- Ask about the teaching, not the teacher. Items anchored to specific, observable behaviours ("The feedback on my assignment helped me improve") are harder to answer with a halo than global personality items ("This instructor is engaging"). See our guide to writing better questions.
- Triangulate. Student ratings are necessary but not sufficient. Pair them with peer observation and self-review, as we argue in triangulation in teaching evaluation.
- Read the distribution and the text, not just the mean. A bimodal split or a thematic pattern tells you far more than a 0.1-point gap, as we explain in why a sentiment score is not insight.
- Collect formatively, mid-cycle, so feedback informs teaching rather than only ranking it (formative vs summative).
How Koji helps — and what it does not claim
Koji does not eliminate age bias. No instrument can remove a student's impression of a person. What Koji does is reduce the room that impression has to masquerade as data. Its AI-moderated conversational interviews probe beyond a number: when a student says a course was "fine" or an instructor "boring," the moderator asks what specifically helped or hindered their learning, surfacing teaching evidence instead of a vibe. The moderation is standardised and bias-aware — every student meets the same patient, consistent interviewer, removing the human-moderator inconsistency that creeps into focus groups. Automatic thematic analysis of open-text feedback keeps every theme traceable to real quotes, so a committee reads what was said rather than an averaged stereotype. Programme- and institution-level reporting lets you see patterns across cohorts without collapsing them into one rankable figure, and it is built for GDPR/AVG-compliant, EU-appropriate data handling.
The same conversational interview engine underpins the main Koji platform used for general user and customer research — the difference here is the education-specific framing, not a different method.
If your promotion and quality processes still lean on a single averaged SET score, the age and gender literature is a good reason to upgrade the instrument. Talk to Koji for Education about collecting feedback that interrogates teaching, not demographics.
A quick self-audit for your own data
Before you trust — or distrust — your institution's evaluation scores, three checks expose whether age and gender are leaking into the numbers. First, plot mean rating against instructor age separately for men and women; if the lines diverge the way the Chatman data predict, you have an interaction, not a fair comparison. Second, hold course type constant: compare only instructors teaching similar class sizes and course levels, because the raw age gap is partly the large required-course assignments that fall to senior staff. Third, read the open text for the vocabulary of bias — comments about "warmth," "approachability," or an instructor being "old-fashioned" that never attach to teaching substance are a tell. None of these audits fixes the bias, but together they stop a contaminated average from being read as a clean verdict on a colleague's teaching — which is the difference between using student voice responsibly and weaponising it.
The bottom line
Age bias in student evaluations is real but modest, and almost always entangled with gender and course allocation. The mature response is neither to dismiss student voice nor to pretend the average is clean. It is to collect feedback rich enough that the teaching, not the teacher's age, is what comes through.