New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Evaluation bias10 min read

When Bias Compounds: The Case for an Intersectional Lens on Student Evaluations of Teaching

Most evidence on bias in student evaluations isolates one axis at a time — gender, or race, or accent. But penalties can compound at the intersections, and a single-axis analysis can leave the most affected instructors statistically invisible. Here is what the evidence shows and what to do about it.

Koji Education Team

Product ·

Bottom line up front: The research literature on bias in student evaluations of teaching (SET) is now substantial, but most of it isolates one axis at a time — gender or race or accent. Real instructors are not single-axis. The evidence increasingly suggests that disadvantages can compound at the intersections: an instructor who is both a woman and a member of a racial minority, or a woman teaching a quantitative subject from a non-native-English background, can face a penalty larger than either factor predicts alone. An institution that corrects for one axis at a time can leave the most affected staff statistically invisible. Taking bias seriously means looking at the intersections — and pairing that with feedback instruments that surface what students actually responded to.

The single-axis evidence is strong — and incomplete

There is little serious doubt that SET scores carry bias. The cleanest demonstrations are experimental. In MacNell, Driscoll and Hunt's much-cited 2015 study, online instructors were presented to students under a male or a female identity while everything else — content, assignments, communication — was held constant. The instructor perceived as male received markedly higher ratings; on promptness, for example, the "male" identity scored about 4.35 versus 3.55 for the "female" identity, despite identical response times. Because teaching was literally held constant, the gap can only be attributed to perceived gender.

Large-scale observational work points the same way. Mengel, Sauermann and Zölitz, analysing 19,952 evaluations at Maastricht University where students were effectively randomly assigned to instructors (Journal of the European Economic Association, 2019), found women received systematically lower evaluations than men even though students' grades and study hours were unaffected by instructor gender. The bias was driven largely by male students, was larger in mathematical courses, and was especially pronounced for junior women. Parallel literatures document penalties tied to race and ethnicity and to non-native accents.

Each of these studies is rigorous. But each, by design, holds the world to one axis. That is exactly where the picture gets incomplete.

Why intersections matter statistically, not just morally

The concept of intersectionality — that overlapping social identities can produce distinct, compounded experiences of disadvantage — is often discussed as an ethical frame. For evaluation, it is also a measurement issue with concrete consequences.

Consider an institution that runs two clean analyses: one comparing men to women, another comparing white instructors to instructors of colour. Each might show a modest, "manageable" average gap. But averages within a group conceal the corners. If the women-of-colour subgroup carries a penalty substantially larger than the gender gap and the race gap added together, neither headline analysis will reveal it — the effect is diluted by the white women and the men of colour in each respective comparison. The people most disadvantaged by the instrument are precisely the ones a main-effects-only model is least equipped to see.

This is not hypothetical. Reviews of the SET literature and analyses of large online rating platforms report that instructors of colour, and women of colour in particular, tend to receive lower ratings than white men on overall quality, clarity and helpfulness even when instruction is held constant, and that effects can be especially pronounced for instructors from non-English-speaking backgrounds in science faculties — an intersection of gender, ethnicity, language and discipline. The signal lives at the crossing, not on either road alone.

The methodological trap: thin cells

Here is the uncomfortable practical reality. The right way to detect compounded bias is to model interaction terms (gender × ethnicity × discipline) or to compare specific intersectional subgroups. But the moment you cross several categories, your subgroups — the "cells" — get small. A single department may have only one or two women-of-colour instructors in a quantitative subject. Estimates for tiny cells are noisy and unstable, which is the same small-sample fragility that plagues all course-evaluation statistics.

This creates a genuine dilemma. Ignore the intersections and you miss the people most affected. Slice too finely and you risk over-interpreting noise — or, worse, exposing identifiable individuals and breaching confidentiality. There is no purely statistical escape. What you can do is (a) pool intersectional data across departments and across years to build cells large enough to estimate stably, (b) report uncertainty honestly rather than point estimates, and (c) stop relying on the rating number alone as your only evidence.

"But isn't this just slicing the data until you find something?"

This is the strongest counterargument and it should be met head-on. Run enough subgroup comparisons and pure chance will hand you a "significant" gap somewhere — the multiple-comparisons problem is real, and intersectional analysis multiplies the comparisons.

Three disciplines keep intersectional analysis honest rather than fishing. First, theory-led, pre-specified hypotheses: you test intersections the literature already predicts (e.g., women of colour, non-native-accent women in STEM), not every possible crossing. Second, correction and humility: adjust for multiple comparisons and treat thin-cell estimates as provisional, never as a personnel verdict. Third, convergent evidence: an intersectional gap in the numbers is a flag to investigate, corroborated against narrative comments, peer observation and student-experience data — not a standalone conclusion. Used this way, intersectionality is not data-dredging; it is a targeted hypothesis about where a known measurement flaw is likely worst.

A second, fair objection: won't this be weaponised to discount all SET data and shield underperformers? No — the goal is not to discard student voice but to stop misreading it. Bias-aware analysis distinguishes "this score reflects the instructor's identity" from "this score reflects the teaching," which protects good teachers from a biased instrument and keeps the genuine signal usable.

What to do about it — and where Koji fits

Mitigating intersectional bias is partly statistical and partly about the instrument itself. On the analysis side: pool across units and time, model interactions where cells allow, adjust for course-structural confounds, and never make a high-stakes decision on a thin cell. None of this eliminates bias — no method can — but it surfaces it instead of averaging it away. And surfacing is the precondition for everything else: an institution cannot mitigate, train against, or contextualise a penalty it has statistically erased through aggregation, which is why the analytical posture matters as much as any single corrective intervention.

The deeper leverage is in how feedback is collected. A bare Likert score is maximally exposed to halo and identity-driven judgement, because it gives the student nowhere to go but a global impression. Richer, behaviourally specific feedback is harder to contaminate with stereotype. This is the design premise of Koji for Education. Its AI-moderated conversational interviews ask students to ground a rating in specifics — what helped them learn, which explanation worked — pulling responses away from diffuse global judgement toward concrete teaching behaviour. Because the same standardized, bias-aware AI moderator runs every interview, you remove the human-moderator inconsistency that adds its own bias layer to traditional interview-based review. Automatic thematic analysis lets QA teams compare the substance of feedback across instructor groups — a far better test of whether a low score tracks teaching or identity than a number alone. Six structured question types (open-ended, scale, single- and multiple-choice, ranking, yes/no) let you separate course-design feedback from instructor feedback, so a poorly-timetabled module is not silently charged to the lecturer. And programme- and institution-level reporting can hold intersectional analysis behind appropriate confidentiality thresholds, in a GDPR/AVG-compliant pipeline built for European norms — connecting naturally to the broader question of how strong the evidence for SET bias really is and the documented gender and racial and ethnic effects.

The same conversational engine powers the main Koji platform for customer and user research, where surfacing why behind a rating is just as decisive.

Single-axis bias analysis was a necessary first step. Treating instructors as whole people — and building feedback that resists stereotype in the first place — is the next one. Koji is built for institutions ready to take it.

Want feedback that resists bias by design? Explore Koji for Education.