New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Evaluation bias8 min read

Sexual Orientation Bias in Student Evaluations: The Least-Studied Demographic Bias

Gender, race, age, accent and attractiveness bias in student evaluations are well documented. Sexual orientation is the demographic bias almost nobody measures — and the experimental evidence that exists should worry quality-assurance teams.

Koji Education Team

Product ·

Student evaluations of teaching (SET) have a well-documented bias literature covering instructor gender, race and ethnicity, age, accent, and physical attractiveness. Sexual orientation is the conspicuous gap: a small experimental literature suggests that when students know an instructor is gay or lesbian, ratings can decouple from actual teaching quality — yet almost no institution examines this, because sexual orientation appears in no registrar dataset and disclosure varies instructor by instructor. This post reviews what the evidence actually shows, why the replication record complicates the story, and what a rigorous quality-assurance office should do about a bias it cannot directly audit.

The experimental evidence

Because sexual orientation is rarely recorded, the evidence base is experimental rather than observational — researchers manipulate whether students believe an instructor is gay or lesbian and observe what happens to ratings.

The most striking design is Ewing, Stukas and Sheehan (2003), published in the Journal of Social Psychology. A male and a female lecturer each delivered either a strong or a weak lecture. Half the students were led to believe the lecturer was gay or lesbian; half received no orientation information. The result: lecture quality strongly influenced ratings only when sexual orientation was unspecified. When students believed the lecturer was gay or lesbian, the quality manipulation stopped mattering — strong and weak lectures earned similar ratings. That is a validity failure in the strictest sense: the instrument stopped measuring the thing it exists to measure.

A year earlier, Russ, Simonds and Hunt (2002) ran a field experiment published in Communication Education. A male instructor delivered the same lecture, with identical delivery and immediacy cues, to 154 undergraduates across eight introductory communication classes; only his disclosed sexual orientation varied. Students who believed the instructor was gay rated him as significantly less credible and reported that they learned considerably less — despite receiving word-for-word the same teaching. The authors memorably described coming out in the classroom as an "occupational hazard."

The replication record — and why honesty matters here

If this blog has one editorial rule, it is that we report the evidence that complicates our own argument. Here it is: the Russ et al. finding has not replicated cleanly. Boren (2018), in Communication Studies, re-ran the 2002 design more than fifteen years later and found that gay instructors were not rated lower on credibility or perceived learning. A 2024 community-college replication by Sundblad and Dansereau likewise found the original effect largely absent.

Three readings are possible, and a careful quality-assurance office should hold all three simultaneously:

  1. Attitudes have shifted. Social acceptance of gay and lesbian people rose substantially between 2002 and 2018 in the populations studied; the bias may genuinely have shrunk.
  2. The effect is heterogeneous. Single-campus experiments with confederate instructors have limited power and narrow samples. Null results in one Californian or Midwestern sample do not establish absence across disciplines, regions, or national cultures — and virtually all of this literature is North American. Whether findings transfer to European classrooms is an open question, as we have argued for SET bias research generally.
  3. The mechanism may have moved. Overt derogation is rarer; subtler penalties — on "professionalism," on perceived rigor, or interacting with gender and other identities — may persist below the resolution of a small experiment. The intersectionality literature shows demographic penalties rarely operate one axis at a time.

What the literature does not support is confident dismissal. The Ewing et al. finding — that quality stops predicting ratings once orientation is salient — has never been directly refuted; it has mostly been left unexamined. A measurement system used for reappointment, tenure and programme review should not need to be proven guilty beyond reasonable doubt before an institution takes precautions. That standard applies nowhere else in quality assurance.

Why this bias is structurally invisible

For gender bias, institutions can — and should — audit their own data: instructor gender is in the HR system. Sexual orientation is different in three ways.

It is invisible to observational audit. No European institution holds (or should hold) a register of instructor sexual orientation — under GDPR it is special-category data with strict processing limits, as is any orientation information surfacing in free-text feedback. You cannot run the regression, so the usual audit playbook fails.

Exposure is selective and asymmetric. Bias operates only where students perceive or infer orientation — through disclosure, inference, or rumor. Openly LGBTQ+ instructors thus carry a rating risk their colleagues do not, which is precisely the "occupational hazard" mechanism: the penalty attaches to openness, and the rational response — concealment — carries its own well-documented costs to wellbeing and to students who benefit from visible role models.

High-stakes use amplifies small effects. When evaluation scores feed tenure and promotion decisions and contingent faculty are renewed on the strength of a decimal point, even a modest, patchy bias becomes a career-relevant lottery.

What institutions can actually do

Since direct auditing is off the table, the response has to work through instrument design and process — reducing the surface area on which any identity-based halo or penalty operates.

Anchor items in observable teaching behaviors. Global items ("rate the overall quality of this instructor") maximize room for affective, identity-linked judgment. Low-inference, behaviorally specific items — about feedback turnaround, clarity of assessment criteria, structure — constrain it.

Weight specifics over sentiment. A rating with no reasoning behind it is exactly where bias hides. Feedback that must articulate what happened — "the weekly problem sets came back annotated within a week" — is harder to contaminate with feelings about who the instructor is.

Triangulate. No personnel decision should rest on student ratings alone; multiple evidence sources — peer observation, teaching portfolios, learning outcomes — dilute any single instrument's biases.

Treat comment streams with duty of care. Derogatory identity-based comments should be filtered before they reach the instructor, with institutional follow-up — the same infrastructure needed for abusive comments generally.

Where Koji fits

Koji for Education was designed around exactly this failure mode: numbers that absorb identity while shedding evidence.

  • AI-moderated conversational interviews probe every rating for the behavioral episode behind it. When a student cannot ground a judgment in anything that happened in the course, that is visible in the data rather than laundered into a clean-looking average.
  • Automatic thematic analysis organizes open-text feedback into teaching-relevant themes — assessment, pacing, materials, support — so committees read what students experienced, not raw comment streams where identity-based remarks sit unflagged. Quality scoring surfaces which responses carry substantiated, actionable content.
  • Standardized, bias-aware moderation asks every student the same questions the same way, with none of the inconsistency of human-led focus groups, and its structured question types support the low-inference item design the measurement literature recommends.
  • GDPR-appropriate handling matters doubly here, given that free-text feedback can surface special-category data about instructors or students.

None of this eliminates orientation bias — no instrument can, and vendors who claim otherwise should be read skeptically. What evidence-anchored, probed, thematically analyzed feedback does is shrink the space in which unexamined identity judgments masquerade as measurements of teaching.

Institutions that also run staff, alumni or applicant research face the same problem on a different population — the main Koji platform applies the same AI interview engine to general user and stakeholder research.

The takeaway

Sexual orientation bias in SET is understudied, not disproven. The strongest experiment on record found that knowing an instructor is gay or lesbian can sever the link between teaching quality and ratings; the replication record suggests the effect is neither universal nor stable, and is essentially unmapped in Europe. Because this bias cannot be audited out of observational data, the only defensible response is structural: behaviorally anchored instruments, evidence-weighted feedback, triangulation, and honest uncertainty in reporting. If your evaluation system still hangs careers on a global satisfaction average, it is exposed to every bias in this series — including the one you cannot see in your data.

Koji for Education replaces the static form with AI-moderated conversational evaluation that asks what happened — and gives your committees evidence instead of averages. Book a demo to see it on your own courses.