New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Evaluation bias10 min read

Does the US Research on Biased Student Evaluations Apply to European Universities?

Most of the famous studies on bias in student evaluations of teaching come from North American campuses. European quality officers are right to ask whether those findings transfer. The honest answer is uncomfortable for both sceptics and believers: the strongest causal evidence on gender bias is actually European — but some other findings really are more US-bound.

Koji Education Team

Product ·

The short answer: It is reasonable to worry that North American findings on biased student evaluations of teaching (SET) may not travel to European higher education — the grading cultures, tenure systems and student populations differ. But the objection cuts less far than sceptics hope. The two cleanest causal studies of gender bias in SET were run at European universities (in the Netherlands and France) and both found it. Other findings — particularly on physical attractiveness and, to a lesser extent, race — rest more heavily on US samples and should be treated as hypotheses to test locally, not settled facts. The right posture for a European QA office is neither "the bias literature is American, ignore it" nor "all of it applies here", but "assume gender bias until you have checked, and gather your own evidence for the rest."

Why the question is legitimate

A common and fair critique of the SET-bias literature is that it is disproportionately North American. US institutions differ from European ones in ways that could plausibly change whether bias appears: heavier reliance on SET for tenure decisions, different grade distributions and grade-inflation norms, "student-as-consumer" framing amplified by high tuition, more racially heterogeneous classrooms, and evaluation instruments with different item wording. If bias is partly a product of those conditions, a study from a large US state university tells a Dutch or Portuguese quality officer relatively little.

This is a real methodological concern — it is the problem of external validity, or in psychometric terms the transportability of a finding across populations. A result can be internally valid (correctly identified within its sample) yet fail to generalise to a different context. Dismissing the concern would be intellectually dishonest. So would using it as a blanket excuse to ignore the evidence.

The inconvenient fact: the best gender-bias evidence is European

Here is what complicates the sceptic's position. The two studies with the strongest causal designs on gender bias in SET are not American at all.

Maastricht University, Netherlands. Mengel, Sauermann and Zölitz analysed 19,952 student evaluations in a business-school setting where students were randomly assigned to instructors — the closest thing to an experiment you can get in a real university. Random assignment removes the usual confounds (better students clustering with certain teachers, self-selection into courses). Women received systematically lower teaching evaluations than men, even though the instructor's gender did not affect students' actual grades or their self-study hours. The bias was driven mainly by male students, was larger in mathematical courses, and fell hardest on junior women (Mengel, Sauermann & Zölitz, 2019, Journal of the European Economic Association).

A selective university in France. Anne Boring used data from a French institution and found that male students rated male professors more highly, and that the teaching dimensions students associated with men (leadership, knowledge) matched gender stereotypes — despite evidence that students learned as much from women as from men (Boring, 2017, Journal of Public Economics).

Two different European countries, two different methods, the same direction of effect — and in the Dutch case, a design strong enough to support a causal claim. When the causal backbone of the gender-bias finding is European, "that is just American research" is not available as a rebuttal.

The often-cited US experiment — MacNell, Driscoll and Hunt's online study, in which the same instructors taught under swapped gender identities and the "female" identity was rated lower across most categories (MacNell, Driscoll & Hunt, 2015, Innovative Higher Education) — then becomes corroboration from a third design in a third country, not the sole pillar. Convergence across the US, the Netherlands and France is exactly the pattern you want before treating a bias as real. We survey this convergence in more depth in our piece on how strong the evidence for biased evaluations actually is.

Where the sceptic has a point

Transportability is not all-or-nothing, and honesty requires conceding where the European evidence is genuinely thinner.

  • Attractiveness bias in SET rests substantially on US and, in some cases, RateMyProfessors-style data. The mechanism (a halo from perceived attractiveness) is plausibly universal, but the magnitude in a European classroom with different evaluation instruments is not well established. Treat it as likely-present-but-unquantified rather than proven.
  • Racial and ethnic bias findings are heavily shaped by the specific racial composition and history of US campuses. Europe's relevant fault lines often run differently — around nationality, migration background, accent and language rather than the US racial categories — which is why accent and language bias for international instructors may be the more locally relevant construct. The underlying phenomenon (in-group favouritism, stereotype-driven judgement) is general; the categories that carry it are local.
  • Grade-driven leniency effects interact with national grading cultures. Where European systems use tighter or externally-moderated grade distributions, the grading-leniency dynamic documented in the US may operate with different force.

So the correct European position is differentiated: gender bias is well-evidenced on European data and should be assumed; attractiveness, race/nationality and leniency effects are plausible but under-measured locally and warrant your own evidence before you rely on — or dismiss — them.

But doesn't this just mean we should throw out student evaluations entirely?

No — and this is the counterargument worth addressing head-on, because it is where many good-faith readers land. Demonstrating that SET carries bias is not the same as showing it carries no information. The bias findings are an argument against naive use — ranking instructors by raw average, treating a 0.2-point gap as a verdict, letting a single number drive a promotion — not against listening to students at all. Student voice is a legitimate and, under the ESG standards for European quality assurance, an expected part of quality systems. The task is to collect it in a way that does not smuggle stereotype-driven judgement into a high-stakes number.

A second fair objection: bias effect sizes in these studies are often modest. True — but "modest on average" and "decisive at the margin" coexist. When the effect is concentrated (male students, quantitative courses, junior women) and the decision is binary (renew or not), a small average bias can flip individual cases. Modest mean effects are precisely why you need confidence intervals and distributions rather than point estimates.

What a European QA office should actually do

  1. Assume gender bias and design against it. Do not compare male and female instructors on raw scores as if the playing field were level; report distributions, use appropriate adjustment, and never let a single averaged number carry a personnel decision.
  2. Gather local evidence for the rest. Run your own analyses on your own data for accent/nationality, discipline and leniency effects before assuming US magnitudes apply — or that they do not.
  3. Shift weight from the number to the reasons. Bias lives most powerfully in the unexplained gap between a male and female instructor's average. It has far less purchase when you can see why students rated as they did.

How Koji helps

Koji for Education is built for that last move — from opaque numbers to interpretable reasons. Its AI-moderated conversational interviews probe beyond a rating, so instead of a bare score you get the reasoning behind it, where stereotype-driven or content-irrelevant judgements become visible rather than baked silently into an average. The moderation is standardised and bias-aware: every student meets the same neutral interviewer, removing the human-facilitator inconsistency that can itself introduce bias, and the AI is designed to probe reasons rather than lead. Automatic thematic analysis lets a QA office run exactly the kind of local, evidence-based check this article recommends — surfacing whether comments about warmth, authority or accent pattern by instructor gender or background in your institution, not in a US dataset. Programme- and institution-level reporting supports the distribution-first, comparison-with-care approach the statistics demand, and everything is GDPR/AVG-compliant.

Koji does not claim to eliminate bias — no instrument can. It mitigates and surfaces it: by capturing reasons, standardising the moderator, and making patterns auditable, it moves bias from an invisible thumb on the scale to something you can see, measure and act on. That is a fairer basis for both improvement and accountability than a legacy Likert form from EvaSys or a Qualtrics survey that records the number and discards the why.

The same conversational interview engine powers the main Koji research platform for teams doing wider user and staff research. If you want to know whether the bias literature applies to your campus rather than someone else's, see how Koji for Education surfaces the reasons behind the ratings.