New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Evaluation bias9 min read

The Accent Penalty: Do Non-Native-Speaking Instructors Get Lower Course Evaluations?

Matched-guise experiments show students can "hear" an accent that isn't there and rate teaching as worse. For Europe's multilingual, English-medium classrooms, that is a validity problem hiding inside every numeric SET score. Here is what the evidence actually shows — and what to do about it.

Koji Education Team

Product · July 9, 2026

Bottom line up front: There is credible experimental evidence that students downgrade instructors they perceive as non-native speakers of the language of instruction — sometimes even when the speech they hear is identical to a native speaker's. The effect is real, it is partly an artefact of expectation rather than comprehension, and it contaminates the numeric scores universities use to compare instructors. It cannot be "averaged away". It can, however, be surfaced and partly mitigated by changing how feedback is collected. For Europe's rapidly expanding English-medium instruction (EMI) sector, this is not a niche concern — it is a structural threat to the fairness of student evaluations of teaching (SET).

Why this matters more in Europe than anywhere else

European higher education has internationalised faster than its evaluation instruments have. Thousands of programmes now teach in English to mixed cohorts, taught by academics who are themselves non-native English speakers. That means both sides of the evaluation — the rater and the rated — are frequently operating in a second language. The instrument that sits in the middle, a five-point Likert scale asking whether the lecturer "communicated clearly", was never validated for this situation.

Start from a sobering baseline. The most-cited recent meta-analysis of multi-section studies, Uttl, White and Gonzalez (2017), concluded that SET ratings explain at most about 1% of the variability in actual student learning once small-sample and publication-bias artefacts are controlled — a non-significant relationship. If ratings barely track learning at their best, then any systematic, learning-irrelevant factor that does move them deserves scrutiny. Perceived accent is exactly such a factor.

The "dialect hallucination" finding

The foundational study is Rubin (1992), published in Research in Higher Education. Using a matched-guise design, Rubin played US undergraduates a recording of a short lecture spoken in standard, unaccented American English. Everyone heard the same recording. The only thing that changed was a photograph projected alongside it: for one group, a Caucasian woman; for another, an Asian woman.

Students who believed they were listening to an Asian instructor reported hearing a foreign accent that was not on the tape, and — more consequentially — scored measurably worse on a comprehension test of the identical audio. Perception of ethnicity manufactured a perception of accent, which manufactured a perceived drop in teaching clarity. Researchers later named this "dialect hallucination" or reverse linguistic stereotyping: the accent is heard because it is expected, not because it is there.

This is the single most important thing a quality-assurance officer needs to understand about accent effects. At least part of the penalty is not a response to real communication difficulty. It is a response to a category the student has assigned the instructor before the teaching is even processed.

What the European evidence adds

The picture is not simply "any accent is punished". Well-designed European studies suggest the effect is graded and context-dependent. Nejjari and colleagues found that Dutch and German students evaluated lecturers with moderate non-native-accented English quite differently from those with strong accents, and that a slight accent did not always incur a status penalty. Hendriks, van Meurs and Usmany (2023), in Language Teaching Research, separated two things that student ratings routinely confuse: intelligibility (can I actually understand this person?) and attitudinal evaluation (how competent/likeable do I judge them?). Strong accents reduced measured intelligibility, but attitudinal downgrades did not map cleanly onto real comprehension loss. A contextualised speaker-evaluation experiment in Belgium similarly found that perceived ethnicity and language variation shaped undergraduates' judgements of instructors independently of content.

The honest reading of this literature is nuanced, and a PhD audience deserves the nuance: sometimes students face a genuine comprehension cost with a heavy accent, and that is legitimate feedback a programme should act on (better acoustics, slower pacing, captioning, language support). But the ratings instrument cannot tell you which case you are in. A low "communication" score could mean "I could not follow the material" or "I decided, before the first sentence, that this person would be hard to follow." Those require opposite institutional responses, and a Likert mean collapses them into one uninterpretable number.

Critics argue: "This is just students reporting a real problem"

The strongest counterargument — and it is a serious one — is that comprehension matters, students are the people best placed to report when they cannot understand a lecturer, and treating accent-related feedback as "bias" risks dismissing legitimate concerns about teaching quality in a second language. If a cohort genuinely cannot follow a course, that is a real quality signal, not prejudice.

This objection is correct on its own terms, and any honest system must preserve that signal. But it does not rescue the numeric score, for three reasons. First, the Rubin result shows comprehension reports can move even when comprehension conditions are held constant — so a low score is not self-certifying evidence of a real problem. Second, the same heavy-accent-driven intelligibility issue and the pure expectation penalty produce the same number, so the score cannot discriminate the actionable case from the unfair one. Third, attribution effects mean students often blame the person rather than the situation (poor room audio, dense material, their own tiredness). The goal is not to ignore comprehension complaints — it is to collect them in a form precise enough to act on and fair enough to defend.

Where Koji fits — collecting the reason, not just the rating

Koji for Education is built on a different premise: that the useful information about clarity lives in why a student says communication was hard, not in the digit they circle. Instead of a static SET form, Koji runs an AI-moderated conversational interview that can probe a vague signal into an actionable one.

  • When a student flags a communication concern, the moderator asks a neutral, standardized follow-up — "Can you point to a specific moment or type of material where it was hard to follow?" — the same probe every student receives, with no human moderator's tone or expectation shaping the answer. That consistency is exactly what human-run oral feedback cannot guarantee.
  • Koji's automatic thematic analysis of open-text and conversational responses separates "the lecture audio was inaudible in that hall" or "the notation was never defined" from unfocused impressions — surfacing whether a genuine, fixable comprehension barrier exists or whether the signal is diffuse and attitudinal.
  • Its six structured question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) let a programme pair a comprehension rating with a concrete, behavioural follow-up, so a number never travels alone.
  • Bias-aware, standardized AI moderation removes the interviewer-inconsistency that plagues focus groups, while quality scoring flags low-information responses. Koji is designed to mitigate and surface accent-linked distortion — never to claim it "eliminates bias", which no instrument can.

Because the same conversational interview engine powers the main Koji platform for general user and customer research, the underlying method — probing beyond a number to the reason behind it — is battle-tested well outside the classroom.

A practical protocol for QA offices

You do not need to wait for a new tool to act on this evidence. Four steps help immediately: (1) never use a raw "communication" mean to compare instructors across an internationalised programme without triangulating against peer observation and learning evidence; (2) collect the specific comprehension barrier, not just a score, so language support and AV fixes can be targeted; (3) report distributions and free-text themes to committees, not just averages, so a bimodal "half loved it, half struggled" pattern is visible; and (4) treat a low clarity score on an EMI course as a question to investigate, never a verdict on the instructor.

Accent effects are one of the clearest cases where the number lies with a straight face. The remedy is not to stop asking about clarity — it is to ask in a way that tells you whether the problem is in the room or in the reception.

Curious what conversational, bias-aware course evaluation looks like on your own programmes? Explore Koji for Education and see how probing the "why" changes what you learn.