Accent and Origin Bias in Course Evaluations: What Rubin (1992) Revealed
Donald Rubin's classic experiment showed students rated an identical recorded lecture as harder to understand when they believed the instructor was Asian — evidence that perceived accent and origin bias course evaluations. What it means for QA in multilingual European universities, and how to design evaluation that resists it.
Koji Education Team
Product
Quick answer
Do students mark down non-native or foreign-perceived instructors on course evaluations because of accent? Yes — and strikingly, the bias operates even when there is no real accent at all. In Donald Rubin's (1992) experiment, undergraduates heard one identical lecture recorded by a native speaker of standard North American English, but were shown a photograph of either a Caucasian or an Asian "instructor." Those shown the Asian face rated the speech as more accented and scored lower on a listening-comprehension test of the same audio. The driver was the listener's expectation, not the speech. For multilingual European universities staffed by international faculty, this means course-evaluation scores can encode students' assumptions about an instructor's origin rather than the quality of teaching — a serious fairness problem that evaluation design must actively counter.
What the research says
The anchor study is Donald L. Rubin, "Nonlanguage factors affecting undergraduates' judgments of nonnative English-speaking teaching assistants," Research in Higher Education (1992), 33(4), 511–531. Rubin's central experiment is a model of how to isolate bias. Undergraduate participants listened to a four-minute recorded mini-lecture. The audio was identical across conditions and was recorded by a native speaker of standard North American English — there was genuinely no foreign accent to hear. What varied was a photograph projected during the lecture: in one condition a Caucasian woman, in the other an Asian woman, matched on other characteristics. Students then rated the instructor and completed a listening-comprehension test on the lecture.
The results were stark. Students who saw the Asian photograph perceived the (identical) speech as more accented and rated the instructor less favourably on teaching-relevant dimensions. More consequentially, their measured listening comprehension was lower — they understood the same recording less well — when they believed the speaker was Asian. The effect, Rubin argued, reflects "reverse" stereotyping: an attribution about the speaker's ethnicity/origin distorts perception of the speech itself, and the expectation of difficulty becomes a self-fulfilling barrier to comprehension. A second study identified attitudinal predictors of these ratings, and a third piloted intervention — undergraduates acting as "teaching coaches" — concluding that the burden cannot fall on non-native instructors alone; students need intercultural sensitization too.
Two further sources establish this as a robust, ongoing finding. Okim Kang and Donald Rubin (2009), "Reverse linguistic stereotyping: Measuring the effect of listener expectations on speech evaluation" (Journal of Language and Social Psychology, 28(4), 441–456), formalised the mechanism: when an "international" identity is ascribed to a voice with standard pronunciation, listeners hear it as less standard and judge it more harshly. The judgment tracks who listeners think is speaking, not what is actually said. This work shows the 1992 result was not a one-off but an instance of a general perceptual phenomenon. The broader bias literature — including Boring, Ottoboni, and Stark (2016), who showed student evaluations are more responsive to students' biases and grade expectations than to teaching effectiveness — situates accent/origin bias alongside the well-documented gender and grading-leniency biases as a reason not to treat raw SET scores as objective measures.
Why it matters for course evaluation in practice
European higher education is profoundly multilingual and international. English-medium instruction has expanded rapidly across the Netherlands, the Nordics, Germany, and beyond, and faculty are recruited globally; a large share of instructors teach in a language that is not their first, to students for whom it is also often not their first. Rubin's finding lands directly on this reality: when an evaluation reduces teaching to a global impression, that impression can absorb a student's assumptions about an instructor's nationality, ethnicity, or "foreignness" — even, as the experiment shows, in the absence of any real comprehension barrier.
The consequences are concrete and unjust. International and minority-ethnic faculty may receive systematically lower scores not because they teach worse but because students expect difficulty and perceive accordingly. If those scores feed contract renewal, probation, or promotion, the institution is importing a perceptual bias into employment decisions — with obvious equality-law and EDI (equality, diversity, inclusion) exposure under European frameworks. It also corrupts the data: a department cannot improve teaching by acting on a "clarity" score that is partly measuring students' stereotypes about origin.
Limitations and honest caveats
A careful reader should weigh several limits. Rubin's study is now three decades old, used a single short audio stimulus, and drew on US undergraduates; the precise effect size may not transfer to a semester-long European course where students accumulate real interactional evidence over months and can recalibrate. The photograph manipulation creates a strong, somewhat artificial cue that a real classroom delivers more gradually and ambiguously. "Accent" and "comprehensibility" are also genuinely multidimensional — some non-native speech really is harder to follow, and the literature does not claim that all perceived difficulty is bias; it claims that a measurable portion of perceived difficulty and lower ratings is driven by listener expectation rather than the speech signal. Disentangling genuine intelligibility issues (which deserve support and training) from pure stereotype effects is hard in field data. Finally, Rubin's own intervention study was a small pilot. The defensible conclusion is bounded but serious: perceived origin demonstrably biases comprehension and ratings independent of actual speech, so raw scores for internationally diverse faculty must be read with this in mind.
How Koji incorporates this
Koji for Education is built to reduce the surface area on which origin/accent stereotypes can attach, and to make any residual pattern visible rather than hidden.
- Behaviourally specific questions instead of a global "communication" rating. Reverse stereotyping fastens onto vague holistic judgments. Koji's AI moderator asks students about specific episodes — "Was there a point where you got lost? What was happening?" — and probes whether the difficulty was about a concept, the pacing, the materials, or genuinely the delivery. Forcing specificity separates real comprehension obstacles (which the department can act on) from a diffuse "hard to understand the lecturer" impression that often encodes expectation.
- Structured, multidimensional capture. Using distinct scale items (organisation, pacing, materials, responsiveness) alongside open_ended and single_choice questions prevents a single origin-laden impression from collapsing onto one summary score, paralleling the multidimensional logic that improves SET validity generally.
- Automatic thematic analysis with bias-aware reporting. Koji clusters open text into themes and lets QA staff compare patterns across cohorts and over time. Persistent, evidence-grounded themes (e.g. "slides went too fast") read very differently from sparse, non-specific complaints — and reading scores in context, rather than ranking faculty on raw means, is exactly what the bias literature recommends. Koji is designed to mitigate origin/accent bias by changing what is collected and how it is interpreted; it cannot eliminate a stereotype that lives in the student's perception.
- Triangulation across sources. Because a single biased cohort impression is fragile, Koji supports comparing student feedback with peer observation and learning-outcome evidence, so a low "clarity" score from one group is not taken at face value.
The same AI-moderated interview engine powers Koji's core research platform at koji.so for product and customer research, where separating a respondent's expectations from their actual experience is an everyday methodological necessity.
A practical protocol for international faculty
For a European university with substantial international staffing, the Rubin findings argue for a specific, defensible protocol around how accent-adjacent feedback is collected and used.
- Distinguish intelligibility support from evaluation. Genuine intelligibility difficulties deserve a constructive, developmental response — pronunciation and communication support offered to any instructor whose students report concrete, repeated comprehension problems. That developmental track should be kept separate from summative evaluation used in renewal decisions, so that a stereotype-driven impression cannot quietly become an employment liability.
- Require specificity before action. A bare low "clarity" or "communication" score should never trigger a consequence on its own. Rubin's result implies that such scores can encode expectation rather than experience. Ask for the moment and the cause: where did comprehension break down, and was it the concept, the pace, the slides, the audio, or the delivery? Specific, reproducible reports across students are actionable; diffuse, non-specific ones often are not.
- Read across cohorts and time. A single cohort's impression is fragile and easily coloured by group expectations. Persistent, consistent patterns across multiple cohorts carry far more evidential weight than one noisy semester, and longitudinal reading lets an instructor's genuine improvement show.
- Share the burden with students. Rubin's third study found that intercultural sensitization for students complements skills training for instructors. Brief framing that reminds students to judge teaching behaviour rather than perceived origin, plus committee guidance on reading these scores, shrinks the space in which reverse stereotyping operates.
The goal is to support instructors who face real communication barriers while refusing to let students' assumptions about nationality or ethnicity become an unexamined input to consequential decisions — a stance aligned with European equality obligations.
Related Resources
- /docs/gender-bias-student-evaluations-teaching
- /docs/response-styles-likert-cross-cultural-evaluation
- /docs/student-written-comments-course-evaluation
- /docs/student-evaluations-teaching-and-learning-meta-analysis
- /docs/why-averaging-likert-scores-misleads-course-evaluation
References
- Rubin, D. L. (1992). Nonlanguage factors affecting undergraduates' judgments of nonnative English-speaking teaching assistants. Research in Higher Education, 33(4), 511–531. https://doi.org/10.1007/BF00973770
- Kang, O., & Rubin, D. L. (2009). Reverse linguistic stereotyping: Measuring the effect of listener expectations on speech evaluation. Journal of Language and Social Psychology, 28(4), 441–456. https://doi.org/10.1177/0261927X09341950
- Boring, A., Ottoboni, K., & Stark, P. B. (2016). Student evaluations of teaching (mostly) do not measure teaching effectiveness. ScienceOpen Research. https://doi.org/10.14293/S2199-1006.1.SOR-EDU.AETBZC.v1
Related articles
Gender Bias in Student Evaluations of Teaching: Evidence and Mitigation
A research-grounded synthesis of gender bias in student evaluations of teaching (SET), anchored to MacNell, Driscoll & Hunt (2015) and corroborated by Boring (2017) and Mengel, Sauermann & Zölitz (2019), with practical mitigations for European quality-assurance teams.
Do Student Evaluations Measure Learning? The Uttl Meta-Analysis Revisited
A research-grounded reading of the Uttl, White & Gonzalez (2017) meta-analysis on the SET–learning relationship, with implications for European course evaluation and quality-assurance policy.
What Can Open-Text Student Comments Tell You That Likert Scores Cannot?
A research-grounded guide to open-ended student comments in course evaluation: what Alhija and Fresko found about who writes them and what they contain, how thematic analysis surfaces issues numbers miss (Stupans et al.), how rare abusive comments actually are (Tucker), and how to turn free text into reliable evidence.
Response Styles and Likert Scales: Why Cross-Cultural Evaluation Needs More Than Numbers
Acquiescence and extreme response styles vary systematically by culture (Harzing, 2006; Baumgartner & Steenkamp, 2001), which means raw Likert averages are not directly comparable across nationalities in Europe's multinational classrooms. This article explains the evidence and how to evaluate fairly across diverse cohorts.