New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Evaluation bias10 min read

Language Bias in Course Evaluation: Are International Students Answering the Same Question?

When 8.4% of EU students come from abroad and many programmes teach in English to non-native speakers, a course evaluation written in one language and one cultural register may not mean the same thing to everyone who answers it. The measurement-invariance problem, explained.

Koji for Education

Research & Editorial Team · June 10, 2026

Bottom line up front: A course evaluation only produces comparable numbers if every respondent interprets the items the same way. With 8.4% of tertiary students in the EU coming from abroad (Eurostat, 2023 data) — and far higher shares in some countries and programmes — that assumption is fragile. Non-native speakers and students from different cultural backgrounds systematically differ in how they read Likert items and in their response styles, which means cross-group comparisons of evaluation scores can mistake language and culture for teaching quality. This is a measurement-invariance problem, and most legacy survey tools do nothing about it.

The scale of the issue in Europe

International and mobile students are not a rounding error in European higher education. Eurostat reported in 2025 that 8.4% of tertiary students in the EU in 2023 came from abroad, with national shares ranging from around 3% in Greece to 52.3% in Luxembourg, 29.6% in Malta, and 22.3% in Cyprus; the Netherlands alone hosts roughly 122,000 international students. Add the large cohorts taught in English at universities where neither staff nor students are native English speakers, plus Erasmus+ mobility (around 386,000 credit-mobile EU graduates in 2023), and a typical European programme is evaluating teaching using an instrument answered by people with very different relationships to the language it is written in.

What measurement invariance means — and why it bites

In psychometrics, measurement invariance is the property that a scale measures the same construct, in the same metric, across groups. If an instrument is invariant, a "4 out of 5" on "the course was well organised" means the same thing for a German native speaker and a visiting student from Indonesia. If it is not invariant, the numbers are not comparable — and averaging or benchmarking across the groups produces artefacts. Invariance is not assumed; it is tested, typically with multi-group confirmatory factor analysis. The point of the concept for evaluation practice is humbling: comparability is an empirical question most institutions never ask.

The research shows invariance cannot be taken for granted across language groups. Studies that test for it — for example, the measurement-invariance analysis of the Student Opinion Scale across English and non-English-language learners (Frontiers in Psychology, 2016, 5,257 students) and invariance work on the Satisfaction with Life Scale for non-native English speakers (2024) — sometimes find invariance holds and sometimes find it does not, which is exactly the point: it has to be checked, scale by scale, rather than assumed.

Response styles travel with culture

Even when respondents understand items identically, how they use a rating scale differs systematically by cultural background. The landmark evidence is Harzing's (2006) 26-country study in the International Journal of Cross Cultural Management, which found major cross-national differences in acquiescence (the tendency to agree) and extreme response style (the tendency to use scale endpoints), correlated with cultural dimensions such as power distance, collectivism, and uncertainty avoidance. The methodological sting, as the cross-cultural literature puts it, is that systematic variance in response style across national or language groups can be mistaken for real differences in the thing you are trying to measure. A programme that compares its average evaluation score across a domestic and an international cohort, and reads the gap as a difference in teaching quality, may simply be reading a difference in how two groups use a 1–5 scale.

Layer on the obvious comprehension issue — academic-register survey items, double-barrelled questions, idioms, and reverse-worded items are all harder for non-native speakers, inflating measurement error and non-response — and the international student's evaluation score carries more noise and more systematic distortion than a domestic student's. Translating the instrument helps with comprehension but, done naively, can worsen invariance unless the translation is validated.

But isn't this a small effect we can ignore?

The strongest counterargument is pragmatic: response-style and language effects are real but small relative to the genuine signal, so for most decisions they wash out. This is sometimes true and should not be dismissed — for a single course with a mostly domestic cohort, language invariance is a second-order concern. But it fails precisely where evaluation is used most consequentially. When scores are benchmarked across programmes with very different international shares, aggregated to rank departments, or compared year on year as cohort composition shifts, small per-respondent biases become systematic between-group biases that move the rankings. An English-taught master's with 60% international students is not being measured on the same ruler as an undergraduate cohort that is 95% domestic — and treating their averages as comparable is the error, not the size of any individual effect. The honest response is not "it is small, ignore it" but "know when it matters, and stop comparing non-comparable numbers."

What good practice looks like

Several moves reduce the distortion. Prefer concrete, behaviourally anchored questions ("Did you receive feedback on your work in time to use it for the next assignment?") over abstract, culturally loaded ones ("Rate the overall quality of teaching"), because concrete items are more robustly understood across languages. Lean on open-ended, qualitative responses, which carry meaning even when the numeric scale does not translate cleanly. Offer the evaluation in the student's stronger language with validated translations where feasible. And, where comparisons across cohorts matter, report and interpret with the composition in mind rather than ranking raw means.

How Koji fits

Koji for Education is structured to reduce exactly these distortions. Its AI-moderated conversational interviews can be conducted in the student's preferred language and adapt their wording to the respondent, so a non-native speaker is not left parsing an academic-register Likert item — and the moderator can rephrase or clarify rather than letting a misread question become silent measurement error. Because Koji emphasises open-ended responses with automatic thematic analysis, it captures what students actually mean even where a 1–5 number would not be cross-culturally comparable, surfacing themes that hold up across language groups. Its standardised, bias-aware moderation applies a consistent approach to every respondent, and its programme- and institution-level reporting makes cohort composition visible rather than hiding it inside a single benchmarked average. To be precise: Koji reduces and surfaces language- and culture-driven measurement distortion — it does not eliminate cultural difference, nor should any tool claim to. (The same multilingual interview engine powers cross-market customer research on the main Koji platform, where comparing feedback across countries raises identical invariance questions.)

A worked example: one gap, three explanations

Consider a department running an English-taught master's (60% international) alongside a domestic-majority bachelor's, which notices the master's scores 0.3 lower on a 5-point "overall quality" item. The tempting reading is that the master's teaching is weaker and needs intervention. But at least three non-teaching explanations are live before that conclusion is warranted. The international cohort may use the scale differently — less extreme-positive responding is common in several of the cultural clusters in Harzing's data. The abstract item "overall quality" may be parsed differently by students educated in systems with other norms about what "quality" denotes. And non-native speakers facing an academic-register survey carry more measurement error, widening the gap purely as noise. Until measurement invariance across the two cohorts is established, the 0.3 difference cannot bear the weight of a teaching-quality verdict — and acting on it could divert resources on the basis of an artefact.

The discipline this implies is modest but rarely practised: before comparing two cohorts' scores, ask whether they were measured on the same ruler, and if you cannot show that they were, compare themes and specifics rather than means. Qualitative evidence degrades far more gracefully across languages than a single contested number does.

International students are among the people a European university most needs to hear from — and the ones a static, monolingual Likert form is least equipped to measure fairly. Asking whether everyone is really answering the same question is not pedantry. It is the difference between an evaluation that informs and one that quietly misleads.

See how multilingual, conversational evaluation gives every student a fair voice — explore Koji for Education.