New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Methodology11 min read

Your International Cohort Is Not Happier — They Just Use the Scale Differently

When you compare course-evaluation scores across nationalities, an English-medium programme, or a transnational campus, you are partly measuring how different cultures use a rating scale — not how good the teaching was. The response-style trap, and how to escape it.

Koji Education Team

Product · August 20, 2026

Bottom line up front: When you compare course-evaluation means across international cohorts — different nationalities in the same room, an English-taught master's, a branch campus abroad — a meaningful chunk of the difference is not teaching quality at all. It is response style: the systematic, culturally patterned way people use a rating scale regardless of what they actually think. Some cultures agree readily (acquiescence), some reach for the endpoints (extreme response style), some cluster on the midpoint. These tendencies vary predictably by country, and answering in a second language pushes responses toward the middle. The upshot is uncomfortable: a lower score from an international cohort may signal a different scale habit, not dissatisfaction — and a higher one may be politeness. If you rank instructors or programmes on raw cross-cultural means, you are partly ranking nationalities. The fix is to stop leaning on the mean and lean on qualitative feedback, which response style contaminates far less.

What response styles are

Response styles are content-independent patterns in how someone answers scaled items. The three that matter most for evaluation are acquiescence response style (ARS) — a tendency to agree with items whatever their content; extreme response style (ERS) — a preference for the ends of the scale (1 and 5) over the middle; and midpoint or middle response style (MRS) — a pull toward the neutral centre. None of these is about the teaching. All of them move the number.

Locally, when everyone shares roughly the same scale habits, response style adds noise but not much bias to comparisons. Across cultures it becomes bias, because the habits themselves differ systematically — and that is precisely the situation European higher education is now in, with internationalised classrooms, English-medium instruction, and transnational, multi-campus provision that invites exactly the cross-group comparisons response style corrupts.

The evidence: culture shapes the scale, and so does language

The clearest evidence comes from Anne-Wil Harzing's 26-country study of response styles in cross-national survey research. It found major, systematic differences between countries, with national-culture characteristics — power distance, collectivism, uncertainty avoidance, and extraversion — significantly predicting both acquiescence and extreme response styles. In broad terms, more collectivist and higher-extraversion cultures tended toward more acquiescent and more extreme responding; the pattern is stable enough to be predictable, which is exactly what makes it a confound rather than random noise.

Harzing's work also surfaces a second, under-appreciated mechanism directly relevant to English-medium programmes: the language of the questionnaire changes the answers. English-language questionnaires elicited more middle responses than native-language ones, and responses in a respondent's own language were closer to their "true" response style. So an international student rating your course in English — a second or third language for many — is nudged toward the midpoint by the language itself, independent of any view about the teaching. The same phenomenon is well documented in international large-scale assessment: cross-national studies such as PISA have to model response-style differences explicitly before comparing self-reported attitudes across countries, precisely because the raw scores are not comparable.

This is a distinct problem from the familiar within-culture biases. It is not the same as generic acquiescence and careless responding, and not the same as central-tendency and range restriction in a single population. The cross-cultural version is more dangerous because it aligns with a group boundary — nationality, language, campus — so it masquerades as a real difference between cohorts.

Why the mean is the wrong statistic here

If two cohorts differ in extreme-response tendency, their means and their standard deviations both shift for reasons that have nothing to do with your course. Averaging Likert scores is already a questionable move on ordinal data; doing it across groups with different scale habits and then comparing the averages compounds the error. A 0.3-point gap between a domestic and an international cohort is not evidence of anything until you have ruled out response style — and usually you cannot, because you did not measure it.

What to do about it

Three defensible responses, in ascending order of usefulness.

First, test before you compare. Formal measurement-invariance and differential-item-functioning checks tell you whether an instrument even means the same thing across groups before you put their scores side by side. If it fails invariance, cross-group mean comparison is not valid, full stop.

Second, anchor the scale. Anchoring vignettes — short descriptions of hypothetical teaching that all respondents rate — let you calibrate how each group uses the scale and adjust accordingly. They are the standard cross-national correction, though they add length and burden.

Third, and most powerfully, shift weight off the number and onto the words. Response style is a property of scaled items. It has far less purchase on open-ended qualitative feedback, where a student describes what actually happened in their own terms. "The feedback on assignments came too late to use for the next one" carries the same information whether the student is culturally extreme, acquiescent, or midpoint-prone. Rich qualitative data routes around the response-style problem instead of trying to statistically correct it after the fact.

But isn't this just an excuse to ignore low scores from international students?

The sharpest objection, and it must be taken seriously: warning that international cohorts "use the scale differently" can become a convenient way to dismiss genuine dissatisfaction from the students most likely to face real problems — language barriers, belonging, inconsistent support. That would be a serious misuse of this argument, and the opposite of what the evidence supports.

The point is not that international students' low scores are artefacts to be explained away. It is that the number alone cannot distinguish a response-style effect from a real problem — which is exactly why you should not act on the number alone in either direction. The correct inference from response-style research is humility about the mean and more attention to the qualitative detail, not less. If anything, it raises the priority of finding out what international students actually experienced, because the scale is precisely the channel that fails them. Dismissing their feedback because "they score differently" gets the lesson exactly backwards: the fix for an unreliable number is better evidence, not a convenient story.

Where Koji fits

Koji for Education is built to reduce reliance on the one statistic response style corrupts most. Instead of a Likert battery whose means drift with nationality and language, Koji runs AI-moderated conversational interviews that elicit specific, described experiences — and it can conduct those interviews and analyse the responses across languages, so a student can speak in the language in which they express their true view rather than being pushed to the midpoint by a second-language English form. Its automatic thematic analysis then turns multilingual open-text responses into structured, comparable themes based on what students say happened, not on how they push a slider. That gives an international programme a basis for comparison that is grounded in content rather than in scale habits.

Koji does not magically make cross-cultural comparison trivial — no tool does, and the responsible move is still to treat cross-group means with caution and to run invariance checks where scores carry weight. What it changes is the centre of gravity of your evidence: away from a fragile number and toward the qualitative detail that survives cultural differences in scale use. Institutions doing multinational user or market research face the identical problem and can use the same conversational interview engine on the main Koji platform.

The uncomfortable truth is simple. If your international cohort scores lower, you do not yet know whether the teaching was worse or the scale was read differently — and the mean will never tell you. The words will.

Frequently asked questions

What is a response style? A content-independent pattern in how someone uses a rating scale — for example acquiescence (agreeing regardless of the item), extreme responding (favouring the endpoints), or midpoint responding (clustering on neutral). It reflects a habit of answering, not an opinion about the teaching.

Do response styles really differ by culture? Yes. Harzing''s 26-country study found that national-culture dimensions such as power distance, collectivism, uncertainty avoidance, and extraversion systematically predict acquiescence and extreme response styles. The differences are large and patterned enough to bias cross-group comparisons.

Does answering in English change the scores? The evidence suggests it does. English-language questionnaires elicit more middle responses than native-language ones, and responses in a person''s own language are closer to their true response style. An international student rating a course in English is nudged toward the midpoint by the language itself.

Does this mean we should ignore low scores from international students? No — the opposite. The lesson is that the number alone cannot separate a response-style effect from a real problem, so you should not act on the number alone in either direction. It raises the priority of understanding what international students actually experienced through qualitative feedback.

How do we compare course evaluations fairly across international cohorts? Run measurement-invariance or differential-item-functioning checks before comparing means, consider anchoring vignettes to calibrate scale use, and — most effectively — shift weight from Likert means to open-ended qualitative feedback, which response style contaminates far less.