New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Sector trends9 min read

Evaluating Across Borders: The Comparability Problem in Transnational and Multi-Campus Programmes

When the same degree runs in Coventry, Dubai, and Kuala Lumpur, can a 4.1 in one mean the same as a 4.1 in another? Transnational and multi-campus education has outgrown the survey instrument it inherited—and the comparability assumptions baked into it.

Koji Education Team

Product ·

Bottom line up front: Transnational education (TNE) is now a structural feature of European higher education, not a fringe activity—UK institutions alone taught 663,970 students for a UK award overseas in 2024/25, up 7.8% year on year (British Council, HESA 2024/25). Yet most institutions evaluate a course delivered across three countries with the same Likert instrument they use at home, then compare the average scores side by side as if they were measuring the same thing. They are not. Cross-cultural differences in response style, language, and the very meaning of "good teaching" make a raw numeric comparison between campuses unsafe—and the regulatory framework (ESG 2015, the European Approach to joint programmes) demands quality equivalence, not score identity. Comparability has to be engineered into the instrument, not assumed from it.

TNE is big, growing, and structurally diverse

The scale is easy to underestimate. Beyond the UK's 663,970 overseas students, transnational provision spans franchised programmes, validated degrees, joint and double degrees, branch campuses, and online delivery, with markets led by China (about 14.3% of UK TNE enrolment), Sri Lanka, Malaysia, Egypt, and Greece (British Council, 2024/25). The most common mode is now students studying for a foreign award without a physical branch campus—through partners or online—which means the teaching, the local academic culture, and the student body can differ enormously while the badge on the certificate stays the same.

That diversity is the whole challenge. A quality office is asked to assure that the "same" programme delivers the "same" standard in Manchester and in Muscat. The instinctive move is to run the same evaluation survey everywhere and compare the means. It feels rigorous. It is statistically naive.

Why a 4.1 in one country is not a 4.1 in another

The problem is measurement non-invariance: the property that the same questionnaire can measure different things, or measure the same thing on different scales, across groups. Three mechanisms drive it in TNE.

First, response styles differ systematically across cultures. Decades of cross-cultural survey research document that respondents in some cultures gravitate to scale extremes (extreme response style) while others cluster around the midpoint, and that acquiescence—the tendency to agree—varies by culture. A cohort that culturally avoids extreme judgements will produce lower means than an equally satisfied cohort that does not, with no difference in actual experience. Comparing raw averages across such cohorts measures culture as much as teaching.

Second, the meaning of "good teaching" is not universal. Expectations about the appropriate distance between staff and students, the acceptability of challenging an instructor, the role of memorisation versus debate, and what counts as "support" all vary across educational cultures. A survey item like "the lecturer encouraged me to challenge ideas" is not a neutral measuring stick; it is a culturally loaded prompt that a student in one system reads as a virtue and a student in another reads as disorganisation.

Third, language and translation fracture equivalence. A questionnaire translated into the local language, or answered in English by non-native speakers, cannot be assumed to carry identical meaning. Subtle shifts in connotation—how strong "satisfied" feels versus its translated equivalent—move scores in ways that have nothing to do with quality. We have written about this effect within single institutions in our piece on language bias and international students; across borders it compounds.

The upshot: ranking campuses by mean evaluation score is, statistically, comparing thermometers calibrated in different units and declaring the warmest room.

What the regulators actually ask for

European quality assurance is clearer-headed about this than many institutional dashboards. The Standards and Guidelines for Quality Assurance in the European Higher Education Area (ESG 2015) are explicitly designed to apply to all higher education "regardless of place or mode of delivery", including transnational and cross-border provision (ESG 2015, ENQA), and the European Approach for Quality Assurance of Joint Programmes was adopted the same year to handle exactly the multi-country case. Crucially, what these frameworks demand is equivalent quality and standards—that students everywhere receive a comparably good education—not that a survey returns an identical number in every country. ENQA has separately emphasised transparent quality assurance to protect the interests of students on TNE programmes (ENQA).

This distinction matters enormously. An institution that chases score parity across campuses—pressuring the lower-scoring site to "improve" its numbers—may be manufacturing a false equivalence while masking a real difference, or erasing a measurement artefact that was never about quality at all. The regulatory ask is to understand difference, not to flatten it.

"But surely we have to compare campuses somehow?"

Yes—and this is the legitimate counterargument. An institution is accountable for equivalent standards across all delivery sites; it cannot simply declare every campus incomparable and stop measuring. Abandoning comparison is not an option.

The resolution is to compare carefully, not naively. Three approaches make cross-campus comparison defensible. Anchoring vignettes—asking respondents to also rate a fixed, described scenario—let you calibrate each cohort's use of the scale and adjust for response style; this is a well-established cross-cultural technique we cover in our vignettes piece. Testing for measurement invariance before comparing means tells you whether a comparison is even meaningful. And most powerfully, shifting weight from numeric scores to thematic evidence: "students at the Dubai campus consistently raised library access" and "students in Manchester raised timetable clashes" are directly comparable findings about real, fixable issues, in a way that "4.1 versus 3.8" never is. Themes survive translation and culture far better than scale midpoints do.

How AI-native evaluation helps

This is where a conversational, analysis-first instrument has a structural advantage over a static survey grid. Koji for Education replaces the single cross-campus Likert average with two things that travel across borders better: standardised conversational probing and automatic thematic analysis. Because the AI moderator is consistent, every campus gets the same quality of follow-up—removing the human-interviewer inconsistency that plagues multi-site qualitative work—while still adapting to what each student actually says. Its thematic analysis surfaces the issues raised at each site, so a programme director compares substance ("assessment workload" in one country, "placement support" in another) rather than culturally-skewed numbers. Multilingual collection lets students answer in the language they think in, and the structured question types (including scale items that can be anchored) let you keep the quantitative signal you need for accreditation while interpreting it in cultural context. Reporting rolls up to programme and institution level, which is exactly the altitude at which TNE equivalence must be demonstrated to bodies working from the ESG.

Koji does not claim to eliminate cross-cultural measurement bias—no instrument can make a midpoint-clustering cohort and an extreme-rating cohort produce identical numbers. It mitigates the danger by reducing how much weight rides on those numbers in the first place, and by making the comparable thing the evidence, not the average. The same conversational engine powers the main Koji platform for global customer research, where multinational teams hit the identical wall: a satisfaction score that means different things in different markets.

Why the stakes are rising, not falling

The comparability problem is becoming more acute, not less. As franchised and partner-delivered provision grows—now the dominant TNE mode—the physical and cultural distance between the awarding institution and the teaching site widens, and so does the regulators' scrutiny of whether quality is genuinely equivalent. Quality offices are increasingly asked to evidence equivalence to external reviewers, not merely assert it. An evaluation system that can only produce campus-by-campus averages offers weak evidence: a reviewer can reasonably object that the numbers are not comparable. A system that produces consistent, cross-site thematic findings—showing that the same issues are surfaced, probed, and acted on everywhere—offers something far more defensible. In a franchising landscape under growing political and regulatory pressure, the ability to demonstrate genuine, like-for-like understanding of the student experience across borders is becoming a compliance asset, not just a pedagogical nicety.

The takeaway

Transnational and multi-campus education has scaled faster than the instruments used to evaluate it. A raw comparison of mean evaluation scores across countries measures response style, language, and cultural expectation as much as teaching quality—and the European regulatory framework asks for equivalent standards, not identical scores. Comparability is something you build: through anchored, invariance-tested measurement and, above all, through thematic evidence that crosses borders intact. The institutions that get TNE evaluation right will stop asking "which campus scored higher?" and start asking "what is each campus's students telling us, in their own words, that we can act on?"