New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods10 min read

Translating a Course Evaluation Is Not Translation: Back-Translation, TRAPD, and Cross-Language Equivalence

Running the same course evaluation in several languages requires more than a good translator. Brislin's back-translation, the TRAPD model, and the ITC Guidelines explain how to keep items equivalent across languages.

Koji Education Team

Product

In brief

Running one course-evaluation instrument across several languages — routine in multilingual and international European universities — is not a translation task; it is a measurement-equivalence task. A literal translation can be linguistically perfect yet measure something subtly different, so that a French and a Dutch cohort's scores are no longer comparable. The methodological consensus has moved from Brislin's (1970) back-translation as a sole check toward committee-based approaches — most prominently TRAPD (Translation, Review, Adjudication, Pre-testing, Documentation) — and is codified in the ITC Guidelines for Translating and Adapting Tests (2nd ed., 2018). The non-negotiable final step is empirical confirmation of equivalence (measurement invariance) before you compare scores across language versions.

What the research says

The historical anchor is Richard Brislin's 1970 paper in the Journal of Cross-Cultural Psychology, "Back-Translation for Cross-Cultural Research." Working with 94 bilinguals across ten languages, Brislin formalised back-translation: one translator renders the source into the target language, an independent translator renders that target version back into the source language, and discrepancies between the original and the back-translated source reveal translation problems. Back-translation was a major advance and remains useful — but the field has learned its limits. A back-translation can look clean while both forward and backward translators share the same misreading, and "translationese" (literal, awkward but technically accurate wording) can pass back-translation yet read unnaturally to real respondents, changing how they answer.

The modern standard is the committee or team approach, best known through Janet Harkness's TRAPD model, developed in the cross-national survey tradition (the European Social Survey among others). The acronym names five steps: Translation (two independent translators each produce a draft), Review (translators and survey-method experts meet to reconcile drafts), Adjudication (a designated adjudicator makes the final binding decisions), Pre-testing (the agreed version is tested with real respondents, ideally via cognitive interviewing), and Documentation (every decision is recorded for transparency and future reuse). The crucial conceptual shift is that translation quality is a team product evaluated against how target respondents will actually understand and answer, not a single bilingual's output checked by reversing it.

The authoritative practice document is the ITC Guidelines for Translating and Adapting Tests (2nd edition, 2018, International Journal of Testing), produced by an International Test Commission committee (Bartram, Berberoglu, Grégoire, Hambleton, Muñiz, van de Vijver and colleagues). Its 18 guidelines are organised into pre-condition, test-development, confirmation, administration, scoring/interpretation, and documentation categories. Two themes matter for course evaluation. First, adaptation, not literal translation: an item should be rendered so it functions equivalently in the target context, which sometimes means changing surface wording. Second, the "confirmation" category insists on empirical evidence of equivalence — that is, statistical measurement invariance testing — before scores from different language versions are pooled or compared. The broader conceptual map for the kinds of non-equivalence (construct bias, method bias, item bias) is van de Vijver and Tanzer's (2004) framework, which links translation choices to the specific threats they create.

Why it matters for course evaluation in practice

For a European university running programmes in, say, English, German, and the national language, the temptation is to treat the language versions as interchangeable and report one combined number per item. Three practical risks follow.

1. Non-equivalent items make cross-language comparison meaningless. If "the lecturer was approachable" is rendered into a target language with a word that connotes informality rather than availability, the two cohorts are answering different questions. A gap between language groups may then reflect translation, not teaching.

2. Response-scale and anchor wording shift in translation. Verbal anchors ("strongly agree", "often") do not have exact cross-language equivalents; their perceived intensity changes, interacting with the cross-cultural response-style differences documented elsewhere (see Response Styles and Likert Scales). The anchoring-vignette and invariance literatures exist precisely because of this (see Anchoring Vignettes and the King Method).

3. Skipping pre-testing imports avoidable ambiguity. Cognitive interviewing in each language (see Before You Field It, Test It) catches the items that are technically translated but practically confusing — the cheapest possible insurance against a year of incomparable data.

The disciplined workflow for a multilingual evaluation is therefore: adapt (not just translate) using a TRAPD-style team; pre-test each version with real students; document every adaptation; and before comparing across languages, test for measurement invariance (configural, metric, scalar — see Measurement Invariance and Differential Item Functioning). Only items that reach at least metric/scalar invariance should be compared quantitatively across language groups; the rest should be reported within-group.

Limitations and honest caveats

Several honest caveats temper the picture. First, full scalar invariance is often unattainable, especially across many languages and cultures; partial invariance (only some items invariant) is the realistic outcome, and it constrains which comparisons are legitimate rather than licensing all of them. Second, the methods are resource-intensive: a full TRAPD process with two translators, a review committee, an adjudicator, and per-language cognitive pre-testing is more than many quality-assurance offices can run for every instrument every cycle — so prioritisation (do it thoroughly once for the core instrument, then reuse the documented version) is essential. Third, back-translation is not worthless — the nuance is that it is a useful supplementary check, not a sufficient sole method; dismissing it entirely overcorrects. Fourth, adaptation introduces a judgement-versus-comparability tension: the more you adapt an item to read naturally in each context, the better it measures within each group but the harder it can be to defend strict cross-group equality. Fifth, evidence generalises imperfectly: much translation methodology comes from large cross-national social surveys and psychological testing, not specifically from short course-evaluation forms, so the elaborate apparatus should be scaled to the stakes — heavier for instruments feeding accreditation or personnel decisions, lighter for low-stakes formative pulse checks.

How Koji incorporates this

Koji is built for multilingual student populations, and its design maps onto the translation-equivalence literature in concrete ways.

  • Adaptation-aware multilingual instruments, with documentation. Koji supports running the same study across languages while keeping a structured record of each language version of every question (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) — the Documentation step of TRAPD operationalised, so adaptations are visible and reusable rather than ad hoc.
  • Conversational probing reduces reliance on a single fragile wording. Because Koji's AI moderator conducts a short interview rather than presenting one fixed sentence, it can detect when a student has misread an item and clarify in the moment — a continuous, lightweight analogue of cognitive pre-testing that catches some translation ambiguity that a static form would silently carry.
  • Equivalence-conscious reporting. Koji reports language/cohort composition alongside results and supports within-group reporting, so an institution is not nudged into pooling language versions before equivalence has been established. This respects the ITC "confirmation" principle that comparability must be demonstrated, not assumed.
  • Triangulation across cohorts. Where cross-language comparison is not defensible, Koji's thematic analysis of open-text answers lets a programme director see what students in each language group raised, which is comparable at the level of themes even when Likert means are not strictly invariant.

Koji frames these as supports for good practice, not substitutes for it: the platform does not certify measurement invariance for you, and a serious multilingual comparison still requires the team-based adaptation and statistical confirmation the literature prescribes. The same multilingual, AI-moderated interview engine powers Koji's core research platform at koji.so, where cross-language equivalence is just as critical for international customer research.

A worked example

Consider a master's programme taught in both English and the national language, evaluating the item "the assessment criteria were clear." A literal translation might render "clear" with a word that leans toward "simple" rather than "transparent," so the national-language cohort effectively answers a subtly easier question and scores higher. Back-translation alone might miss this if both the forward and backward translators share the same slanted reading of "clear." A TRAPD team, by contrast, would surface the ambiguity at the review meeting; the adjudicator would settle on wording that unambiguously conveys transparency of criteria; and cognitive pre-testing with a handful of students in each language would confirm that respondents actually interpret the item as intended rather than as ease of the task. Only then would the programme test scalar invariance across the two versions. If the item reached invariance it could be compared directly across cohorts; if not, the two language groups would be reported separately rather than pooled into a misleading combined mean. The half-day of additional process prevents a full year of quietly incomparable data — and pre-empts the entirely reasonable question an accreditation panel will ask: "are these two numbers measuring the same thing?"

Related resources

References

  • Brislin, R. W. (1970). Back-translation for cross-cultural research. Journal of Cross-Cultural Psychology, 1(3), 185–216. https://doi.org/10.1177/135910457000100301
  • International Test Commission. (2018). ITC Guidelines for Translating and Adapting Tests (Second Edition). International Journal of Testing, 18(2), 101–134. https://doi.org/10.1080/15305058.2017.1398166
  • Harkness, J. A. (2003). Questionnaire translation. In J. A. Harkness, F. J. R. van de Vijver, & P. Ph. Mohler (Eds.), Cross-Cultural Survey Methods (pp. 35–56). Wiley. (TRAPD model)
  • van de Vijver, F. J. R., & Tanzer, N. K. (2004). Bias and equivalence in cross-cultural assessment: an overview. Revue Européenne de Psychologie Appliquée / European Review of Applied Psychology, 54(2), 119–135. https://doi.org/10.1016/j.erap.2003.12.004

Related articles

analysis-reporting

Measurement Invariance: Can You Compare Course-Evaluation Scores Across Groups at All?

Comparing average ratings across departments, languages or online vs paper assumes the questionnaire means the same thing to every group. Measurement invariance is the test that assumption usually fails — and almost nobody runs it.

analysis-reporting

Can You Compare Course Ratings Across Cultures? Anchoring Vignettes and the King Method

When students from different countries interpret the same rating scale differently, their scores are not comparable. King, Murray, Salomon and Tandon (2004) introduced anchoring vignettes to correct this. Here is how the technique works and what it means for international and multi-campus course evaluation.

research-methods

Response Styles and Likert Scales: Why Cross-Cultural Evaluation Needs More Than Numbers

Acquiescence and extreme response styles vary systematically by culture (Harzing, 2006; Baumgartner & Steenkamp, 2001), which means raw Likert averages are not directly comparable across nationalities in Europe's multinational classrooms. This article explains the evidence and how to evaluate fairly across diverse cohorts.

research-methods

Should Every Point on a Course-Evaluation Scale Be Labelled? The Evidence on Verbal Anchors

Labelling every response option, not just the endpoints, tends to raise reliability and tame response styles — but it also interacts with the number of categories and the language of your students. What the rating-scale-format research says for course evaluation.