Can You Compare Course Ratings Across Cultures? Anchoring Vignettes and the King Method
When students from different countries interpret the same rating scale differently, their scores are not comparable. King, Murray, Salomon and Tandon (2004) introduced anchoring vignettes to correct this. Here is how the technique works and what it means for international and multi-campus course evaluation.
Koji Education Team
Product
Answer (BLUF): If two cohorts use a 1–5 teaching-quality scale differently — say, German exchange students reserve "5" for the exceptional while some peers hand it out freely — their average scores are not directly comparable, and ranking units by raw means imports that difference as if it were teaching quality. Anchoring vignettes, introduced by King, Murray, Salomon and Tandon (2004), fix this by asking every respondent to rate the same short hypothetical scenarios on the same scale; differences in how they rate the fixed scenarios reveal how each respondent uses the scale, which you then use to rescale their self-reports. It is the most rigorous published method for making cross-group rating comparisons defensible — and it carries real assumptions you must test before trusting it.
What the research says
A persistent problem in survey research is what Gary King and colleagues call interpersonal incomparability: different respondents attach different meanings to the same response categories. A "good" rating from a severe judge and a "very good" from a lenient one may reflect identical underlying reality. When this varies systematically by group — nationality, language, discipline, prior expectation — comparing group means measures the groups' scale use as much as the thing being rated. This is the same family of problem covered in our notes on measurement invariance and reference bias — but those tools mostly detect the problem. Anchoring vignettes try to correct it.
In their American Political Science Review paper Enhancing the Validity and Cross-Cultural Comparability of Measurement in Survey Research (2004), King, Murray, Salomon and Tandon proposed the method. Alongside the self-assessment ("How would you rate the clarity of your lecturer?"), respondents rate several vignettes — short descriptions of hypothetical people fixed for everyone (e.g., a lecturer who reads slides verbatim and never takes questions; a lecturer who checks understanding and adapts pace). Because the vignette is identical for every respondent, any variation in how respondents rate it is variation in scale use, not in the target. You then re-express each person's self-rating relative to where they placed the vignettes — anchoring their answer to a common, externally fixed yardstick.
The method rests on two assumptions the authors state explicitly. Response consistency: a respondent uses the rating scale the same way when judging the vignettes as when judging themselves (or their own course). Vignette equivalence: the level of the underlying construct portrayed in a vignette is perceived as the same by all respondents. When these hold, the rescaled self-assessments become comparable across groups.
The approach was developed and validated in the World Health Organization's cross-national health surveys. Salomon, Tandon and Murray (2004), in the BMJ, showed that self-rated health was badly incomparable across countries until anchoring vignettes were used to adjust for differing reporting styles, after which apparently paradoxical international rankings became coherent. King and Wand (2007), in Political Analysis, then supplied practical tools for selecting and evaluating vignettes and a nonparametric estimator, making the method usable without heavy parametric assumptions. Across these sources the message is consistent: cross-group comparisons of subjective ratings are untrustworthy by default, and vignettes are one of the few principled ways to repair them.
Why it matters for course evaluation in practice
European higher education is structurally multi-cultural: Erasmus mobility, joint and multi-campus programmes, English-medium degrees with globally mixed cohorts, and international branch campuses all routinely compare teaching ratings across nationality and language groups.
- Raw cross-cohort league tables are suspect. If your international Master's scores 4.1 and the domestic equivalent scores 4.4, the gap may be reporting style, not teaching. Acting on it — in promotion, programme review, or resource allocation — risks penalising staff for their cohort's scale habits.
- Cross-cultural response styles are documented and directional. Some cultures exhibit more acquiescence or more extreme-response style; others cluster toward the midpoint (see response styles and Likert scales). Vignettes turn that nuisance into something estimable.
- Vignettes separate "harsh raters" from "worse teaching." By fixing the scenario, you can ask: given two cohorts that rate the same described lecturer differently, how much of their difference on the real lecturer is just scale use? That decomposition is exactly what accreditation reviewers should want before comparing units.
- It strengthens evidence for quality assurance. Under ESG and national frameworks, the credibility of evaluation evidence matters. "We anchored cross-cohort comparisons to common vignettes" is a far stronger methodological claim than "we compared the raw means."
Limitations & honest caveats
Anchoring vignettes are powerful but not magic, and a critical reader should hold both assumptions to the light.
- Response consistency can fail. A student might judge a described lecturer by different criteria than they judge their own, breaking the link the method depends on. Where consistency fails, the correction is biased.
- Vignette equivalence is itself cross-cultural. The fix assumes everyone reads the scenario as the same level of quality — but interpretation of a vignette can vary by culture and language too, partially reintroducing the problem it aims to solve. Careful vignette wording and piloting are essential, and King and Wand's (2007) selection tools exist precisely because some vignettes behave badly.
- Respondent burden. Each anchored item needs several vignettes, multiplying questionnaire length. That collides directly with survey fatigue and falling response rates — you may buy comparability at the cost of completion.
- Construct describability. Some teaching qualities (warmth, inspiration) are hard to pin down in a short, equivalent vignette, limiting where the method applies.
- Limited validation in teaching evaluation specifically. The strongest evidence is in health and political-efficacy surveys. Applying vignettes to course evaluation is well-justified by analogy but is not yet backed by a large teaching-specific literature, so claims should be appropriately measured.
The reasonable posture is to use vignettes selectively — for high-stakes cross-cohort comparisons where the comparability threat is real — rather than bolting them onto every routine module survey.
How Koji incorporates this
Koji treats cross-group comparability as a first-class measurement problem rather than assuming raw means are comparable.
- Vignette-style anchoring as structured items. Koji's structured question types (
scale,single_choice,ranking) can present a fixed described scenario and ask the respondent to rate it, giving you the common yardstick King et al. require. Used on key items, these anchors let analysis separate scale use from teaching signal across cohorts. - Conversational elicitation of each respondent''s frame. Where formal vignettes are too burdensome, Koji's AI-moderated interview can probe what a student means by "good teaching" — a qualitative analogue to anchoring that surfaces the respondent''s reference point directly, rather than inferring it statistically.
- Bias-aware, cohort-segmented reporting. Koji reports distributions by cohort and flags cross-group comparisons as requiring care, so a committee is not handed a single league-table number that silently conflates reporting style with quality. This is designed to mitigate — not eliminate — incomparability.
- Triangulation across cohorts and sources. Rather than resting a cross-cultural comparison on one anchored scale, Koji combines structured ratings, open-text themes, and repeated cycles, reducing the weight any single comparison must bear.
The same engine powers Koji's core research platform at koji.so, where cross-market customer research faces the identical challenge: a "satisfaction" score means different things in different countries, and anchoring or conversational framing is needed before the numbers can be compared.
A worked example
Suppose a joint European Master''s runs the same module on two campuses, and the Barcelona cohort returns a mean clarity rating of 4.4 while the Munich cohort returns 4.0. Before concluding that the Munich delivery is weaker, embed one anchoring vignette: "A lecturer reads directly from slides, rarely pauses, and does not check whether students follow." If the Munich students rate that fixed vignette 1.5 and the Barcelona students rate it 2.5, the Munich cohort is simply using the lower end of the scale more readily. Rescaling the self-reports against the vignette can shrink — or even reverse — the apparent 0.4-point gap. The decision rule is straightforward: reserve vignette anchoring for comparisons that actually cross a plausible response-style boundary (nationality, language of instruction, discipline culture) and that carry real consequences. For same-cohort, term-on-term tracking of a single lecturer, the comparability threat is small and the extra vignette burden is rarely justified.
Related Resources
- Measurement Invariance: Can You Compare Course-Evaluation Scores Across Groups at All?
- Response Styles and Likert Scales: Why Cross-Cultural Evaluation Needs More Than Numbers
- Why "I Learned a Lot" Can''t Be Compared Across Courses: Reference Bias
- Interpreting and Reporting Student Ratings Responsibly
- Generalizability Theory and the Reliability of Student Ratings
- Response-Shift Bias: Why Self-Reported Learning Gains Can Mislead
References
- King, G., Murray, C. J. L., Salomon, J. A., & Tandon, A. (2004). Enhancing the Validity and Cross-Cultural Comparability of Measurement in Survey Research. American Political Science Review, 98(1), 191–207. https://gking.harvard.edu/files/gking/files/vign.pdf
- Salomon, J. A., Tandon, A., & Murray, C. J. L. (2004). Comparability of self rated health: cross sectional multi-country survey using anchoring vignettes. BMJ, 328(7434), 258. https://www.semanticscholar.org/paper/02a50a32b2c256a6812ebbb9054ac5342d5b6089
- King, G., & Wand, J. (2007). Comparing Incomparable Survey Responses: Evaluating and Selecting Anchoring Vignettes. Political Analysis, 15(1), 46–66. https://gking.harvard.edu/files/abs/c-abs.shtml
- Schwarz, N., Knäuper, B., Hippler, H.-J., Noelle-Neumann, E., & Clark, L. (1991). Rating Scales: Numeric Values May Change the Meaning of Scale Labels. Public Opinion Quarterly, 55(4), 570–582. https://doi.org/10.1086/269282
Related articles
Generalizability Theory and the Reliability of Student Ratings: Why One Class Is Not Enough
Reliability is not one number. Generalizability theory (Gillmore, Kane & Naccarato 1978; Marsh 1984) decomposes the variance in student ratings into student, teacher, course and occasion components — and shows that a single class can be reliable for the course yet a poor estimate of the teacher. What that means for fair evaluation.
Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees
The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.
Measurement Invariance: Can You Compare Course-Evaluation Scores Across Groups at All?
Comparing average ratings across departments, languages or online vs paper assumes the questionnaire means the same thing to every group. Measurement invariance is the test that assumption usually fails — and almost nobody runs it.
Response Styles and Likert Scales: Why Cross-Cultural Evaluation Needs More Than Numbers
Acquiescence and extreme response styles vary systematically by culture (Harzing, 2006; Baumgartner & Steenkamp, 2001), which means raw Likert averages are not directly comparable across nationalities in Europe's multinational classrooms. This article explains the evidence and how to evaluate fairly across diverse cohorts.