Should Every Point on a Course-Evaluation Scale Be Labelled? The Evidence on Verbal Anchors
Labelling every response option, not just the endpoints, tends to raise reliability and tame response styles — but it also interacts with the number of categories and the language of your students. What the rating-scale-format research says for course evaluation.
Koji Education Team
Product
In brief: The evidence favours labelling every response category with words, not just the two endpoints. Fully verbally labelled scales tend to produce higher test-retest reliability (Weng, 2004) and shift response-style behaviour — Weijters, Cabooter and Schillewaert (2010) found endpoint-only labelling increases acquiescence and extreme responding, while full labelling reduces them. The catch: full labelling interacts with how many points you use and with the language proficiency of your respondents, so it is a strong default rather than a universal rule.
A design choice hiding in plain sight
Every course-evaluation question that uses a rating scale forces a quiet decision: do you write a word next to each point ("Strongly disagree / Disagree / Neither / Agree / Strongly agree"), or do you label only the ends and leave the middle as bare numbers (1–7, with "Poor" at one end and "Excellent" at the other)? Most institutions never deliberate this; the format is inherited from whatever template the previous system used. Yet a substantial psychometric literature shows that this choice changes the meaning of the numbers students return — how reliable they are, and how systematically they tilt toward agreement or extremity. For an evaluation programme whose scores feed module review and accreditation evidence, that is not a cosmetic detail.
What the research says
The most directly relevant experimental work is Weijters, Cabooter and Schillewaert (2010), in the International Journal of Research in Marketing. They manipulated two features of rating-scale format — the number of response categories and whether categories were fully or only partially (endpoint) labelled — and measured the effect on three response styles: net acquiescence (the tendency to agree regardless of content), extreme responding (clustering at the scale ends), and misresponse to reverse-worded items. Their headline result is that label format matters as much as the number of points. Endpoint-only labelling, especially combined with more categories, was associated with more acquiescence and more extreme responding; fully labelling each category gave respondents a shared verbal frame that disciplined these tendencies. In other words, bare numeric points invite students to impose their own idiosyncratic interpretation of "what a 5 means," and that idiosyncrasy shows up as systematic bias.
This dovetails with Weng (2004) in Educational and Psychological Measurement, a study of 1,247 college students — a population much closer to course evaluation than a marketing panel. Weng crossed the number of response categories (3 to 9) with the labelling approach (every option labelled vs endpoints only) and examined both coefficient alpha and test-retest reliability. Two findings stand out. First, scales with very few categories produced lower reliability, particularly lower test-retest reliability. Second, and central here, scales with all options clearly labelled yielded higher test-retest reliability than endpoint-only scales. Verbal anchors appear to stabilise how a given student maps an internal judgment onto a number from one occasion to the next — exactly the property you want when comparing the same module across cohorts.
The authoritative synthesis is Krosnick and Presser (2010) in the Handbook of Survey Research. Their review of decades of questionnaire-design experiments concludes that fully labelled scales are generally preferable: verbal labels on every point improve reliability and validity and reduce the cognitive burden of translating a feeling into an unlabelled number. They also note the practical ceiling — it becomes hard to write distinct, ordered verbal labels for very long scales (nine or eleven points), which is one reason their guidance pairs full labelling with a moderate number of categories (commonly five to seven).
Read together, the three sources converge on a clear direction: verbal labels on every category are a better default than endpoint-only numeric scales because they reduce response-style contamination and improve the consistency of the measure — provided you do not demand more categories than you can sensibly label.
Why it matters for course evaluation in practice
The implications are concrete and, for many institutions, mean editing an existing instrument rather than buying a new one.
1. Label the middle, not just the ends. A 1–5 scale showing only "Poor" and "Excellent" leaves the meaning of 2, 3, and 4 to each student's imagination. The same nominal score then means different things to different students — noise that no amount of downstream analysis can recover. Writing "Poor / Fair / Good / Very good / Excellent" gives every respondent the same yardstick.
2. Response styles are not random noise — they bias comparisons. Acquiescence and extreme responding are systematic. If endpoint-only labelling inflates extreme responses, two instructors with genuinely similar teaching can diverge on the reported mean simply because their students used the unanchored scale differently. Full labelling shrinks that artefact, which directly improves the fairness of any cross-module or cross-instructor comparison — a recurring concern for quality committees.
3. Labelling interacts with scale length. Because legible verbal labels are hard to write beyond about seven points, the labelling decision and the number-of-points decision are linked. A fully labelled five- or seven-point scale is a defensible, evidence-aligned default; an eleven-point endpoint-only scale is the configuration the research most warns against.
4. It matters more, not less, for international cohorts. Where students differ in language background and in cultural response style, a shared verbal frame does more work. Anchoring every point reduces the room for divergent interpretations of bare numbers — relevant to any European programme teaching a multilingual student body.
Limitations and honest caveats
A careful reader should resist over-generalising from these studies.
- Population transfer. Weijters et al. (2010) studied consumer-marketing respondents; their response-style findings are robust but were not collected in an end-of-term teaching-evaluation context, where motivation and stakes differ. Weng (2004) used college students, which is reassuring, but a single instrument and setting.
- Labels themselves can be unequal. Fully labelling a scale only helps if the verbal anchors are ordered with roughly equal psychological spacing. Poorly chosen labels ("Excellent / Very good / Good / Acceptable / Terrible") create their own distortions; the benefit assumes competent label construction, not labelling per se.
- Translation risk. The advantage of verbal anchors can erode across languages: "Good" and its nearest translation may not sit at the same point on the latent scale. For multilingual evaluations, labels must be validated per language, not merely translated literally — otherwise you trade numeric ambiguity for translation non-equivalence.
- Effect sizes are moderate. These are reliability and response-style improvements, not transformations. Full labelling will not rescue a badly conceived item or a leading question; it is one component of sound scale design.
- Reverse-worded items remain double-edged. Weijters et al. link misresponse to reversed items with scale format, but reverse-wording has its own well-documented hazards (see the companion article on acquiescence). Full labelling reduces, but does not eliminate, the need to handle reversed items with care.
The honest summary: full verbal labelling is a well-supported default that improves consistency and curbs systematic response styles, conditional on good label-writing and per-language validation. It is not a guarantee of validity.
How Koji incorporates this
Koji is an AI-native, AI-moderated course-evaluation platform, and its design reflects the rating-scale evidence rather than treating scale points as an afterthought.
- Fully labelled scale items by default. When a scale question is used, Koji is designed to present verbally anchored response options rather than bare numeric points, so that every student shares the same frame — the configuration Weng (2004) and Krosnick and Presser (2010) associate with higher reliability.
- Reducing reliance on the single Likert number. The deeper mitigation is structural. Because response styles contaminate closed ratings, Koji's AI-moderated conversational interview probes beyond the number: when a student selects a rating, the platform can ask why, capturing open-ended reasoning that does not inherit acquiescence or extreme-response bias. A "4 — because the labs were excellent but the lectures dragged" is worth more than a bare 4, and it is robust to how the student reads the scale.
- Bias-aware reporting and quality scoring. Koji's analysis layer is built to flag patterns consistent with straightlining and extreme responding, and its quality scoring distinguishes substantive answers from low-effort ones — so a scale tilted by response style is less likely to be read at face value.
- Multilingual anchoring. For international cohorts, the conversational layer elicits reasoning in the student's own words, reducing dependence on whether a translated label sits at exactly the same latent point — a direct hedge against the translation-equivalence problem the caveats raise.
- Structured question types matched to purpose. Across open_ended, scale, single_choice, multiple_choice, ranking, and yes_no formats, designers can choose the lightest format that captures a construct and rely on conversational follow-up where a number alone would be ambiguous.
The claim is deliberately modest: Koji is designed to mitigate response-style contamination by anchoring scales and triangulating every rating with open-text reasoning — not to eliminate bias, which no instrument can. The same AI-moderated interview engine underpins product and customer research on the core platform at koji.so, where rating-scale response styles are an equally live concern.
Related Resources
- How Many Scale Points Should a Course-Evaluation Question Have?
- Acquiescence Bias and Reverse-Worded Items: Should Course Evaluations Flip the Question?
- Response Styles and Likert Scales: Why Cross-Cultural Evaluation Needs More Than Numbers
- Why Students Click Straight Down the Middle: Satisficing in Course Evaluations
- Question Order and Context Effects in Course-Evaluation Surveys
- Should Course Evaluations Offer a "Don't Know" or "Not Applicable" Option?
References
- Weijters, B., Cabooter, E., & Schillewaert, N. (2010). The effect of rating scale format on response styles: The number of response categories and response category labels. International Journal of Research in Marketing, 27(3), 236–247. https://doi.org/10.1016/j.ijresmar.2010.02.004
- Weng, L.-J. (2004). Impact of the Number of Response Categories and Anchor Labels on Coefficient Alpha and Test-Retest Reliability. Educational and Psychological Measurement, 64(6), 956–972. https://doi.org/10.1177/0013164404268674
- Krosnick, J. A., & Presser, S. (2010). Question and Questionnaire Design. In P. V. Marsden & J. D. Wright (Eds.), Handbook of Survey Research (2nd ed., pp. 263–313). Emerald.
Related articles
Response Styles and Likert Scales: Why Cross-Cultural Evaluation Needs More Than Numbers
Acquiescence and extreme response styles vary systematically by culture (Harzing, 2006; Baumgartner & Steenkamp, 2001), which means raw Likert averages are not directly comparable across nationalities in Europe's multinational classrooms. This article explains the evidence and how to evaluate fairly across diverse cohorts.
How Many Scale Points Should a Course-Evaluation Question Have?
What the measurement literature — Preston & Colman (2000), Weng (2004), Dawes (2008) — says about the optimal number of response categories on rating scales, and why the answer for course evaluation is not just a number but a question about what a Likert item can and cannot capture.
Why Students Click Straight Down the Middle: Satisficing in Course Evaluations
A large share of students do not answer evaluation questions carefully — they straightline, speed through, and pick the easy option. Barge and Gehlbach (2012) showed this "satisficing" doesn't just add noise; it inflates your reliability and validity statistics, making bad data look good.
Question Order and Context Effects: How the Sequence of Items Shapes Course-Evaluation Answers
The order in which you ask evaluation questions changes the answers you get. Drawing on Schwarz (1999), Strack, Martin & Schwarz (1988) and Tourangeau, Rips & Rasinski (2000), we explain part-whole and assimilation/contrast effects, what they do to your data, and how conversational evaluation reduces the damage.