Do the Numbers on Your Rating Scale Change the Score? The Evidence on Numeric Anchors
Schwarz and colleagues showed that the numeric values printed on a rating scale (0 to 10 vs minus 5 to plus 5) systematically shift responses even when the verbal labels are identical. Here is what that means for course-evaluation design, comparability, and reporting.
Koji Education Team
Product
Answer (BLUF): The numbers printed on a rating scale are not neutral packaging — they are part of the question. Schwarz, Knäuper, Hippler, Noelle-Neumann and Clark (1991) showed that an 11-point scale running −5 to +5 yields systematically higher ratings than a verbally identical scale running 0 to 10, because negative numbers are read as "the opposite of the quality," not merely "less of it." For course evaluation this means two units using the same words but different numeric anchors are not comparable, and a scale's numbering can quietly inflate or deflate an instructor's mean. Use unipolar numbering (0–10 or 1–5), keep it byte-for-byte identical across every instrument and cohort, and report the distribution rather than treating the mean as an absolute truth.
What the research says
In a now-classic split-ballot experiment, Norbert Schwarz and colleagues (1991) asked a representative sample of German adults how successful they had been in life, using an eleven-point rating scale whose endpoints were labelled not at all successful and extremely successful. The wording never changed. The only thing the experimenters manipulated was the numbers attached to the eleven response boxes: one version ran from 0 to 10, the other from −5 to +5. Verbally, the third box from the bottom means exactly the same thing in both versions. Numerically it is labelled "2" in one and "−3" in the other.
The resulting distributions diverged sharply. On the 0–10 scale, 34% of respondents placed themselves in the lower half of the scale (the values 0 through 5). On the −5 to +5 scale, only 13% placed themselves in the formally equivalent lower half (−5 through 0). Changing the numbers — and nothing else — nearly tripled the proportion of people who appeared to rate themselves below the midpoint.
Schwarz's explanation is that respondents treat the numbers as a source of information about what the question means. A 0-to-10 range implies a unipolar continuum: zero signals the absence of the quality and higher numbers signal more of it. A −5-to-+5 range implies a bipolar continuum: the negative numbers are read as the presence of the opposite (here, failure), and 0 becomes a neutral middle. Because very few people will describe themselves as actively failing at life, bipolar numbering pushes the whole distribution upward. The labels promise equivalence; the digits deliver a different question.
A third experiment in the same paper extended the finding from responding to interpreting: recipients of a respondent's report drew more negative inferences from a numerically negative value than from a verbally equivalent positive one. The number colours meaning at both ends of the communication.
Two further sources show this is not a one-off curiosity. Schwarz and Hippler (1995, International Journal of Public Opinion Research) replicated the manipulation and compared its impact across mail surveys and telephone interviews, confirming the effect is robust to mode. Most directly for our audience, Armitage and Deeprose (2004), in a paper titled Changing Student Evaluations by Means of the Numeric Values of Rating Scales, carried the manipulation into higher education and reported that the numeric values attached to an otherwise identical student-evaluation scale changed the evaluations students gave. The broader lesson of Schwarz's research programme is blunt: a rating scale is not a neutral measuring stick but an active part of the communication between question-writer and respondent.
Why it matters for course evaluation in practice
Most quality-assurance teams obsess over the wording of evaluation items and never look twice at the numbering. The evidence says the numbering deserves equal scrutiny.
- Cross-unit comparability collapses if numbering varies. If the engineering faculty uses a 1–5 scale and the medical school uses a −2 to +2 "bipolar" agreement scale, their means are measuring subtly different constructs. Benchmarking one against the other — or against an institutional average — imports the artefact directly into personnel and accreditation conversations.
- Bipolar agree/disagree numbering inflates apparent positivity. Scales numbered around a zero midpoint (e.g. −3 to +3 "strongly disagree" to "strongly agree") will tend to produce more favourable distributions than the same labels numbered 1 to 7, simply because students avoid the "negative" region. An instructor's score can move without their teaching moving.
- Small mean differences are already over-interpreted. As covered in our guidance on interpreting and reporting student ratings responsibly, committees routinely treat a 0.2-point gap as meaningful. If part of that gap is a numbering artefact rather than a teaching signal, the over-interpretation becomes actively unfair.
- Year-on-year trends break when instruments are "modernised." Redesigning the evaluation form and quietly renumbering the scale will produce a discontinuity that looks like a change in teaching quality but is pure measurement.
The practical takeaway is conservative and cheap: pick unipolar numbering (0–10 or 1–5), label the endpoints clearly, hold the numeric anchors constant across every department and every cohort, and never renumber a scale without flagging the break in your time series.
Limitations & honest caveats
A careful reader should not over-extrapolate from a single 1991 life-satisfaction study to every teaching-evaluation item.
- Construct distance. Schwarz's flagship demonstration used a self-evaluation of life success, a highly bipolar, ego-involving construct where "the opposite" (failure) is psychologically vivid. Many course-evaluation items (e.g. "the lecturer explained concepts clearly") are more naturally unipolar, and the numeric effect may be smaller for them.
- Effect size depends on design. The dramatic 34% vs 13% contrast is specific to the 0–10 versus −5/+5 comparison. A choice between, say, 1–5 and 0–4 unipolar numbering is unlikely to produce anything like that magnitude. The strong warning is really about unipolar versus bipolar numbering, not about every cosmetic digit choice.
- Replication is mixed across formats. A wider methodological literature (e.g. Dodou and de Winter, 2014, on social-desirability and format effects) finds that some feared scale-format effects are smaller than early studies implied. The honest position is that numbering can bias responses, especially when it changes the unipolar/bipolar reading, not that every numeric tweak materially moves scores.
- Ordinal-data caveat compounds the problem. Because evaluation responses are ordinal, comparing means across differently numbered scales is doubly questionable — see our note on why averaging Likert scores can mislead. The numbering effect is one more reason to lead with distributions.
None of these caveats rescues the practice of mixing numbering schemes across an institution. They simply tell you where the effect is large (unipolar vs bipolar) and where it is probably modest (cosmetic digit changes within a unipolar scale).
How Koji incorporates this
Koji is designed to remove numeric-anchor artefacts as a source of noise rather than to pretend they do not exist.
- Controlled, consistent scale construction. Koji's structured
scalequestion type uses unipolar numbering with clearly labelled endpoints by default, and the same anchors are reused across cohorts and units. This is a deliberate design choice to keep the construct stable so that a change in the number reflects a change in teaching, not a change in the form. - The number is never the whole measurement. Because Koji runs AI-moderated conversational interviews, a
scalerating is routinely followed by an open, probing follow-up ("you said 3 out of 5 on clarity — what specifically made it a 3?"). The qualitative answer is analysed thematically and is far less sensitive to whether the box was labelled 3 or −1. This triangulation is precisely the mitigation Schwarz's work implies: do not let a numeral carry the full weight of the judgment. - Bias-aware reporting. Koji reports full response distributions and the underlying comments, not just a mean, so a committee can see whether a score is genuinely low or merely sitting in the lower numeric region of a particular scale design. This is designed to mitigate — not eliminate — the over-interpretation of small mean differences.
- Instrument versioning. When an evaluation instrument changes, Koji preserves the prior version so that year-on-year comparisons flag a methodological break rather than silently reporting a renumbering as a quality shift.
The same AI-moderated interview engine powers Koji's core research platform at koji.so for product and customer research, where the identical principle applies: a Likert number is a starting point for a conversation, not the conclusion of one.
Related Resources
- How Many Scale Points Should a Course-Evaluation Question Have?
- Should Every Point on a Course-Evaluation Scale Be Labelled?
- Agree/Disagree or Item-Specific? The Scale Choice That Quietly Degrades Course-Evaluation Data
- Sliders, Visual-Analogue, or Radio Buttons? The Evidence on Response Formats
- Response Styles and Likert Scales: Why Cross-Cultural Evaluation Needs More Than Numbers
- Interpreting and Reporting Student Ratings Responsibly
References
- Schwarz, N., Knäuper, B., Hippler, H.-J., Noelle-Neumann, E., & Clark, L. (1991). Rating Scales: Numeric Values May Change the Meaning of Scale Labels. Public Opinion Quarterly, 55(4), 570–582. https://doi.org/10.1086/269282
- Schwarz, N., & Hippler, H.-J. (1995). The Numeric Values of Rating Scales: A Comparison of Their Impact in Mail Surveys and Telephone Interviews. International Journal of Public Opinion Research, 7(1), 72–74. https://doi.org/10.1093/ijpor/7.1.72
- Armitage, C. J., & Deeprose, C. (2004). Changing Student Evaluations by Means of the Numeric Values of Rating Scales. Psychology Learning & Teaching, 3(2), 122–127. https://doi.org/10.2304/plat.2003.3.2.122
- Dodou, D., & de Winter, J. C. F. (2014). Social desirability is the same in offline, online, and paper surveys: A meta-analysis. Computers in Human Behavior, 36, 487–495. https://doi.org/10.1016/j.chb.2014.04.005
Related articles
Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees
The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.
Response Styles and Likert Scales: Why Cross-Cultural Evaluation Needs More Than Numbers
Acquiescence and extreme response styles vary systematically by culture (Harzing, 2006; Baumgartner & Steenkamp, 2001), which means raw Likert averages are not directly comparable across nationalities in Europe's multinational classrooms. This article explains the evidence and how to evaluate fairly across diverse cohorts.
How Many Scale Points Should a Course-Evaluation Question Have?
What the measurement literature — Preston & Colman (2000), Weng (2004), Dawes (2008) — says about the optimal number of response categories on rating scales, and why the answer for course evaluation is not just a number but a question about what a Likert item can and cannot capture.
Should Every Point on a Course-Evaluation Scale Be Labelled? The Evidence on Verbal Anchors
Labelling every response option, not just the endpoints, tends to raise reliability and tame response styles — but it also interacts with the number of categories and the language of your students. What the rating-scale-format research says for course evaluation.