New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Methodology10 min read

Five, Seven, or Ten Points? How Your Rating Scale Silently Shapes Course-Evaluation Results

The number of points on your scale, whether it has a midpoint, and whether you label every option are not cosmetic choices. They change the reliability of your data, the means you report, and whether two courses can be compared at all. Here is what the psychometric evidence actually says.

Koji for Education

Editorial Team · June 30, 2026

Answer up front: The design of your rating scale — how many points it has, whether there is a neutral midpoint, and whether every point is labelled — measurably affects the quality and the comparability of course-evaluation data. The evidence converges on a practical sweet spot: scales with five to seven points have better reliability and validity than two-to-four-point scales, with little to gain beyond seven and some loss of stability past ten. But the more consequential lesson is about consistency: a 5-point course score and a 7-point one are not interchangeable, and rescaling them to a common range does not make them so. If your institution mixes scale formats across faculties, your benchmarking is built on sand.

Why the number of points is a measurement question

A rating scale is a measuring instrument, and like any instrument it has a resolution. Too few points and you throw away real distinctions — a student who thinks the course was "quite good" and one who thinks it was "excellent" are forced into the same box. Too many points and you exceed students' ability to discriminate meaningfully, adding noise rather than signal as people make essentially arbitrary choices between "7" and "8."

The most cited empirical work here is Preston and Colman's study on the optimal number of response categories. Across reliability, validity, and discriminating power, they found that two-, three-, and four-point scales performed worse than five-, six-, and seven-point scales; that indices improved up to around ten points and then began to decline; and that the best overall psychometric characteristics belonged to the seven-point scale. Respondents themselves preferred the 5-, 7-, and 10-point versions for ease of use. Lozano, García-Cueto and Muñiz (2008) reached a compatible conclusion from simulation: reliability and factorial validity climb as you add categories, with the optimal range between four and seven, and minimal gains thereafter.

So the familiar 5-point Likert that dominates course evaluation is defensible — but it is at the lower edge of the good range. A 7-point scale typically buys you a little more reliability and discrimination, at the cost of slightly higher cognitive load. Anything below five points is actively costing you information, which is one reason crude "thumbs up / thumbs down" or 3-point "below/met/exceeded expectations" forms are weaker than they look.

The means are not comparable across formats

Here is the finding that should worry anyone running cross-faculty comparisons. Dawes (2008) showed that 5- and 7-point scales produce similar means once rescaled, but a 10-point scale produces a systematically lower rescaled mean than either. In other words, the same underlying attitude expressed on a 10-point scale does not map linearly onto its 5-point equivalent. A department evaluating on 1–10 and reporting a "7.8" is not reporting the same thing as a department on 1–5 reporting a rescaled "3.9," even though the arithmetic says they match.

This is the scale-design version of a problem we have written about repeatedly: that you cannot compare a 4.1 in engineering to a 4.4 in history when the conditions differ. Format is one of those conditions. It also compounds the ceiling effects that already crush course-evaluation distributions into the top of the range: a shorter scale hits its ceiling faster, so a 5-point form manufactures more 4s-and-5s pile-up than a 10-point form would, independent of any difference in teaching.

Midpoints, labels, and the smaller design decisions

Beyond length, two other choices carry weight:

The midpoint. An odd-numbered scale offers a neutral middle ("neither agree nor disagree"); an even-numbered scale forces a lean. Removing the midpoint can reduce central-tendency clustering, where ambivalent or disengaged students park themselves in the safe middle. But forcing a choice also fabricates an opinion from students who genuinely have none, converting honest neutrality into noise. There is no free lunch: the midpoint trades one bias (central tendency) against another (forced choice). The right call depends on whether genuine neutrality is informative in your context.

Verbal labels. Scales where every point carries a verbal anchor ("poor / fair / good / very good / excellent") tend to be more reliable than scales that label only the endpoints and leave the middle as bare numbers, because the labels give every respondent the same interpretation of each point. Unlabelled interior points invite each student to invent their own spacing — is the gap from 6 to 7 the same as 2 to 3? — which injects exactly the kind of construct-irrelevant variance that erodes validity. Fully labelled scales also help international students, for whom an unlabelled number may carry different connotations than the designer assumed.

All of this sits on top of the more fundamental objection we have made elsewhere — that averaging Likert responses at all treats ordinal categories as if they were equal-interval numbers. Scale design does not solve that; it just determines how badly the ordinal-as-interval shortcut misleads you.

"But this is just fiddling at the margins — teaching quality swamps scale format"

This is the strongest objection, and it is partly true. The variance in course-evaluation scores attributable to scale format is smaller than the variance attributable to who is teaching, what the subject is, and how students experienced the course. If you only ever used one scale, for one purpose, internally, you could reasonably ignore most of this.

But two things rescue the topic from triviality. First, almost nobody uses scale format that cleanly. Real institutions have layered EvaSys forms, faculty-bespoke surveys, and legacy paper instruments with different lengths and labels, and then aggregate them into institution-level reports as if the numbers were commensurable. The format effect is small per item but systematic, and systematic small effects do not wash out when you average — they bias the comparison in a fixed direction. Second, the differences institutions act on are themselves small. When a promotion committee treats a 0.3-point gap as meaningful (a gap that is often within measurement error to begin with), a format-induced shift of similar size is decision-relevant, not academic.

The honest framing, then, is not "scale design is the biggest problem in course evaluation" — it is not. It is "scale design is a controllable source of non-comparability that most institutions leave uncontrolled, and that interacts with every downstream comparison they make."

Where a conversational approach fits

A response scale exists to compress a complex judgement into a single number that can be aggregated. Every design choice above is an attempt to manage the information you lose in that compression — but you are still compressing. The alternative is not to abandon scales (they remain useful for tracking trends) but to stop relying on them as the only instrument.

Koji for Education treats the scale as one of six structured question types — open_ended, scale, single_choice, multiple_choice, ranking, and yes_no — rather than the whole evaluation. When a scale rating is ambiguous, the AI moderator can probe it conversationally: a student who selects "3 out of 5" can be asked what would have made it a 5, turning a flat midpoint into a specific, actionable reason. That probing is standardised across every interview by the same AI, so the interpretation of the scale does not drift the way it does across human administrators or self-invented interior points. And because Koji applies automatic thematic analysis to the open text, the judgement is reconstructed from what students mean, not only from where they tapped on a line.

Koji is precise about the claim: better instrumentation reduces the information lost to scale compression and mitigates format-driven non-comparability; it does not eliminate the need for sensible scale design where scales are used. If you must use a scale, the evidence says use five to seven fully labelled points, decide the midpoint deliberately, and — above all — use the same scale everywhere you intend to compare. The same multi-format, AI-moderated interview engine powers the main Koji platform for teams doing customer and product research, where the limits of a single rating number are just as real.

Practical recommendations

  • Use 5 to 7 points. Below five loses information; above seven adds little; past ten, stability declines.
  • Label every point, not just the ends, so all students interpret the scale the same way.
  • Decide the midpoint on purpose. Keep it if genuine neutrality is informative; drop it if central-tendency clustering is your bigger problem — but know you are trading one bias for another.
  • Standardise the format across every course you benchmark. Mixing 5- and 10-point scales makes your league tables meaningless before any other bias enters.
  • Do not rescale and pretend. A rescaled 10-point mean is not equivalent to a 5-point one; Dawes's evidence says the 10-point version runs lower.
  • Pair the number with a reason. A scale tells you where; only the open text tells you why — and the why is what lets you act.

A rating scale feels like the most neutral, objective part of a course evaluation. It is one of the most quietly consequential design decisions you make.


Koji for Education combines structured scales with AI-moderated probing and thematic analysis, so a rating always comes with its reason. See how it works.