How Many Scale Points Should a Course-Evaluation Question Have?
What the measurement literature — Preston & Colman (2000), Weng (2004), Dawes (2008) — says about the optimal number of response categories on rating scales, and why the answer for course evaluation is not just a number but a question about what a Likert item can and cannot capture.
Koji Education Team
Product
Quick answer
How many points should a course-evaluation rating scale have — 5, 7, or 10? The measurement evidence converges on a band rather than a single magic number: 2-, 3-, and 4-point scales perform poorly on reliability, validity, and discriminating power, while scales with roughly 7 to 10 categories maximise those properties, and reliability tends to fall off again beyond about 10–11 points. Preston and Colman (2000) found the most reliable scores came from 7–10 categories, with respondents themselves preferring 10-, 7-, and 9-point scales. But the deeper lesson for course evaluation is that no number of points fixes the core limitation of a Likert item — it still asks a person to compress a complex judgment into one ordinal position — so the better question is how to complement well-built scales with evidence scales cannot capture.
What the research says
The anchor is Carolyn Preston and Andrew Colman, "Optimal number of response categories in rating scales: Reliability, validity, discriminating power, and respondent preferences," Acta Psychologica (2000), 104(1), 1–15. Respondents rated service elements of a recently visited store or restaurant on scales that were identical except for the number of response categories, which ranged from 2 to 11 (plus a 101-point format). Because only the number of points varied, the design cleanly isolates the effect of scale length on data quality.
Their findings are now a standard reference. On multiple indices of reliability, validity, and discriminating power, the 2-, 3-, and 4-point scales performed relatively poorly, and these indices rose significantly as categories increased up to about seven. The most reliable scores came from scales with 7 to 10 response categories. Test–retest reliability tended to decrease for scales with more than about 10 categories, suggesting diminishing or negative returns once a scale becomes so fine-grained that respondents cannot use the distinctions consistently. Notably, respondent preferences did not match the very long 101-point format: people most preferred the 10-point scale, closely followed by 7- and 9-point scales. So the scale length that yields good data and the length people find comfortable overlap in the 7–10 range.
Two corroborating studies strengthen and qualify this. Li-Jen Weng (2004), "Impact of the number of response categories and anchor labels on coefficient alpha and test–retest reliability" (Educational and Psychological Measurement, 64(6), 956–972), studied scales with 3 to 9 categories and found that few categories produced lower reliability, especially lower test–retest reliability — and added a second, often-overlooked lever: scales with every response option clearly labelled (not just the endpoints) achieved higher test–retest reliability. In other words, how you label the points matters alongside how many there are. John Dawes (2008), "Do data characteristics change according to the number of scale points used?" (International Journal of Market Research, 50(1), 61–104), experimentally compared 5-, 7-, and 10-point formats and found that the number of points affects the distributional characteristics of the data (means, skew, and the comparability of scores rescaled across formats) — a practical warning that you cannot naively compare or pool results collected on different scale lengths.
The synthesis across these three: more points than four meaningfully improves reliability and discrimination up to about seven; gains flatten and can reverse beyond ten; fully labelling categories helps; and the chosen format shapes the data itself, so consistency and rescaling caution are required.
Why it matters for course evaluation in practice
Most institutional course-evaluation instruments default to a 5-point Likert scale, often out of habit. The literature suggests this is a defensible but not obviously optimal choice: a 5-point scale sits just below the reliability sweet spot, and moving to a 7-point fully labelled scale would, on this evidence, modestly improve the reliability and discriminating power of each item. "Discriminating power" matters acutely in course evaluation, where administrators routinely try to distinguish instructors whose mean scores differ by a tenth of a point — a distinction a coarse scale simply cannot support reliably.
Three practical implications follow. First, avoid very short scales (2–4 points) for anything used in decisions; they throw away discriminating power. Second, prefer fully labelled categories where feasible — Weng's result says labelling each point, not just the ends, buys reliability. Third, never silently change scale length and then compare across years: Dawes shows the data characteristics shift with format, so a move from 5 to 7 points breaks longitudinal comparability unless explicitly handled. These are low-cost design choices that materially affect whether a number means anything.
Limitations and honest caveats
The caveats are important and, for course evaluation, decisive. First, all three anchor studies are about general rating scales — service quality, attitudes — not course evaluation specifically; the optimal band is likely transferable, but the effect sizes were established in other domains. Second, "optimal" here means optimal for the psychometric properties of a Likert item — it presupposes that a Likert item is the right tool. It says nothing about the far larger threats to course-evaluation validity documented elsewhere: gender, attractiveness, accent, and grading-leniency biases are not cured by adding scale points. A perfectly calibrated 7-point scale measuring a biased construct just measures the bias more reliably. Third, more categories increase cognitive burden and can invite central-tendency or response-style effects, especially across cultures where response styles differ. Fourth, treating ordinal Likert responses as interval data for averaging is a separate, contested assumption that scale length does not resolve. The honest conclusion: scale length is a real, tunable lever with a known sweet spot (about 7, fully labelled), but it is a second-order fix that cannot rescue a fundamentally limited instrument.
How Koji incorporates this
Koji for Education treats good scale design as table stakes and then goes beyond the Likert item where the evidence says you must.
- Evidence-based scale construction. Koji's structured scale questions are configured in the reliability-supported band (around 7 fully labelled points rather than a bare 5-point or coarse 3-point scale), and support clearly labelled categories — directly applying Preston & Colman and Weng. Scale definitions are held stable so longitudinal comparisons are not silently broken, addressing Dawes's warning about cross-format comparability.
- The right question type for the construct. Koji offers open_ended, scale, single_choice, multiple_choice, ranking, and yes_no questions, so designers are not forced to bend every judgment into a Likert item. A ranking or single_choice question often captures a comparative judgment more honestly than a fine-grained scale.
- AI-moderated interviews that go where scales cannot. The literature's deepest caveat is that adding points cannot make a Likert item capture why. Koji's conversational moderator probes the reasoning behind a rating — turning a "4 out of 7 on clarity" into a specific account of what was unclear — so the instrument captures meaning, not just a more reliable point estimate.
- Automatic thematic analysis and triangulation. Open-text responses are clustered into themes and read alongside scale results, and Koji's reporting is designed to mitigate over-reading of small scale differences by presenting distributions and context rather than ranking on hairline mean gaps. Koji does not claim a scale format eliminates bias — only that combining a well-built scale with conversational evidence yields a more trustworthy picture.
The same AI-moderated interview engine powers Koji's core research platform at koji.so, where pairing well-designed scale items with probing follow-ups is the standard recipe for high-quality product and customer research.
A practical configuration for course-evaluation scales
Turning the measurement evidence into an instrument specification yields a small set of concrete, low-cost rules.
- Default to a 7-point, fully labelled scale for evaluative items. Seven points sit at the reliability and discriminating-power optimum identified by Preston and Colman, and labelling every category — not just the endpoints — adds the test–retest reliability gain Weng documented. This is strictly better than the habitual 5-point endpoints-only scale, at no extra respondent cost.
- Reserve very short scales for genuinely binary judgments. A yes_no or 2–3 point item is fine for a categorical fact ("Did you receive feedback on your assignments?") but should never carry an evaluative judgment that will be averaged and compared, where it discards discriminating power.
- Freeze the format across cycles. Because Dawes showed that means, skew, and comparability shift with scale length, the scale definition should be held stable year over year. If a change is unavoidable, flag the break explicitly and avoid presenting a continuous trend line across the discontinuity.
- Match the question type to the construct. Not every judgment belongs on a Likert scale. A comparative priority is better captured by a ranking item; a discrete choice by single_choice or multiple_choice. Forcing every question into a rating scale is a common source of low-information data.
- Treat the scale as the floor, not the ceiling. Even an optimally configured scale cannot tell you why a student rated as they did, and it cannot neutralise gender, attractiveness, or accent bias. Pair every consequential scale item with at least one open-text or conversational follow-up so the instrument captures reasons, not only a more reliable number.
Get these five choices right and each rating carries more signal; treat them as the whole solution and you simply measure a limited construct more precisely.
Related Resources
- /docs/why-averaging-likert-scores-misleads-course-evaluation
- /docs/response-styles-likert-cross-cultural-evaluation
- /docs/student-written-comments-course-evaluation
- /docs/response-rate-bias-course-evaluations
- /docs/student-evaluations-teaching-and-learning-meta-analysis
References
- Preston, C. C., & Colman, A. M. (2000). Optimal number of response categories in rating scales: Reliability, validity, discriminating power, and respondent preferences. Acta Psychologica, 104(1), 1–15. https://doi.org/10.1016/S0001-6918(99)00050-5
- Weng, L.-J. (2004). Impact of the number of response categories and anchor labels on coefficient alpha and test–retest reliability. Educational and Psychological Measurement, 64(6), 956–972. https://doi.org/10.1177/0013164404268674
- Dawes, J. (2008). Do data characteristics change according to the number of scale points used? An experiment using 5-point, 7-point and 10-point scales. International Journal of Market Research, 50(1), 61–104. https://doi.org/10.1177/147078530805000106
Related articles
Response Styles and Likert Scales: Why Cross-Cultural Evaluation Needs More Than Numbers
Acquiescence and extreme response styles vary systematically by culture (Harzing, 2006; Baumgartner & Steenkamp, 2001), which means raw Likert averages are not directly comparable across nationalities in Europe's multinational classrooms. This article explains the evidence and how to evaluate fairly across diverse cohorts.
What Can Open-Text Student Comments Tell You That Likert Scores Cannot?
A research-grounded guide to open-ended student comments in course evaluation: what Alhija and Fresko found about who writes them and what they contain, how thematic analysis surfaces issues numbers miss (Stupans et al.), how rare abusive comments actually are (Tucker), and how to turn free text into reliable evidence.
Do Student Evaluations Measure Learning? The Uttl Meta-Analysis Revisited
A research-grounded reading of the Uttl, White & Gonzalez (2017) meta-analysis on the SET–learning relationship, with implications for European course evaluation and quality-assurance policy.
Racial and Ethnic Bias in Student Evaluations of Teaching: What the Evidence Shows
Peer-reviewed evidence shows students rate instructors of colour systematically lower than White peers on the same teaching. We review Reid (2010), Chávez & Mitchell (2020) and Bavishi et al. (2010), the limitations, and how to design evaluation so race-linked bias does not contaminate quality decisions.