Agree/Disagree or Item-Specific? The Scale Choice That Quietly Degrades Course-Evaluation Data
"The lecturer was well organised: Strongly disagree to Strongly agree" feels natural, but the agree/disagree format invites acquiescence and lower data quality. Saris et al. (2010) on why item-specific scales measure better.
Koji Education Team
Product
In short: The familiar "Strongly disagree → Strongly agree" (agree/disagree, or A/D) format is convenient and easy to write, but the survey-methodology evidence says it produces lower-quality data than item-specific scales that label the construct directly (for example, "How clear was the lecturer's explanation? Not at all clear → Extremely clear"). Saris, Revilla, Krosnick and Shaeffer (2010) showed empirically that A/D options diminish item quality, largely because they invite acquiescence — the tendency to agree regardless of content. For course evaluations, where a half-point shift can change a personnel or programme decision, converting A/D items to construct-specific scales is one of the cheapest validity upgrades available.
What the research says
The anchor study is Willem Saris, Melanie Revilla, Jon Krosnick and Eric Shaeffer's "Comparing Questions with Agree/Disagree Response Options to Questions with Item-Specific Response Options" (Survey Research Methods, 2010, 4(1), 61–79). Agree/disagree batteries are ubiquitous in social science because they are quick to assemble: write any statement, bolt on the same five-point agree scale, repeat. The authors asked whether that convenience comes at a measurement cost.
Using a multitrait-multimethod (MTMM) design — the gold-standard approach for estimating the reliability and validity of survey items by comparing multiple traits measured by multiple methods — they compared A/D items against item-specific (IS) items that name the dimension and label the scale accordingly. The empirical result was clear: agree/disagree response options diminish item quality. IS questions yielded higher-quality measurement of the same underlying attitudes. A central mechanism is acquiescence: a content-independent bias toward agreeing, which inflates correlations among A/D items and contaminates the signal.
Two corroborating sources reinforce and bound the finding:
- Revilla, Saris and Krosnick (2014), "Choosing the Number of Categories in Agree–Disagree Scales" (Sociological Methods & Research, 43(1), 73–97), extends the programme to scale length, finding that even within A/D formats the number of categories interacts with quality — and continuing to recommend item-specific construction where feasible.
- The broader acquiescence literature — including the long-standing finding that agreement bias is stronger among respondents with lower education or motivation — explains why the A/D format underperforms, and connects directly to our existing note on acquiescence bias and reverse-worded items. The traditional "fix" of inserting reverse-worded A/D items to cancel acquiescence introduces its own problems (confused respondents, lowered reliability), which is part of why item-specific scales are the cleaner solution.
The convergent message: how you label the response scale is not cosmetic — it changes the quality of the number you get back.
Why it matters for course evaluation in practice
Most institutional SET instruments are built almost entirely from A/D statements: "The teaching was well organised," "The assessment was fair," "I learned a lot," each rated Strongly disagree → Strongly agree. This design is attractive because one scale fits every item, the report looks uniform, and trends are easy to compute. But it carries three practical costs:
- Acquiescence inflation. A meaningful share of students lean toward agreement regardless of the statement. That systematically pushes scores upward and, worse, correlates the items artificially — making a multidimensional instrument look more unidimensional than the underlying experience is (a problem that then distorts factor analyses and Cronbach's alpha).
- Lower validity per item. Saris et al.'s MTMM evidence means each A/D item carries more method variance and less trait variance — you are measuring the format as much as the teaching.
- Harder cross-group comparison. Because acquiescence varies by group (education, language, culture — see response styles across cultures), A/D items make comparisons between, say, domestic and international cohorts less trustworthy.
The remedy is concrete: rewrite items so the scale names the construct. Instead of "The lecturer explained concepts clearly (disagree–agree)," ask "How clearly did the lecturer explain concepts? (Not at all clearly → Extremely clearly)." Instead of "The workload was appropriate (disagree–agree)," ask "How was the workload? (Far too light → Far too heavy)." The question now measures the dimension directly, with no agreement to acquiesce to.
Limitations & honest caveats
Several reservations keep this honest:
- Item-specific scales are harder to write and to standardise. Every dimension needs its own carefully chosen endpoints, which is more design work than reusing one agree scale, and a badly chosen anchor can reintroduce ambiguity.
- The effect sizes, while consistent, are not enormous. Converting to IS scales improves quality at the margin; it will not rescue a survey crippled by a 10% response rate or end-of-exam timing. It is a refinement, not a cure-all.
- MTMM estimates are model-dependent. The quality coefficients in Saris et al. rest on assumptions about the multitrait-multimethod model; different specifications can shift the numbers, though the direction of the A/D penalty is robust across the programme.
- Trend disruption. Switching an established instrument from A/D to IS breaks comparability with historical data — a real cost for institutions trending scores over many years, which must be managed with a bridging period.
- Generalisability. Much of this evidence comes from general social-attitude surveys, not course evaluations specifically; the mechanism (acquiescence) transfers cleanly, but published course-evaluation MTMM studies remain comparatively scarce.
How Koji incorporates this
Koji is designed to reduce acquiescence and method bias rather than bake them in:
- Construct-specific scale support by default. Koji's
scalequestion type lets you set item-specific endpoints (Not at all clear → Extremely clear; Far too light → Far too heavy) instead of forcing every item onto one agree/disagree track — directly applying the Saris et al. recommendation. - Less Likert dependence overall. The AI-moderated conversational interview asks open, construct-specific questions and probes the answer, so the most important judgments are not collected as agreement with a pre-written statement at all. There is nothing to acquiesce to when the question is "Walk me through how the assessment was explained to you."
- Acquiescence-aware analysis. Because Koji captures both structured scales and open text on the same construct, it can triangulate: if a student "agrees" with everything on the scale but their open-text answers are lukewarm or critical, that discrepancy surfaces in the analysis rather than being averaged into a falsely positive mean (a form of bias-aware reporting).
- Quality scoring and straightlining detection catch the non-differentiating responding that acquiescence and satisficing produce (see satisficing), so low-quality agree-everything responses can be flagged.
Koji frames these as designed to mitigate acquiescence, never to eliminate response bias entirely. The same conversational engine powers product and customer research on Koji's core platform at koji.so, where agree/disagree brand-statement batteries create the very same quality penalty.
Related Resources
- Acquiescence bias and reverse-worded items
- Should every point on a course-evaluation scale be labelled?
- Response styles and Likert scales across cultures
- Cronbach's alpha and course-evaluation reliability
- How many scale points should a question have?
- Why students click straight down the middle: satisficing
Converting your instrument: before and after
The cheapest way to apply Saris et al. (2010) is to rewrite your existing agree/disagree statements into item-specific scales. The pattern is mechanical once you see it: take the statement, find the dimension it is really about, and let the scale endpoints name that dimension directly so there is no proposition to agree or disagree with.
| Agree/disagree (A/D) item | Item-specific (IS) rewrite |
|---|---|
| "The lecturer explained concepts clearly." (Strongly disagree → Strongly agree) | "How clearly did the lecturer explain concepts?" (Not at all clearly → Extremely clearly) |
| "The assessment was fair." (Strongly disagree → Strongly agree) | "How fair was the assessment?" (Not at all fair → Completely fair) |
| "The workload was appropriate." (Strongly disagree → Strongly agree) | "How was the workload for this course?" (Far too light → Far too heavy) |
| "I received useful feedback." (Strongly disagree → Strongly agree) | "How useful was the feedback you received?" (Not at all useful → Extremely useful) |
Notice the workload example: the IS version also fixes a hidden flaw in the A/D form. "Appropriate" is evaluative and one-directional, so an A/D scale cannot distinguish too light from too heavy — both register as "disagree". The bipolar IS scale recovers that information, telling you which way to adjust.
Three guardrails when you convert:
- Choose endpoints carefully. A vague anchor ("Poor → Good") can reintroduce the ambiguity you were trying to remove. Name the construct concretely.
- Bridge the trend. Switching format breaks comparability with historical A/D data. Run both versions for one cycle if the trend matters, or annotate the break clearly in your reporting.
- Don't convert everything at once. Prioritise the items that feed real decisions — the ones where acquiescence inflation could tip a personnel or programme judgement — and leave low-stakes items for a later revision.
Done this way, the conversion is a low-risk, high-return validity upgrade: each item now measures its dimension rather than a student's general willingness to agree.
References
- Saris, W. E., Revilla, M., Krosnick, J. A., & Shaeffer, E. M. (2010). Comparing questions with agree/disagree response options to questions with item-specific response options. Survey Research Methods, 4(1), 61–79. https://doi.org/10.18148/srm/2010.v4i1.2682
- Revilla, M. A., Saris, W. E., & Krosnick, J. A. (2014). Choosing the number of categories in agree–disagree scales. Sociological Methods & Research, 43(1), 73–97. https://doi.org/10.1177/0049124113509605
- Krosnick, J. A. (1991). Response strategies for coping with the cognitive demands of attitude measures in surveys. Applied Cognitive Psychology, 5(3), 213–236. https://doi.org/10.1002/acp.2350050305
Related articles
Cronbach's Alpha and Course-Evaluation Reliability: What It Does and Doesn't Tell You
A high Cronbach's alpha on your course-evaluation instrument is widely read as proof of a "reliable" survey. The psychometric literature says that reading is largely wrong. What alpha actually measures, where it misleads, and what to report instead.
Why Students Click Straight Down the Middle: Satisficing in Course Evaluations
A large share of students do not answer evaluation questions carefully — they straightline, speed through, and pick the easy option. Barge and Gehlbach (2012) showed this "satisficing" doesn't just add noise; it inflates your reliability and validity statistics, making bad data look good.
Response Styles and Likert Scales: Why Cross-Cultural Evaluation Needs More Than Numbers
Acquiescence and extreme response styles vary systematically by culture (Harzing, 2006; Baumgartner & Steenkamp, 2001), which means raw Likert averages are not directly comparable across nationalities in Europe's multinational classrooms. This article explains the evidence and how to evaluate fairly across diverse cohorts.
How Many Scale Points Should a Course-Evaluation Question Have?
What the measurement literature — Preston & Colman (2000), Weng (2004), Dawes (2008) — says about the optimal number of response categories on rating scales, and why the answer for course evaluation is not just a number but a question about what a Likert item can and cannot capture.