Acquiescence Bias and Reverse-Worded Items: Should Course Evaluations Flip the Question?
Reverse-worded items are the classic survey-design fix for yea-saying. But Weijters and Baumgartner show negated and reversed items create their own measurement problems. What the evidence means for course-evaluation questionnaire design.
Koji Education Team
Product
In short: Acquiescence — the tendency to agree with statements regardless of content — is a real and well-documented threat to Likert-based course evaluations, and the textbook remedy is to mix in reverse-worded ("negatively keyed") items so agreement no longer maps cleanly onto a positive view. But Weijters and Baumgartner's review shows that reversed and negated items introduce their own errors: lower reliability, spurious "method" factors, and confused respondents who misread the flip. The defensible conclusion is not "always reverse-word" or "never reverse-word", but: control acquiescence through balanced, carefully written items and validation, and lean on conversational and open-text methods that do not depend on agree/disagree at all.
What the research says
Acquiescence response bias (also called yea-saying) is the disposition to endorse the agreement pole of a Likert item independently of its substantive content. In a course evaluation where almost every statement is phrased positively — "The instructor explained concepts clearly", "Assessment was fair", "I would recommend this course" — an acquiescent respondent inflates every score, and a uniformly positive profile can reflect a response style rather than a genuinely excellent course.
The standard psychometric defence is to include reverse-worded items: statements phrased so that disagreement indicates the positive view (for example, "The instructor often left concepts unexplained"). If a respondent agrees with both "explained clearly" and "left concepts unexplained", their acquiescence is exposed and can be modelled out.
The anchor source is Bert Weijters and Hans Baumgartner's "Misresponse to Reversed and Negated Items in Surveys: A Review" (Journal of Marketing Research, 2012, 49(5), 737–747; DOI: 10.1509/jmr.11.0368). Reviewing a large body of evidence, they document that reversed and negated items frequently misfire. Respondents skim, miss the "not" or the negative framing, and answer as if the item were positively keyed. The consequences are measurable: reverse-keyed items load on a separate method factor in factor analyses (an artefact of wording, not of the construct), they tend to have lower item-total correlations, and they reduce the apparent reliability and clean dimensionality of the scale.
Weijters, Baumgartner and Schillewaert develop this into a formal account in "Reversed item bias: An integrative model" (Psychological Methods, 2013, 18(3), 320–334; DOI: 10.1037/a0032121), distinguishing careless responding (skimming past the reversal) from acquiescence proper and showing how both distort balanced scales. The classic demonstration that negatively worded items create artefactual factors is Herbert Marsh's work on self-esteem scales, where "negative self-esteem" emerged as a separate factor purely because of item wording (Marsh, 1996, Journal of Personality and Social Psychology, 70(4), 810–819).
On the modelling side, Jaak Billiet and McKee McClendon's "Modeling Acquiescence in Measurement Models for Two Balanced Sets of Items" (Structural Equation Modeling, 2000, 7(4), 608–628; DOI: 10.1207/S15328007SEM0704_5) shows that a balanced set of positive and negative items lets you estimate a separate acquiescence style factor — but only if the balance and the model are set up correctly. The remedy works, in other words, only when executed with care.
Why it matters for course evaluation in practice
Three implications follow for a quality-assurance office that runs Likert questionnaires.
First, all-positive item sets are vulnerable to inflation. If every statement points the same way, an acquiescent or satisficing respondent (see satisficing and straightlining) produces a uniformly high profile that is indistinguishable from a genuinely excellent course. Mean scores drift upward and ceiling effects set in, which undermines the comparisons accreditation panels want to make.
Second, naively "fixing" this by sprinkling in reverse-worded items can make data worse, not better. A reversed item that half your cohort misreads adds noise, fractures your factor structure, and produces a counter-intuitive correlation that a committee may over-interpret. The Weijters–Baumgartner review is a direct warning against the cargo-cult use of reverse wording as a box-ticking exercise.
Third, the cleanest signal often comes from outside the agree/disagree frame entirely. Acquiescence is a property of the Likert agreement scale. Open-ended prompts, behaviourally anchored questions ("How often did you receive feedback you could act on?"), and conversational follow-ups do not offer an "agree" pole to lean on, so they sidestep the bias rather than trying to model it after the fact. This is the same logic behind preferring open-text comments to bare numbers.
Limitations and honest caveats
A careful reader should resist over-correcting in the other direction.
Reverse wording is not always harmful. When items are written clearly, piloted, and the resulting data modelled with an explicit acquiescence factor (as Billiet and McClendon show), balanced scales can control yea-saying effectively. The problem Weijters and Baumgartner identify is poor execution, not the principle. Abandoning reverse items entirely throws away a legitimate tool.
Acquiescence is hard to separate from real positivity. A student body may rate a course uniformly high because the course is genuinely excellent. Distinguishing true positivity from response style requires either balanced items, anchored behavioural questions, or triangulation with other evidence — you cannot diagnose acquiescence from a single positive mean.
Cultural variation. Acquiescence levels differ systematically across cultures and languages, which matters for cross-border European institutions and exchange cohorts; see our note on cross-cultural response styles. A design tuned to one population may behave differently in another.
Most evidence is from general survey and consumer research, not specifically from course evaluation. The mechanisms (careless reading, agreement tendency) are general enough to transfer, but the precise magnitudes in a student-evaluation setting are less well quantified, and we should say so rather than imply a settled effect size.
How Koji incorporates this
Koji for Education is designed to reduce dependence on the very agree/disagree mechanics that make acquiescence a problem — framed as mitigation, not a cure-all.
- AI-moderated conversational interviews replace agree/disagree with description. Instead of "Do you agree the instructor explained clearly?", Koji asks an open question and follows up: Walk me through a time the explanation worked — or did not. There is no agreement pole to default to, so yea-saying has far less purchase. The moderator can also gently probe a suspiciously uniform set of answers in real time.
- Mixed structured question types reduce single-format bias. Using
open_ended,scale,single_choice,multiple_choice,ranking, andyes_no, an evaluation can ask students to rank what helped them learn or choose among concrete options rather than agreeing with a battery of positive statements. Ranking and forced-choice formats are structurally resistant to acquiescence because agreement is not an available shortcut. - Behaviourally anchored scale items. Where a
scalequestion is used, Koji encourages frequency- and behaviour-anchored wording ("How often could you act on the feedback you received?") rather than attitude statements, which lowers the acquiescence surface. - Quality scoring and bias-aware reporting flag response patterns consistent with acquiescence or careless responding — uniformly high answers with thin or contradictory open-text — so a QA committee does not read a response-style artefact as a quality verdict. This is the practical equivalent of modelling out an acquiescence factor, but surfaced transparently.
- Automatic thematic analysis of open text gives a content-grounded counterweight to any Likert summary, so conclusions rest on what students actually described.
Koji does not claim to eliminate acquiescence — no instrument can. The design goal is to stop relying on a single agreement-keyed scale as the system of record. Koji's core research platform at koji.so applies the same conversational engine to customer and product research, where acquiescence and yea-saying are equally well-known threats to satisfaction surveys.
A practical checklist for questionnaire design
If you run Likert-based evaluations and want to limit acquiescence without importing the problems of careless reverse wording, the evidence points to a few disciplined moves.
- Do not rely on an all-positive item battery. A questionnaire where every statement points the same way is the easiest target for yea-saying and produces ceiling-bound means that resist comparison.
- If you use reverse-keyed items, pilot them. Check that students read the flip correctly, keep the wording simple and concrete, avoid double negatives, and inspect the factor structure for an artefactual method factor before trusting the results.
- Prefer behaviourally anchored questions to attitude statements. "How often did you receive feedback you could act on?" gives a frequency answer with no agreement pole to lean on, which structurally lowers the acquiescence surface.
- Add ranking and forced-choice formats. Asking students to rank what most helped their learning forces discrimination and resists the uniform-agreement shortcut entirely.
- Cross-check uniformly positive profiles against open text. A high mean with thin or contradictory comments is a signal to interpret the numbers cautiously rather than at face value, and to triangulate before drawing conclusions.
The underlying principle is consistent with the rest of this knowledge base: the more an evaluation depends on a single agreement-keyed scale, the more vulnerable it is to response style rather than substance.
Related Resources
- Response Styles and Likert Scales: Cross-Cultural Evaluation
- Satisficing and Straightlining in Course Evaluations
- How Many Scale Points Should a Question Have?
- Question Order and Context Effects
- What Open-Text Comments Tell You That Likert Scores Cannot
- What Do Student Evaluations Actually Measure?
References
- Weijters, B., & Baumgartner, H. (2012). Misresponse to reversed and negated items in surveys: A review. Journal of Marketing Research, 49(5), 737–747. https://doi.org/10.1509/jmr.11.0368
- Weijters, B., Baumgartner, H., & Schillewaert, N. (2013). Reversed item bias: An integrative model. Psychological Methods, 18(3), 320–334. https://doi.org/10.1037/a0032121
- Billiet, J. B., & McClendon, M. J. (2000). Modeling acquiescence in measurement models for two balanced sets of items. Structural Equation Modeling, 7(4), 608–628. https://doi.org/10.1207/S15328007SEM0704_5
- Marsh, H. W. (1996). Positive and negative global self-esteem: A substantively meaningful distinction or artifactors? Journal of Personality and Social Psychology, 70(4), 810–819. https://doi.org/10.1037/0022-3514.70.4.810
Related articles
What Can Open-Text Student Comments Tell You That Likert Scores Cannot?
A research-grounded guide to open-ended student comments in course evaluation: what Alhija and Fresko found about who writes them and what they contain, how thematic analysis surfaces issues numbers miss (Stupans et al.), how rare abusive comments actually are (Tucker), and how to turn free text into reliable evidence.
Response Styles and Likert Scales: Why Cross-Cultural Evaluation Needs More Than Numbers
Acquiescence and extreme response styles vary systematically by culture (Harzing, 2006; Baumgartner & Steenkamp, 2001), which means raw Likert averages are not directly comparable across nationalities in Europe's multinational classrooms. This article explains the evidence and how to evaluate fairly across diverse cohorts.
How Many Scale Points Should a Course-Evaluation Question Have?
What the measurement literature — Preston & Colman (2000), Weng (2004), Dawes (2008) — says about the optimal number of response categories on rating scales, and why the answer for course evaluation is not just a number but a question about what a Likert item can and cannot capture.
What Do Student Evaluations Actually Measure? Marsh, the SEEQ, and the Case for Multidimensional Feedback
Herbert Marsh spent decades showing that well-built student evaluations are multidimensional, reliable, and stable — and that a single global score throws away most of what they can tell you. What his work establishes, where critics push back, and how to design feedback that is actually usable.