Should a Course-Evaluation Scale Have a Neutral Midpoint? What "Neither Agree Nor Disagree" Really Captures
A neutral midpoint looks harmless, but the evidence shows it often functions as a hidden "don't know." Here is what the research says about including or omitting the middle option in course evaluations.
Koji Education Team
Product
In brief
The "neither agree nor disagree" midpoint on a course-evaluation scale looks like it captures genuine neutrality — but the strongest evidence suggests it often functions as a face-saving "don't know." Sturgis, Roberts and Smith (2014) showed experimentally that most respondents who pick the midpoint do not hold a truly neutral attitude; they use it to avoid the effort of forming or reporting a directional opinion. The practical implication for course evaluation is not a blanket "always include" or "always omit," but a design decision: distinguish genuine neutrality from no-opinion, offer an explicit "not applicable / no basis to judge" route, and beware that adding or removing a midpoint changes your distribution in ways that matter for comparison.
What the research says
The central study is Patrick Sturgis, Caroline Roberts and Patten Smith's 2014 article in Sociological Methods & Research, "Middle Alternatives Revisited: how the neither/nor response acts as a way of saying 'I don't know'." Using a split-ballot experiment, they compared a scale with an explicit midpoint against one without, and then probed why midpoint-choosers landed there. Their key finding: the bulk of respondents who select "neither agree nor disagree" are not reporting a balanced, considered neutrality. Instead, the midpoint operates as a low-effort refuge for people who lack a crystallised opinion or do not want to invest the cognitive effort to retrieve one — a "face-saving" alternative to admitting "I don't know." Removing the midpoint pushed some of these respondents to express a (weak) direction and others toward a "don't know" option, but it did not simply destroy real information.
This sits inside a well-developed survey-methodology literature. Jon Krosnick's work on satisficing (1991) provides the mechanism: when motivation or ability is low, respondents take the easiest defensible route through a question, and a midpoint is a conspicuously easy route. Nadler, Weston and Voyles (2015), in the Journal of General Psychology ("Stuck in the Middle"), asked respondents what they meant by the midpoint and found strikingly heterogeneous interpretations — "no opinion," "don't care," "unsure," "neutral," "equal/both," and "neither" — demonstrating that a single midpoint conflates several distinct mental states into one ambiguous data point. Chyung and colleagues (2017), reviewing the evidence for practitioners, frame it as a genuine trade-off: omitting the midpoint can force false directionality from truly ambivalent respondents, while including it invites satisficing and no-opinion contamination — so the right choice depends on whether neutrality is a substantively meaningful position for the question being asked.
A crucial distinction runs through all of this: neutrality is not the same as no-opinion. "I have a clear view and it is exactly balanced" is a different state from "I have no basis to answer." A single midpoint cannot tell them apart, which is why the methodological recommendation is increasingly to separate the two — keep a midpoint only where true neutrality is plausible, and provide an explicit, clearly-worded "no opinion / not applicable" route for the no-basis case (the subject of Should Course Evaluations Offer a "Don't Know" or "Not Applicable" Option?).
Why it matters for course evaluation in practice
This is not the same question as "how many scale points?" (covered in How Many Scale Points Should a Course-Evaluation Question Have?). You can have a 6-point scale with no midpoint or a 5-point scale with one; the presence of a neutral category is the design lever here, and it has specific consequences.
1. The midpoint shifts your distribution and your mean. A scale with a midpoint concentrates ambivalent and no-opinion responses in the centre, which interacts with the ceiling effects and skew institutions already battle. Switching an instrument from a 5-point (with midpoint) to a 4-point (forced choice) between cycles will move the numbers even if teaching did not change — a comparability trap for any year-over-year reporting.
2. A contaminated midpoint pollutes interpretation. If 30% of a class sits at "neither," a committee cannot tell whether the course genuinely produced balanced reactions or whether a third of students simply had no view and took the easy route. The number looks like signal; it may be mostly absence of signal.
3. For items some students cannot judge, the midpoint is the wrong tool. First-year students asked about "the relevance of the course to my career" may have no basis to answer; pushing them to a midpoint manufactures fake neutrality, whereas an explicit "not applicable" preserves honesty and keeps the genuine-neutral midpoint clean.
The defensible design, then: include a midpoint only where balanced opinion is a real possibility; always pair it with a distinct no-opinion/NA route rather than letting the midpoint absorb both; keep the choice stable across cycles for comparability; and, where you do report midpoint-heavy items, treat a large central mass as a flag to investigate rather than as evidence of consensus.
Limitations and honest caveats
A careful reader should resist over-reading the evidence. First, the finding is "often," not "always." Some midpoint responses are genuine neutrality, and forcing those respondents to fake a direction (by omitting the midpoint) introduces its own error — the choice trades one bias for another rather than eliminating bias. Second, much of this evidence comes from attitude surveys (political and social attitudes), not specifically from course evaluations; the cognitive dynamics are likely to transfer, but the exact proportions will differ for evaluative judgements about a course a student has just experienced. Third, item type matters: for a concrete behavioural item ("the lecturer returned feedback within two weeks") a midpoint makes less sense than for an attitudinal one, so a blanket scale policy across heterogeneous items is itself a compromise. Fourth, removing the midpoint does not cleanly recover real opinions — Sturgis and colleagues found it redistributes responses partly to weak directions and partly to "don't know," so the gain is "less ambiguity" rather than "more truth." Fifth, cross-cultural response styles complicate any fixed rule: midpoint usage varies systematically across cultures (see Response Styles and Likert Scales), so a midpoint decision that suits one cohort may distort comparisons across a multilingual student body. The honest position is that the midpoint is a genuine trade-off to be made deliberately and documented — not a settled best practice.
How Koji incorporates this
Koji's conversational model changes the terms of the midpoint debate rather than just picking a side of it.
- Disambiguating the middle in the moment. When a student gives a neutral or hedged answer, Koji's AI moderator can follow up — "you said it was about average; was there something specific that worked and something that didn't, or did you not really have a strong view?" — which separates genuine balanced opinion from no-basis-to-judge at the point of collection. That is precisely the distinction Sturgis, Roberts and Smith show a static midpoint cannot capture.
- Distinct structured options instead of one overloaded category. Koji's question types (scale, single_choice, yes_no, open_ended) let an instrument offer a true neutral point and a separate, explicitly worded "not applicable / no basis to judge" path, so the two mental states the literature warns about are not collapsed into one ambiguous datum.
- Surfacing the central mass honestly. Where many respondents cluster at neutral, Koji's reporting flags it and pairs it with the open-text the moderator elicited, so an evaluation committee sees why the centre is heavy rather than reading a midpoint-laden mean as consensus.
- Stability and documentation for comparability. Because Koji records the exact instrument configuration per cycle, an institution can keep midpoint design stable over time and detect when a change would break year-over-year comparison — guarding against the comparability trap above.
Koji frames this as mitigation, not magic: it cannot read minds, and a hedged conversational answer can still be ambiguous. What it does is reduce how often a single midpoint silently absorbs several different meanings. The same AI-moderated interview engine powers Koji's core research platform at koji.so, where the neutral-midpoint problem is just as live in customer and product surveys.
A simple decision rule
Faced with a specific item, ask three questions. First, is balanced opinion a substantively meaningful answer here? For an attitudinal item ("the course was intellectually stimulating") it usually is, so a midpoint earns its place; for a concrete behavioural item ("feedback was returned within two weeks") it rarely is, and a midpoint mostly invites satisficing. Second, can every respondent legitimately answer? If some students have no basis to judge — common for forward-looking items about career relevance — add an explicit "not applicable / no basis to judge" option so the genuine-neutral midpoint stays uncontaminated. Third, will this instrument be compared across cycles or cohorts? If so, fix the midpoint decision once and hold it constant, because changing it will move the numbers independently of any change in teaching. Applied item by item rather than as a single blanket scale policy, these three questions turn an unresolved methodological debate into a defensible, documented design choice — which is exactly what a sceptical accreditation reviewer or institutional-research analyst will expect to see.
Related resources
- How Many Scale Points Should a Course-Evaluation Question Have?
- Should Course Evaluations Offer a "Don't Know" or "Not Applicable" Option?
- Should Every Point on a Course-Evaluation Scale Be Labelled? The Evidence on Verbal Anchors
- Agree/Disagree or Item-Specific? The Scale Choice That Quietly Degrades Course-Evaluation Data
- Why Students Click Straight Down the Middle: Satisficing in Course Evaluations
- Do the Numbers on Your Rating Scale Change the Score? The Evidence on Numeric Anchors
References
- Sturgis, P., Roberts, C., & Smith, P. (2014). Middle alternatives revisited: how the neither/nor response acts as a way of saying "I don't know"? Sociological Methods & Research, 43(1), 15–38. https://doi.org/10.1177/0049124112452527
- Nadler, J. T., Weston, R., & Voyles, E. C. (2015). Stuck in the middle: the use and interpretation of mid-points in items on questionnaires. The Journal of General Psychology, 142(2), 71–89. https://doi.org/10.1080/00221309.2014.994590
- Krosnick, J. A. (1991). Response strategies for coping with the cognitive demands of attitude measures in surveys. Applied Cognitive Psychology, 5(3), 213–236. https://doi.org/10.1002/acp.2350050305
- Chyung, S. Y., Roberts, K., Swanson, I., & Hankinson, A. (2017). Evidence-based survey design: the use of a midpoint on the Likert scale. Performance Improvement, 56(10), 15–23. https://doi.org/10.1002/pfi.21727
Related articles
Why Students Click Straight Down the Middle: Satisficing in Course Evaluations
A large share of students do not answer evaluation questions carefully — they straightline, speed through, and pick the easy option. Barge and Gehlbach (2012) showed this "satisficing" doesn't just add noise; it inflates your reliability and validity statistics, making bad data look good.
Do the Numbers on Your Rating Scale Change the Score? The Evidence on Numeric Anchors
Schwarz and colleagues showed that the numeric values printed on a rating scale (0 to 10 vs minus 5 to plus 5) systematically shift responses even when the verbal labels are identical. Here is what that means for course-evaluation design, comparability, and reporting.
Agree/Disagree or Item-Specific? The Scale Choice That Quietly Degrades Course-Evaluation Data
"The lecturer was well organised: Strongly disagree to Strongly agree" feels natural, but the agree/disagree format invites acquiescence and lower data quality. Saris et al. (2010) on why item-specific scales measure better.
How Many Scale Points Should a Course-Evaluation Question Have?
What the measurement literature — Preston & Colman (2000), Weng (2004), Dawes (2008) — says about the optimal number of response categories on rating scales, and why the answer for course evaluation is not just a number but a question about what a Likert item can and cannot capture.