New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Methodology9 min read

Yes, Yes, Yes: Acquiescence, Straightlining and the Data-Quality Problem in Course Evaluation

A meaningful share of every course-evaluation dataset is not real opinion — it is acquiescence, straightlining, and careless responding. Here is the evidence on how common it is, why it quietly distorts your results, and how to design feedback that resists it.

Koji Education Team

Product ·

Bottom line up front: Not every response in your course-evaluation dataset reflects a considered opinion. Some students agree with whatever is put to them (acquiescence, or "yea-saying"), some click the same point down every row (straightlining), and some answer essentially at random to finish quickly (careless or insufficient-effort responding). These are not rare edge cases. In one large study of student self-report data, depending on the detection method, between roughly 14% and 37% of respondents showed at least one marker of careless responding. Because these patterns are systematic rather than purely random, they bias your scales — usually upward and toward the middle — and they are easy to miss in an aggregate mean. The most durable defence is not a longer survey with more attention checks; it is a feedback mode that makes thoughtless responding harder than thoughtful responding.

Three distinct failure modes, often lumped together

It helps to separate the phenomena, because they have different signatures and different fixes.

Acquiescence bias is the tendency to agree. Faced with "The instructor explained concepts clearly — agree or disagree?", a yea-sayer leans toward "agree" regardless of the actual experience. As one Cambridge Core analysis in Political Analysis puts it, acquiescence is "common, systematic, and easy to miss" — and part of what looks like acquiescence is in fact an artifact of inattentive responding. Crucially, attention checks alone do not fully resolve it.

Straightlining is choosing the same scale point down a whole battery of items — all 4s, all 5s — without differentiating between questions that genuinely deserve different answers. It is fast, it is effortless, and on a typical ten-item teaching-evaluation grid it is almost invisible once averaged.

Careless or insufficient-effort responding is the umbrella: random clicking, pattern-following, or rushing. It includes acquiescence and straightlining as special cases but also covers genuinely arbitrary answers.

How common is it? The numbers are sobering

The most useful recent evidence comes from Chauliac, Willems, Gijbels and Donche (2023), published in Frontiers in Education. They examined self-report questionnaires on student learning from 12,578 students (mean age 17) and applied five different indicators of careless responding. The headline figures:

  • Using a lenient criterion, 14.17% of respondents were flagged as careless by at least one of the five indicators.
  • Using a strict criterion, that rose to 37.45%.
  • Individual detection methods flagged anywhere from 0.6% to 23% of respondents, depending on the method.

The study also confirmed the consequence: careless respondents showed significantly lower internal-consistency (Cronbach's alpha) on the scales than careful respondents. In other words, a measurable fraction of the data was adding noise — and worse, structured noise.

This matters because course evaluation is, methodologically, exactly this kind of self-report questionnaire, often administered at the end of a long term to tired students who have several forms to clear. The conditions that produce careless responding are baked into how most universities collect feedback.

Why "structured noise" is worse than random noise

A tempting reassurance is that careless responses cancel out — random clicks scatter in all directions and wash out in the average. That is true only for genuinely random responding. Acquiescence and straightlining are directional. Yea-saying pushes systematically toward agreement; straightlining tends to cluster on the positive-but-not-extreme points (lots of 4s). These do not cancel; they bias. They compress variance, inflate apparent consensus, and make a mediocre course look quietly fine.

That compression interacts with a problem we have written about before: the mean hides the distribution. When a chunk of your respondents are straightlining 4s, your standard deviation shrinks and your dashboard reports a falsely confident "students broadly agree." It also feeds the ceiling effect: yet another reason almost every course lands around 4 out of 5.

Critics argue: "Just add attention checks and reverse-coded items"

This is the standard textbook prescription, and it is not wrong — but it is incomplete, and in course evaluation it can backfire.

Reverse-coded items (deliberately phrasing some questions negatively so that a straightliner contradicts themselves) do help detect careless responding. But research shows reverse-wording introduces its own problems: confused respondents, lower reliability, and method artifacts that can be mistaken for real factor structure. Attention-check items ("Select 'strongly agree' for this question") catch the most blatant cases but, as the Political Analysis work above notes, do not fully resolve acquiescence — and they irritate the conscientious students you most want to keep engaged.

There is also a deeper issue: these are detection and deletion strategies. They help you throw out bad data after the fact, shrinking your already-strained sample (a problem we explore in response rates and non-response bias). They do nothing to prevent careless responding in the first place. And the more you bolt detection items onto a survey, the longer and more tedious it becomes — which increases the fatigue that drives careless responding. You can end up chasing your own tail.

The real lever: make thoughtless answering harder than thoughtful answering

Careless responding thrives in a specific environment: a long grid of similar-looking Likert items, low stakes, no interaction, and an obvious "just click down the middle and submit" escape route. Change that environment and you change the behaviour.

  • Ask fewer, better questions. A short, well-targeted instrument gives the straightliner less to straightline. See our guide to writing better questions.
  • Break the grid. Acquiescence and straightlining are partly products of the matrix format. Mixing question types and asking one thing at a time disrupts the autopilot.
  • Make it conversational. It is hard to "yea-say" your way through a question that asks you to describe something in your own words and then asks a genuine follow-up.

Where Koji fits

Koji for Education is built around that last lever. Instead of a static matrix of agree/disagree items, Koji runs an AI-moderated conversational interview. The format itself resists acquiescence and straightlining: there is no row of identical Likert items to click straight down, and an open-ended answer cannot be "yea-said." When a student gives a thin or reflexive answer, the AI asks a neutral follow-up — which is precisely the kind of engagement that converts a careless respondent into a thoughtful one, rather than just flagging them for deletion.

Koji's quality scoring then assesses the substance of each response, surfacing low-effort or contradictory contributions so they can be weighted appropriately — a detection layer that complements, rather than replaces, the prevention built into the conversational format. Its automatic thematic analysis works on what students actually say, so the signal comes from content, not from the position of a click on a five-point scale. And its six structured question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) let you deploy a scale where a scale is genuinely the right tool — without building the endless grid that breeds straightlining. As always, the honest framing: Koji reduces and surfaces careless responding; it cannot guarantee every respondent engages fully.

The same conversational engine underpins the main Koji platform for customer and market research, where careless responding in panel surveys is a notorious data-quality threat — so these mechanisms are tested at scale well beyond higher education.

The takeaway for institutional researchers

Before you act on a course-evaluation mean, ask how much of it is real opinion. If a fifth to a third of your respondents may be acquiescing, straightlining, or clicking through to be done, then small differences between courses or instructors are well within the noise floor — a point we make in is a 0.3-point difference real?. The durable response is not more attention-check items grafted onto a tired survey. It is a feedback mode designed so that giving a thoughtful answer is easier than faking one.


Koji for Education collects student feedback in a format that resists acquiescence and straightlining by design. See how Koji works.