New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods9 min read

Grid or One Question at a Time? Matrix Formats and Straightlining in Course Evaluations

Rendering rating items as a single grid instead of one question per screen quietly degrades course-evaluation data. The controlled evidence on item non-response, breakoff and straightlining — and what to do about it.

Koji Education Team

Product

Answer in brief. Presenting several rating items as a single grid (matrix) rather than one question at a time makes a course-evaluation survey look shorter, but it measurably lowers data quality: matrix layouts increase item non-response and — under some conditions — non-differentiation ("straightlining"), and the harm grows with more scale columns and on mobile devices. Controlled web-survey experiments (Couper, Tourangeau, Conrad & Zhang 2013; Liu & Cernat 2018) show these effects are driven by the format, not just the respondent. For course evaluations, prefer item-by-item delivery or narrow "split" grids, keep matrices short, render mobile-first, and monitor straightlining as a live data-quality indicator rather than assuming your averages are clean.

What the research says

Almost every institutional course evaluation asks students to rate a battery of statements — "The lecturer explained concepts clearly," "The assessment criteria were fair," "The workload was manageable" — on one shared response scale. The default way to present that battery is a grid (or matrix): items are rows, scale points are columns, and the student answers many items on a single screen. It is compact and feels efficient. That efficiency is exactly the problem.

The theoretical starting point is Jon Krosnick's satisficing account of survey response (Krosnick, 1991). Answering a question well requires four cognitive steps — comprehending the item, retrieving relevant information, forming a judgement, and mapping it onto the response options. A motivated respondent does all four (optimising); a fatigued or unmotivated one shortcuts them (satisficing). A grid is an engine for strong satisficing: once a student has picked a plausible value in the first row, the visual layout invites them to simply repeat that value straight down the column. Tourangeau, Couper and Conrad's earlier work on the visual features of survey questions showed that respondents read meaning into layout itself — items placed near each other are assumed to be related, and the middle of a scale is read as the "typical" or expected value — heuristics that a grid amplifies.

The anchor study is Couper, Tourangeau, Conrad and Zhang (2013), "The Design of Grids in Web Surveys," Social Science Computer Review, 31(3), 322–345. Across two web-survey experiments the authors manipulated grid design — baseline grids, dynamic shading that highlighted the active row, and "split" grids that broke one large matrix into smaller blocks. The clearest result concerned missing data: the baseline grid produced a mean of about 0.107 missing items per respondent, while split-grid and dynamic-feedback versions cut that to roughly 0.010–0.029 — a large relative reduction from a simple layout change. Notably, and to their credit as careful researchers, they found straightlining was rare in their data: only about 1.1% of respondents gave identical answers to all 13 items and 2.3% to twelve or more. Adding visual complexity (alternating colours, varied typefaces) did nothing. The practical lesson: breaking a big grid into smaller pieces, or moving toward one-item-at-a-time, recovers real data.

A strong corroborating study is Liu and Cernat (2018), "Item-by-item Versus Matrix Questions: A Web Survey Experiment," Social Science Computer Review, 36(6), 690–706. They varied the number of response columns (2, 3, 4, 5, 7, 9 and 11) and compared matrix against item-by-item presentation. Their headline finding: data quality for matrix questions deteriorates as the number of columns grows, especially at 9 and 11 options, and item non-response is consistently higher for matrix than for item-by-item questions — particularly among mobile respondents. In their experiment straightlining and response time were similar across formats, which again cautions against treating straightlining as the guaranteed cost; the robust, replicated cost is item non-response and its interaction with wide scales and small screens.

Taken together, the literature converges on a defensible synthesis: grids are not automatically ruinous, but they carry format-driven risks — skipped rows (especially in the middle and bottom of a long matrix, which respondents literally do not notice), breakoff, and a heightened invitation to non-differentiate — and those risks scale with the width of the scale and the smallness of the screen. Item-by-item and split grids reduce them.

Why it matters for course evaluation in practice

Course evaluations sit in a uniquely hostile measurement environment: response rates are often already low, the incentive to answer carefully is weak, and completion has migrated overwhelmingly to phones. Every format-driven loss compounds an existing problem.

  • Non-response bias gets worse. A student who abandons a long grid halfway does not contribute a partial-but-usable record on the items that matter most; they inflate missingness precisely on the later, often more diagnostic, items.
  • Straightlining destroys the multidimensional signal you paid to collect. If a well-designed instrument distinguishes clarity, fairness, workload and organisation, a straightlined column of 4s collapses all of that into noise — and, perversely, makes your instrument look more internally consistent (a spuriously high Cronbach's alpha), giving false reassurance about reliability.
  • The mobile penalty lands where your data lives. Because students overwhelmingly complete evaluations on phones, the matrix disadvantage that Liu and Cernat found strongest on mobile is not an edge case — it is your median respondent.
  • Wide scales multiply the cost. Institutions fond of 7- or 11-point agreement scales inside a grid are combining the two conditions the evidence flags as worst.

Limitations and honest caveats

A PhD reader should push back on any "grids are bad, item-by-item is good" slogan, and the evidence rewards that scepticism.

First, straightlining is not a guaranteed consequence. The flagship Couper et al. study found it rare, and Liu and Cernat found it comparable across formats. The reliably replicated harm is item non-response and breakoff, not non-differentiation. Overstating the straightlining claim would itself be a misreading of the literature.

Second, most of this evidence comes from general web panels, not university course evaluations. Generalisability to a captive student population completing a familiar institutional survey is plausible but not established; a motivated cohort answering a short, well-designed grid may show none of these problems.

Third, item-by-item is not free. Splitting a battery into many single-question screens lengthens the perceived survey, adds clicks, and can itself raise dropout; it also removes the side-by-side comparison that some respondents use meaningfully to calibrate their answers. The optimum is usually short, narrow, well-signposted blocks, not an atomised one-item-per-screen slog.

Fourth, layout co-varies with length, progress indication and device, so isolated "format effects" in any single study are partly confounded. The safest reading is directional, not a precise effect size to import into your own context.

How Koji incorporates this

Koji for Education is built around a delivery model that sidesteps the matrix penalty rather than trying to decorate a grid into safety.

  • One thing at a time. Koji's AI-moderated conversational evaluation presents each structured question in turn rather than as a wall of rows. This is, in effect, an item-by-item delivery — the format the evidence associates with lower item non-response — without the fatigue of an endless single-question survey, because the moderator paces and threads the questions naturally.
  • Structured question types, asked sequentially. Scale, single_choice, multiple_choice, ranking, yes_no and open_ended items are administered in a flow the moderator controls, so a student cannot skim a column of rows and copy a value downward.
  • Non-differentiation becomes a probe, not a silent data point. When a respondent gives uniformly flat ratings, the moderator is designed to follow up ("You rated most things around a 4 — was there anything that felt notably weaker or stronger?"), converting a satisficing signal into usable qualitative detail. This is the opposite of a grid, where straightlining is invisible until analysis.
  • Mobile-first by construction. Because the interview is conversational, it renders naturally on a phone, avoiding the horizontal-scroll and mis-tap problems that drive mobile matrix non-response.
  • Quality scoring and bias-aware reporting flag low-effort or contradictory responses so that QA officers can weight or set them aside, rather than averaging them in unnoticed.

None of this eliminates satisficing — a determined student can still answer thinly, and Koji is designed to mitigate, not abolish, the problem. But it removes the specific structural trap the grid literature identifies. Koji's core research platform at koji.so applies the same AI-moderated, one-question-at-a-time engine to product and customer research, where matrix fatigue is an equally well-documented threat to data quality.

Related resources

References

  • Couper, M. P., Tourangeau, R., Conrad, F. G., & Zhang, C. (2013). The design of grids in web surveys. Social Science Computer Review, 31(3), 322–345. https://doi.org/10.1177/0894439312446510
  • Liu, M., & Cernat, A. (2018). Item-by-item versus matrix questions: A web survey experiment. Social Science Computer Review, 36(6), 690–706. https://doi.org/10.1177/0894439316674459
  • Tourangeau, R., Couper, M. P., & Conrad, F. G. (2004). Spacing, position, and order: Interpretive heuristics for visual features of survey questions. Public Opinion Quarterly, 68(3), 368–393. https://doi.org/10.1093/poq/nfh035
  • Krosnick, J. A. (1991). Response strategies for coping with the cognitive demands of attitude measures in surveys. Applied Cognitive Psychology, 5(3), 213–236. https://doi.org/10.1002/acp.2350050305