New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods9 min read

When Students Round to the Nearest Five: Digit Preference and Heaping in Numeric Course-Evaluation Answers

Typed numbers in evaluations pile up on multiples of five and ten. Whipple's and Myers' indices measure this heaping; it biases means and tails and flags data quality.

Koji Education Team

Product

In brief

When a course evaluation asks students to type a number — hours of independent study per week, lectures attended, an expected grade, days since the last assignment — the answers do not scatter smoothly. They pile up on multiples of five and ten: "10 hours," "20 hours," "50 percent." This is digit preference, or heaping, a well-studied form of measurement error in which respondents round to psychologically attractive values. Demographers quantify it with Whipple's and Myers' indices; survey statisticians model it directly. Heaping biases means and, more sharply, variances and tails, and it is a readable signal of low recall or low engagement — which is why it is worth measuring rather than ignoring.

What the research says

Digit preference has been recognised in demography for over a century. Whipple's index measures the excess of ages reported at multiples of five over what a smooth distribution would predict; Myers (1940), in "Errors and bias in the reporting of ages in census data," refined this into a blended index that detects preference or avoidance across all ten terminal digits (0–9) while cancelling the confounding effect of a genuinely sloping age distribution. These indices remain the standard demographic tools for grading the quality of self-reported numeric data, and international statistical agencies use them to appraise census age data.

A'Hearn, Baten and Crayen (2009), writing in the Journal of Economic History, gave the phenomenon a second life. In "Quantifying Quantitative Literacy," they showed that the degree of age heaping in a population is a robust proxy for numeracy and human capital — the more numerate a population, the less it rounds — and validated a simplified measure now known as the ABCC index against literacy at both individual and aggregate levels. The key inferential move is that heaping is not merely noise to be scrubbed; its amount carries information about the respondent.

On the modelling side, Wang and Heitjan (2008), in Statistics in Medicine, treated heaping as a coarsening mechanism rather than a nuisance to delete. Studying self-reported cigarette counts — where some respondents give exact figures and others round to multiples of twenty, ten or five — they built a likelihood model that separates the true underlying count from the rounding behaviour, showing that ignoring heaping biases estimates of the mean and that the propensity to round is itself patterned (heavier users round more). Their broader point generalises to any numeric self-report: heaping is coarsened data, and coarsening-aware estimation recovers less biased quantities than naive averaging.

Why it matters for course evaluation in practice

Course evaluations increasingly ask for numbers, not just Likert agreement. "How many hours per week did you spend on this course outside class?" "How many of the twenty lectures did you attend?" "What grade do you expect?" These items feed workload analytics, engagement dashboards, and accreditation evidence about student effort. If the answers heap, three practical harms follow.

First, biased central estimates. When students round study hours to the nearest five, the reported mean can drift away from the truth in a direction that depends on the rounding rule, and comparisons across cohorts or courses inherit that drift. Second, distorted spread and tails. Heaping artificially concentrates mass on round values and empties the values between them, so the variance, the percentiles, and any tail-based metric — the share studying "under 5 hours," say — become unreliable. Any method that relies on the shape of the distribution, from quantile summaries to workload thresholds, is affected. Third, and most usefully, a data-quality signal. Following A'Hearn and colleagues, the amount of heaping in a set of numeric responses is a marker of how carefully students engaged with the question. A cohort whose study-hours answers are 60 percent multiples of five is telling you something about recall effort and item comprehension, not just about study time. Computing a Whipple- or Myers-style index on numeric items gives a quantitative, defensible quality flag alongside the more familiar signals of straightlining and speeding.

Limitations and honest caveats

Heaping indices were designed for large populations and single-year age distributions; applied to a small class's study-hours answers they are noisy and can over-read chance clustering, so they are most trustworthy at programme or cohort scale. Some apparent heaping is rational, not careless: a student who genuinely studied "about ten hours" is giving an honest, appropriately uncertain answer, and punishing that rounding as an error would be wrong — the informative case is heaping far in excess of the true precision of the quantity. The indices also assume the underlying distribution is reasonably smooth; if real study hours genuinely cluster (many courses meet in two-hour blocks), a smoothness assumption will misclassify structure as heaping. Modelling approaches such as Wang and Heitjan's recover less biased estimates but require assumptions about the rounding mechanism that are hard to verify from the data alone, and they are heavier machinery than most quality offices will deploy. Finally, heaping is only one of several numeric-response pathologies; it does not address deliberate misreporting, non-response on the numeric item, or the deeper question of whether self-reported effort measures anything valid at all. It is a lens on how numbers were given, not a warrant that the numbers mean what you hope.

How Koji incorporates this

Koji's approach to numeric items is shaped by the recognition that a typed number is coarsened, self-reported data. Its conversational, AI-moderated design lets the platform do what a static form cannot: when a student answers "about ten hours," a follow-up probe can gently establish whether that is a precise figure or a rounded gesture, disambiguating the very heaping that would otherwise bias the aggregate. On the reporting side, Koji can compute heaping diagnostics — the concentration of numeric answers on round values — as a data-quality indicator for a cohort's numeric items, sitting alongside its existing screens for careless responding, straightlining and implausible speed, so that a workload figure comes with an honest quality caveat rather than a false air of precision. Where heaping is heavy, Koji flags the affected metric as coarse and steers reporting toward robust summaries and toward the qualitative signal in students' own words, rather than presenting a spuriously exact mean. The same measurement discipline runs through Koji's core research platform at koji.so, where product and customer-research teams collect numeric self-reports — usage frequency, willingness to pay, time spent — that heap for exactly the same psychological reasons and benefit from the same conversational disambiguation. Throughout, Koji frames these tools as designed to mitigate heaping and surface it, not to eliminate a phenomenon that is partly an honest expression of uncertainty.

Frequently asked questions

Is heaping the same as careless or insufficient-effort responding?

They overlap but are distinct. Careless responding shows up as straightlining and inconsistent answers across items; heaping is specifically the rounding of numeric answers to attractive values, and it can occur even in an otherwise careful respondent who simply does not remember an exact figure. Our note on careless responding covers the broader screen.

How do I actually measure heaping in my data?

For a quick appraisal, compute the share of numeric answers falling on multiples of five and ten and compare it with what a smooth distribution would give — the logic behind Whipple's index. Myers' blended index extends this across all ten terminal digits while correcting for a sloping underlying distribution. Both are simple to calculate on a cohort's responses.

Does a little rounding really matter?

A little is harmless and often honest. The concern is systematic heaping large enough to distort the mean, empty the values between the round numbers, and make percentiles and tail shares unreliable. The size of the effect, not its mere presence, is what matters.

Which course-evaluation items are most affected?

Any open numeric field: self-reported study or preparation hours, number of classes attended or missed, expected grade or mark, and counts of resources used. Bounded Likert items are not affected in the same way, though they have their own response-style issues.

Can I just delete or round the heaped answers?

Deleting throws away real information and can introduce its own bias, as our missing-data guidance explains. Modelling the heaping (as Wang and Heitjan do) or reporting robustly and flagging the item as coarse is preferable to silently "cleaning" it.

Is heaping ever useful rather than just a problem?

Yes. Following A'Hearn, Baten and Crayen, the amount of heaping is itself a proxy for how numerate and engaged respondents were with a numeric task. A sharp rise in heaping across cohorts can flag declining engagement or a poorly understood item, independent of the numbers themselves.

References

  • Myers, R. J. (1940). Errors and bias in the reporting of ages in census data. Transactions of the Actuarial Society of America, 41(2), 395–415.
  • A'Hearn, B., Baten, J., & Crayen, D. (2009). Quantifying Quantitative Literacy: Age Heaping and the History of Human Capital. Journal of Economic History, 69(3), 783–808. https://doi.org/10.1017/S0022050709001120
  • Wang, H., & Heitjan, D. F. (2008). Modeling heaping in self-reported cigarette counts. Statistics in Medicine, 27(19), 3789–3804. https://doi.org/10.1002/sim.3281

Related resources