New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods9 min read

How Often Is "Often"? Why Vague Quantifiers Quietly Distort Course-Evaluation Data

Words like "often", "usually" and "sometimes" mean different things to different students, which makes course-evaluation responses hard to compare. Here is the survey-methodology evidence and what to do about it.

Koji Education Team

Product

In brief

Response options built from vague quantifiers — "always", "often", "sometimes", "rarely" — are interpreted differently by different respondents and across different question contexts, so two students who behaved identically can choose different options, and two students who choose "often" may mean very different frequencies. The classic evidence is Schaeffer's (1991) demonstration that such phrases shift meaning across subgroups, and higher-education-specific replications on the National Survey of Student Engagement (NSSE) show the same problem in our own sector. For course evaluation this matters because group comparisons (by gender, discipline, cohort, campus) can be partly an artefact of differential interpretation rather than real differences in teaching. The practical fix is to anchor frequency questions to concrete referents or numeric ranges, or to probe what the student actually means — which is where AI-moderated follow-up has a role.

What the research says

The foundational study is Nora Cate Schaeffer's "Hardly Ever or Constantly? Group Comparisons Using Vague Quantifiers" (Public Opinion Quarterly, 1991). Using a national probability sample (N≈1,172), Schaeffer examined frequency reports of excitement and boredom and showed that the mapping between a vague quantifier (e.g. "often") and an underlying frequency is not constant: it varies systematically with the respondent's race, sex, education and age, and with whether the scale is framed in relative ("often") or absolute ("once a week") terms. Crucially, conclusions about group differences changed depending on whether absolute or relative frequency language was used — meaning the measurement instrument, not the respondents, partly produced the "difference".

This is not an isolated finding. A broad survey-methodology literature (Bradburn, Sudman & Wansink; Schaeffer & Presser, 2003, Annual Review of Sociology) documents that vague quantifiers carry no fixed numeric value and are pulled around by context, the range of other options offered, and conversational norms about what a "normal" answer looks like.

For our sector specifically, Rocconi, Dumford and Butler (2020), "Examining the Meaning of Vague Quantifiers in Higher Education: How Often is 'Often'?" (Research in Higher Education) tested the exact quantifiers used on the NSSE — "never", "sometimes", "often", "very often". They found the numeric meanings students assign to these labels are unequal and non-linear: the perceived distance between "sometimes" and "often" is not the same as between "often" and "very often". Treating such a scale as equal-interval (e.g. averaging it 1–2–3–4) therefore mis-states the data. Earlier work by the same research stream (Laird and colleagues, 2008) reached compatible conclusions about NSSE quantifiers.

Why it matters for course evaluation in practice

Most course-evaluation instruments are riddled with vague quantifiers, usually without anyone noticing:

  • "The instructor often related material to real-world examples."
  • "I regularly received useful feedback."
  • "Class time was usually well organised."

Three concrete problems follow from the research:

  1. Between-group comparisons are unsafe. If international students interpret "often" more conservatively than home students (a plausible cross-cultural difference — see our note on response styles), a department can read a "gap" that reflects language, not teaching. Schaeffer's result is precisely that subgroup conclusions flip with quantifier framing.

  2. Trends over time can be artefacts. If you change a scale's wording — or even the surrounding items — the meaning of "often" can drift, contaminating year-on-year comparisons that quality committees treat as real movement.

  3. Averaging is mathematically dubious. Rocconi et al. show the intervals between labels are unequal, so the mean of a "never/sometimes/often/very often" scale has no clean interpretation. This compounds the well-known caution about averaging Likert scores.

Limitations and honest caveats

A PhD reader will rightly push back, and the honest position is nuanced:

  • Vagueness is not always fatal. Some research (e.g. work finding vague quantifiers show limited frame-of-reference effects in certain designs) suggests that for within-person, within-instrument comparisons the noise can partly cancel out. The strongest warnings apply to between-group and between-instrument comparisons.
  • Numeric alternatives have their own problems. Replacing "often" with "3–4 times per term" assumes the respondent can recall and count accurately — a strong assumption given memory limits and end-of-term recall decay. You may trade interpretation variance for recall error.
  • Generalizability. Schaeffer's data are from a 1970s US sample about emotions, not 2020s European students about teaching; the NSSE studies are US engagement surveys. The mechanism (labels lack fixed meaning) generalizes well, but exact effect sizes for a given course-evaluation item should be established locally, ideally through cognitive interviewing.
  • Construct fit. For genuinely fuzzy constructs (overall impressions), a vague label may match how students actually think better than a false-precision count would.

How Koji incorporates this

Koji is designed to mitigate, not magically eliminate, the vague-quantifier problem, through several concrete mechanisms:

  • Behaviourally specific question design. Koji's structured question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) let you replace "The instructor often gave feedback" with a concrete, countable referent — for example a single_choice item with options anchored to identifiable events ("feedback on every assignment / most assignments / some / none"), reducing reliance on a free-floating "often".
  • Conversational disambiguation. This is the distinctive part. When a student selects a vague option, Koji's AI-moderated interview can probe ("You said feedback came 'often' — roughly how often, and on which kinds of work?"), converting a fuzzy label into an interpretable, evidence-bearing statement. A static survey cannot ask this follow-up; the conversational layer is built precisely to recover the meaning behind the quantifier.
  • Thematic analysis over labels. Because Koji captures open text and transcripts, its automatic thematic analysis can surface what students mean by "regular feedback" rather than counting checkbox selections, triangulating the Likert number against the student's own words.
  • Bias-aware reporting. Koji's reporting is built to flag when apparent subgroup "gaps" rest on small samples or ambiguous items, discouraging the over-reading of differences that Schaeffer warns are partly artefactual.

We frame these as mitigations: conversational probing reduces, but does not remove, interpretation variance, and good instrument design upstream still matters most. Koji's core research platform at koji.so applies the same AI-moderated interview engine to product and customer research, where vague-quantifier ambiguity is an equally well-known headache.

Practical checklist

  • Audit every evaluation item for words like often, usually, regularly, sometimes, frequently.
  • Where the construct is a countable behaviour, anchor to concrete referents or event counts.
  • Avoid cross-group league tables built on vague-quantifier items unless you have evidence the groups interpret the labels equivalently.
  • Do not average a "never/sometimes/often/very often" scale as if it were equal-interval; report distributions, and pair it with open text.
  • Pre-test new items with a handful of students using cognitive interviewing before fielding.
  • When you must compare cohorts or campuses, establish first that the groups read the labels equivalently; otherwise report each group's distribution separately rather than a single ranked table.
  • Document any wording change to a frequency item, because altering the label or its neighbours can shift its meaning and silently break year-on-year trend comparisons.

Related Resources

Worked example: rewriting a vague-quantifier item

Take a common item: "The instructor often provided helpful feedback." A student answers on a 5-point agree/disagree scale. Two students who received feedback on every single assignment may answer differently — one reads "often" as "nearly always" and agrees strongly; the other reserves "often" for daily contact and only somewhat agrees. The word, not the teaching, splits them.

A behaviourally anchored rewrite removes the vague quantifier entirely:

"On how many of your assessments did you receive written feedback from the instructor?" All of them / Most of them / About half / A few / None

This is countable, has a concrete referent, and supports honest aggregation (a proportion, not a fuzzy mean). A second item can capture the quality dimension separately — note that bundling frequency and helpfulness into one item would re-introduce a double-barreled problem.

Where the underlying construct is genuinely about subjective frequency and a count is unrealistic, the next-best option is to define the quantifier in the question stem ("By 'regularly' we mean roughly weekly or more") so that at least the intended meaning is shared, even if recall remains imperfect.

A useful institutional habit is to keep a banned-words list for instrument authors — often, usually, regularly, frequently, sometimes, occasionally, rarely — that triggers a review whenever one appears. Pairing that list with a short cognitive-interviewing pass before each survey cycle catches the quantifiers that slip through and reveals how your specific student population reads them, which is the only way to know whether a label travels safely across your cohorts and campuses.

References

  • Schaeffer, N. C. (1991). Hardly ever or constantly? Group comparisons using vague quantifiers. Public Opinion Quarterly, 55(3), 395–423. https://doi.org/10.1086/269270
  • Rocconi, L. M., Dumford, A. D., & Butler, K. (2020). Examining the meaning of vague quantifiers in higher education: How often is "often"? Research in Higher Education, 61, 1019–1041. https://doi.org/10.1007/s11162-020-09587-8
  • Schaeffer, N. C., & Presser, S. (2003). The science of asking questions. Annual Review of Sociology, 29, 65–88. https://doi.org/10.1146/annurev.soc.29.110702.110112
  • Bradburn, N. M., Sudman, S., & Wansink, B. (2004). Asking Questions: The Definitive Guide to Questionnaire Design. Jossey-Bass.

Related articles