New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods11 min read

Measuring What Students Will Not Admit Directly: The List Experiment for Sensitive Course-Evaluation Questions

Some of the most important evaluation questions - Did you actually attend? Did you use AI on assessments? Did the grade you expected shape your rating? - are exactly the ones students answer dishonestly. The list experiment (item-count technique) estimates their prevalence without ever asking any student to admit anything.

Koji Education Team

Product

The short answer

Some of the most consequential course-evaluation questions are the ones students are least willing to answer truthfully: Did you attend most classes? Did you use AI to complete assessments? Did the grade you expected influence the rating you gave? Direct questions on these topics are contaminated by social-desirability bias, so the numbers you get are wrong in a predictable direction. The list experiment — also called the item-count technique (ICT) — is a survey design that estimates the prevalence of a sensitive attitude or behaviour across a cohort without any individual ever revealing where they stand on it.

BLUF: In a list experiment you randomly split respondents into two groups. A control group sees a short list of innocuous items and reports only how many apply to them — not which ones. A treatment group sees the identical list plus one sensitive item and again reports only the total count. Because the two groups are randomly equivalent, the difference in their mean counts estimates the proportion of the cohort for whom the sensitive item is true — while every respondent enjoys complete deniability, since a raw number never exposes any single answer. It trades individual-level data and statistical precision for honesty on questions a direct survey cannot measure.

What the research says

The technique has a mature methodological literature in political science and survey methodology, where it is used for topics like racial prejudice, vote-buying, and turnout overreporting.

Glynn (2013), "What Can We Learn with Statistical Truth Serum?" (Public Opinion Quarterly) is the standard design reference. Glynn shows that the quality of a list experiment is made or broken by the choice of control items, and derives the practical rules: pick control items whose prevalences are negatively correlated (so respondents rarely hit the floor of 0 or the ceiling of all items), avoid items that are near-universal or near-absent, and prefer the double list experiment — where every respondent serves as treatment for one list and control for another — to cut variance roughly in half. He also provides sample-size formulas, and they deliver the technique's central warning: because you are estimating a prevalence from a difference of two noisy means, you need substantially larger samples than a direct question to reach the same precision.

Blair & Imai (2012), "Statistical Analysis of List Experiments" (Political Analysis) moved the field beyond the naïve difference-in-means estimator. They developed maximum-likelihood and regression estimators that let you model who holds the sensitive attitude as a function of covariates, and — crucially — diagnostic tests for the technique's two behavioural assumptions: "no design effect" (adding the sensitive item does not change how respondents answer the control items) and "no liars" (respondents answer the aggregate count truthfully). Their methods ship in the openly available R package list.

Does it actually reduce bias? Holbrook & Krosnick (2010) (Public Opinion Quarterly) applied the item-count technique to voter-turnout overreporting — a textbook social-desirability problem — and found that ICT-based turnout estimates were meaningfully lower than direct self-reports, consistent with the technique stripping out socially-desirable lying. And the largest synthesis, Blair, Coppock & Moor (2020) (American Political Science Review), meta-analysed roughly three decades of list experiments and offered a "social reference theory" that predicts when sensitivity bias is large enough to justify the technique's cost — a reminder that the list experiment is a targeted instrument, not a default.

Why it matters for course evaluation in practice

Several questions a modern quality-assurance office genuinely wants answered are sensitive by construction:

  • Attendance honesty: "I attended fewer than half the live sessions." Self-attendance is systematically over-reported, which distorts any attempt to relate attendance to ratings (see attendance as a confound where available).
  • Academic-integrity and AI use: "I used a generative-AI tool to complete assessed work in this course." A direct question invites denial; a list experiment can estimate cohort-level prevalence to inform policy without accusing anyone.
  • Grade-driven rating: "My rating was influenced by the grade I expected." Students are reluctant to admit this, yet it is central to the grading-leniency debate.
  • Experiences of bias or misconduct: "I witnessed or experienced inappropriate conduct in this course." Here deniability is not just about honesty but about safety.

For these, the list experiment gives a defensible prevalence estimate at the cohort or programme level — the unit at which QA policy is actually made — while structurally protecting respondents.

How this differs from the Randomized Response Technique

Koji's knowledge base already covers the Randomized Response Technique (RRT), and the two are easily confused. The difference is fundamental: RRT still asks each respondent the sensitive question directly, masking the answer with a randomising device (a coin or die) so that any individual "yes" is deniable. The list experiment never asks the individual the sensitive question at all — it only ever collects a count. That makes the list experiment more intuitive for respondents (no dice, no instructions to distrust) and often more protective, at the cost of being purely aggregate: RRT can, with modelling, support individual-level analysis in ways the classic list experiment cannot, and RRT can be more sample-efficient. Choose RRT when you must retain a per-respondent indicator; choose the list experiment when maximal deniability and respondent comprehension matter most.

Limitations and honest caveats

  • It is aggregate-only. You get a prevalence, not a flag on any student, and not a clean cross-tab unless you use the Blair–Imai regression estimators with adequate sample. It cannot drive individual or instructor-level decisions.
  • It is statistically hungry. Estimating a difference of means to useful precision can require several hundred to a few thousand respondents per item — often infeasible for a single small class, and best reserved for programme- or institution-level measurement.
  • Design effects and "ceiling/floor" leakage. If a respondent's true count is 0 or "all", the design can inadvertently reveal their sensitive answer, breaking deniability and biasing the estimate. Careful control-item selection (Glynn) and the Blair–Imai diagnostics are mandatory, not optional.
  • It can produce impossible estimates. Naïve estimators sometimes yield negative prevalences in subgroups — a signal of assumption violations, not a usable number.
  • It measures prevalence, not mechanism. A list experiment tells you how many, never why. It must be paired with qualitative inquiry to be actionable.

How Koji incorporates this

Koji treats sensitive measurement as a design problem, not an afterthought.

  • List-experiment templates. Koji's structured-question engine can field a randomised two-arm (or double-list) item-count block, randomising respondents to control and treatment lists and computing the difference-in-means prevalence with confidence intervals — including the double-list variance reduction Glynn recommends.
  • Built-in diagnostics. Reports surface the Blair–Imai style checks (implied negative proportions, ceiling/floor exposure) so an analyst is warned when the "no design effect" or "no liars" assumptions look violated, rather than quietly reporting a broken estimate.
  • Triangulation with the conversational interview. Because a list experiment yields prevalence but not mechanism, Koji's AI-moderated conversational interviews run alongside to explore why a behaviour is common — for example probing, in a fully separate and non-identifying strand, how students think about AI use — so the QA office gets both the magnitude and the reasoning. This is designed to mitigate social-desirability bias on the count, not to eliminate dishonesty everywhere.
  • Governance-aware deployment. Koji's anonymity and small-cell protections (cell suppression, minimum response thresholds) are applied so that a list-experiment block cannot be combined with other data to re-identify a respondent.

The same engine powers sensitive prevalence questions in product and customer research on Koji's core platform at koji.so, where honest measurement of behaviours people are reluctant to admit is just as valuable.

Related resources

Frequently asked questions

How does a list experiment protect a student's privacy if the survey still records answers?

Each respondent reports only a total count of how many items on a list apply to them, never which ones. Because the sensitive item is one of several and the individual answers are never disaggregated, no single response reveals a student's position on the sensitive item — deniability is built into the design.

How is the list experiment different from the Randomized Response Technique?

The Randomized Response Technique still asks each student the sensitive question directly and masks the answer with a randomising device. The list experiment never asks the individual the sensitive question — it only collects a count and estimates prevalence across the group. The list experiment is usually more intuitive and more protective; RRT can better support individual-level analysis.

How many responses do I need?

More than a direct question. Because you estimate a difference between two noisy means, useful precision often needs several hundred to a few thousand respondents, which makes the technique suited to programme- or institution-level measurement rather than a single small class. Use published sample-size formulas (Glynn, 2013) during design.

What is a design effect and why does it matter?

A design effect occurs when adding the sensitive item changes how respondents answer the control items, or when a respondent whose true count is zero or the maximum is forced to reveal their sensitive answer. Both bias the estimate and can break deniability, so control items must be chosen so that extreme totals are rare.

Can a list experiment tell me which instructors or cohorts are affected?

Only if you use regression-based estimators (Blair and Imai, 2012) with sufficient sample, and even then at the subgroup rather than individual level. The classic estimator yields a single prevalence. It is not a tool for instructor-level decisions.

When should I not use a list experiment?

When the question is not genuinely sensitive (a direct question is cheaper and more precise), when you need per-respondent data, or when your sample is too small to estimate a difference of means reliably. It is a targeted instrument for high-stakes, socially-loaded questions.

References

  • Glynn, A. N. (2013). What can we learn with statistical truth serum? Design and analysis of the list experiment. Public Opinion Quarterly, 77(S1), 159–172. https://doi.org/10.1093/poq/nfs070
  • Blair, G., & Imai, K. (2012). Statistical analysis of list experiments. Political Analysis, 20(1), 47–77. https://doi.org/10.1093/pan/mpr048
  • Holbrook, A. L., & Krosnick, J. A. (2010). Social desirability bias in voter turnout reports: Tests using the item count technique. Public Opinion Quarterly, 74(1), 37–67. https://doi.org/10.1093/poq/nfp065
  • Blair, G., Coppock, A., & Moor, M. (2020). When to worry about sensitivity bias: A social reference theory and evidence from 30 years of list experiments. American Political Science Review, 114(4), 1297–1315. https://doi.org/10.1017/S0003055420000374

Related articles

best-practices

Does Anonymity Make Students More Honest? Social Desirability and Course Feedback

Joinson (1999) showed people report more candidly when anonymous and online. What the social-desirability evidence means for whether your course evaluations capture honest student views, and the tension between candour and accountability.

research-methods

Asking the Questions Students Won't Answer Honestly: The Randomized Response Technique for Sensitive Course-Evaluation Items

When a course or climate survey asks about harassment, discrimination, academic misconduct, or truancy, ordinary anonymity is not enough — students still under-report. The randomized response technique adds provable privacy so respondents can answer honestly. What Warner (1965) proposed, what the validation evidence shows, and where it fails.

research-methods

Does a Conversational Course Evaluation Make Students Less Honest? Mode Effects and Social Desirability

Survey mode shapes honesty: interviewer-administered surveys invite more social-desirability bias than self-administered ones. What Tourangeau and Yan (2007) and the mode-effects literature mean for anonymous, AI-moderated conversational course evaluations.

research-methods

Total Survey Error: The Framework That Connects Every Course-Evaluation Quality Decision

The Total Survey Error framework organises the whole zoo of course-evaluation biases — coverage, sampling, nonresponse, measurement, processing — into one map, and tells you where to spend your limited effort for the biggest gain in accuracy.