Asking the Questions Students Won't Answer Honestly: The Randomized Response Technique for Sensitive Course-Evaluation Items
When a course or climate survey asks about harassment, discrimination, academic misconduct, or truancy, ordinary anonymity is not enough — students still under-report. The randomized response technique adds provable privacy so respondents can answer honestly. What Warner (1965) proposed, what the validation evidence shows, and where it fails.
Koji Education Team
Product
In brief: Some questions a university genuinely needs answered — did an instructor behave inappropriately, did you experience discrimination, did you actually attend, did you use unauthorised help — are exactly the questions students will not answer truthfully, even on an "anonymous" form, because the honest answer feels risky or shameful. The randomized response technique (RRT), introduced by Warner (1965), breaks this deadlock by injecting deliberate, known randomness into the answering process: a coin flip or die roll determines whether the student answers the sensitive question or a harmless dummy one, so no individual answer can ever be pinned to the sensitive behaviour, yet the population rate can still be estimated exactly because the mechanism's probabilities are known. A validation meta-analysis (Lensvelt-Mulders et al., 2005) found RRT yields more valid prevalence estimates than direct questioning — though later work (John et al., 2018) shows it is no panacea. RRT is the tool for the handful of course- and climate-survey items where ordinary anonymity is not enough.
The problem ordinary anonymity does not solve
Most course-evaluation items are not sensitive: "the lecturer explained concepts clearly" carries no social risk. But quality-assurance and student-experience surveys increasingly need to ask items that do: experiences of harassment or discrimination, whether a placement supervisor behaved unprofessionally, academic-integrity behaviours, or even mundane-but-embarrassing facts like non-attendance. For these, the promise of anonymity — however sincere — is not enough. Students reason, correctly, that a small class plus a few demographic fields can make them re-identifiable, and that "anonymous" data still passes through institutional systems. The result is social-desirability bias on a scale that ordinary anonymity assurances only partly reduce: sensitive behaviours are under-reported and desirable ones over-reported, so the prevalence numbers a committee relies on are simply wrong.
RRT attacks the problem at its root. Instead of promising privacy and asking the respondent to trust the institution, it guarantees privacy mathematically, so that answering honestly costs the student nothing — because not even the researcher holding the raw data can tell what any individual actually did.
What the research says
Warner's original design (1965). Stanley Warner, writing in the Journal of the American Statistical Association, proposed a device that sounds paradoxical and turns out to be rigorous. The respondent uses a private randomiser — a spinner, coin, or die the surveyor never sees — that with known probability p directs them to answer statement A ("I have done X") and with probability 1−p its negation ("I have not done X"). The interviewer records only "yes" or "no" and never learns which statement was answered. Because p is known, the observed proportion of "yes" answers is a simple linear function of the true prevalence, which can be algebraically recovered for the group. The individual is protected absolutely; the aggregate is estimated without bias. Later variants (the "unrelated question" design, forced-response designs) refined the idea to make it easier for respondents to understand and to improve statistical efficiency.
Does it actually work? Lensvelt-Mulders, Hox, van der Heijden, and Maas (2005) conducted the definitive validation: two meta-analyses covering thirty-five years of RRT studies, including rare "known-truth" validation studies where the true individual status could be checked against records. Their conclusion was that RRT produces more valid population estimates of sensitive behaviours than conventional question formats — the more sensitive the topic, the larger RRT's advantage — while noting that RRT still tends to under-estimate true prevalence somewhat (it reduces, rather than abolishes, evasive answering).
Modern design and analysis. Blair, Imai, and Zhou (2015) put RRT on a rigorous modern footing, deriving efficient estimators, showing how to incorporate covariates through multivariate regression on randomized-response data, and providing open-source tools. Their work matters for evaluation because it lets you go beyond a single prevalence number — for example, estimating whether the rate of a sensitive experience differs across programmes or years while preserving each respondent's protection.
The honest limitation. John, Loewenstein, Acquisti, and Vosgerau (2018), in Organizational Behavior and Human Decision Processes, provide the essential counterweight: RRT often fails to elicit the truth because respondents who do not understand or trust the mechanism default to the safe answer (denying the sensitive behaviour) rather than following the randomiser. When comprehension is low or trust is absent, RRT can perform no better than — and sometimes worse than — a direct question. The technique's guarantee is only as good as the respondent's belief in it.
Why it matters for course evaluation in practice
RRT is a specialist tool, not a replacement for the ordinary Likert form. Used surgically, it enables things a standard survey cannot.
-
Credible prevalence estimates for climate and integrity items. When a faculty needs a defensible figure for the rate of harassment, discrimination, or academic misconduct — the kind of number that appears in an accreditation self-evaluation or an equality report — RRT provides an estimate that withstands the objection "students were too afraid to tell you the truth."
-
Reducing under-report on embarrassing-but-benign items. Non-attendance, not doing the reading, or using unauthorised study aids are all systematically under-reported. Where these matter for interpreting other results, an RRT item gives a truer base rate.
-
A privacy story students can verify. Unlike an anonymity promise, RRT's protection is something a student can reason about themselves — the coin, not the institution, decides. That transparency can rebuild trust with cohorts who have learned to give guarded answers, a problem related to demand characteristics.
-
Aggregate-only by design. RRT is intrinsically incapable of producing individual-level sensitive data, which aligns neatly with data-minimisation obligations: you literally cannot leak what you never collected.
Limitations and honest caveats
A PhD reader will — rightly — treat RRT with caution, and so should any QA office adopting it.
-
Statistical cost. The injected randomness is noise, so RRT estimates have much larger variance than direct questions. You need substantially larger samples to reach the same precision — often two to four times as many respondents — which makes RRT impractical for small classes and best suited to programme-, faculty-, or institution-level surveys.
-
Comprehension and trust are prerequisites, not givens. As John et al. (2018) show, the whole edifice collapses if respondents do not understand or believe the mechanism. RRT requires clear instructions, ideally a practice item, and a genuinely credible randomiser; deployed carelessly it adds noise without adding honesty.
-
It answers "how many," not "who" or "why." RRT yields a protected prevalence rate and nothing else. It cannot support follow-up, cannot be linked to other responses at the individual level, and produces no qualitative detail about the experiences behind the number.
-
It only fits binary or simple categorical items. RRT is designed for yes/no sensitive attributes. It does not translate to the graded, multidimensional judgements that make up most of teaching evaluation.
-
Alternatives may fit better. The related list experiment (item-count technique) and the crosswise model address the same problem with different trade-offs and are sometimes easier for respondents. RRT is one option in a small family of sensitive-question methods, not the only one.
Because of these costs, RRT belongs on the two or three items where honesty genuinely cannot be assumed — not on a whole questionnaire.
How Koji incorporates this
Koji's design philosophy — remove the reasons a respondent has to shade the truth — is closely aligned with the goal RRT serves, and Koji addresses it through complementary means.
-
Structured item types that can carry a protected design. Koji's
yes_noandsingle_choicequestion objects are the natural container for a forced-response or unrelated-question RRT item, with the randomising instruction delivered in the question text and the known probabilities recorded for later estimation. Because Koji stores responses against defined objects, the aggregate recovery that RRT requires is straightforward to compute. -
Trust built through experience, not just assurance. The failure mode identified by John et al. (2018) is broken trust. Koji's AI-moderated conversational interviews are designed to establish rapport and explain why a question is being asked, which is precisely the comprehension-and-trust foundation an RRT item needs to work — students who understand the protection are the ones who use it honestly.
-
A qualitative route where RRT cannot go. RRT gives a number and stops. For the why behind a sensitive prevalence rate, Koji's AI-moderated interviews can explore experiences in a respondent-controlled, non-judgemental way, and its automatic thematic analysis surfaces patterns across open text — the depth RRT structurally lacks. Used together, RRT-style items for defensible rates and conversational interviews for meaning cover both halves of a sensitive topic.
-
Data minimisation by architecture. Koji's anonymity and re-identification safeguards reinforce the same principle RRT embodies: the safest sensitive data is the individual-level data you never hold.
Koji does not claim to make students unconditionally honest, and RRT is not a built-in wizard — deploying it well requires design care and adequate sample sizes. What Koji provides is the structured, trust-oriented collection environment in which sensitive-question methods have the best chance of working. (Koji's core research platform at koji.so applies the same rapport-building interview engine to sensitive product and customer research, from pricing honesty to reports of workplace experience.)
Related Resources
- Does Anonymity Make Students More Honest? Social Desirability and Course Feedback
- Does a Conversational Course Evaluation Make Students Less Honest? Mode Effects and Social Desirability
- Demand Characteristics: When Students Guess What Your Course Evaluation Wants to Hear
- Can a Student Be Re-Identified From Their Course Feedback? k-Anonymity and GDPR
- Abusive Open-Text Comments in Course Evaluations: The Evidence and the Duty of Care
References
- Warner, S. L. (1965). Randomized response: A survey technique for eliminating evasive answer bias. Journal of the American Statistical Association, 60(309), 63–69. https://doi.org/10.1080/01621459.1965.10480775
- Lensvelt-Mulders, G. J. L. M., Hox, J. J., van der Heijden, P. G. M., & Maas, C. J. M. (2005). Meta-analysis of randomized response research: Thirty-five years of validation. Sociological Methods & Research, 33(3), 319–348. https://doi.org/10.1177/0049124104268664
- Blair, G., Imai, K., & Zhou, Y.-Y. (2015). Design and analysis of the randomized response technique. Journal of the American Statistical Association, 110(511), 1304–1319. https://doi.org/10.1080/01621459.2015.1050028
- John, L. K., Loewenstein, G., Acquisti, A., & Vosgerau, J. (2018). When and why randomized response techniques (fail to) elicit the truth. Organizational Behavior and Human Decision Processes, 148, 101–123. https://www.sciencedirect.com/science/article/abs/pii/S0749597817300523
Related articles
Abusive Open-Text Comments in Course Evaluations: The Evidence and the Duty of Care
Lakeman et al. (2022) found that 91% of surveyed Australian academics received non-constructive anonymous comments — insults, threats, remarks on appearance. A research-grounded look at abusive student feedback, its effect on staff wellbeing, and how to moderate open text responsibly.
Does Anonymity Make Students More Honest? Social Desirability and Course Feedback
Joinson (1999) showed people report more candidly when anonymous and online. What the social-desirability evidence means for whether your course evaluations capture honest student views, and the tension between candour and accountability.
Can a Student Be Re-Identified From Their Course Feedback? Small-Class Anonymity, k-Anonymity and GDPR
"Anonymous" course evaluations in small classes often are not. What re-identification research and GDPR actually require — and how to report small-cohort feedback without breaching either confidentiality or law.
Does a Conversational Course Evaluation Make Students Less Honest? Mode Effects and Social Desirability
Survey mode shapes honesty: interviewer-administered surveys invite more social-desirability bias than self-administered ones. What Tourangeau and Yan (2007) and the mode-effects literature mean for anonymous, AI-moderated conversational course evaluations.