Does Anonymity Make Students More Honest? Social Desirability and Course Feedback
Joinson (1999) showed people report more candidly when anonymous and online. What the social-desirability evidence means for whether your course evaluations capture honest student views, and the tension between candour and accountability.
Koji Education Team
Product
In short: Students do not always say what they think. Social-desirability bias — the tendency to give answers that look acceptable rather than truthful — suppresses candid criticism in course evaluations, and Joinson (1999) showed experimentally that anonymity and self-administered online formats reduce it: anonymous web respondents reported the lowest social-desirability scores, while non-anonymous pen-and-paper respondents reported the highest. The practical implication is that perceived anonymity is a precondition for honest feedback, but it sits in tension with the institution's need to detect abuse, follow up, and act — a tension that must be designed for, not ignored.
What the research says
Social-desirability bias is the well-documented tendency of survey respondents to present themselves favourably — over-reporting socially approved attitudes and under-reporting disapproved ones. In a course evaluation, it pushes students toward bland, agreeable answers: reluctance to criticise a likeable but ineffective instructor, softened complaints, and the omission of genuine but awkward problems.
The anchor study is Adam Joinson's "Social desirability, anonymity, and Internet-based questionnaires" (Behavior Research Methods, Instruments, & Computers, 1999, 31(3), 433–438; DOI: 10.3758/BF03200723). Joinson had participants complete measures of self-esteem, social anxiety, and social desirability either on the web or on paper, and either anonymously or non-anonymously. The result was a clean crossed pattern: respondents reported lower social desirability and lower social anxiety, and higher self-esteem, when anonymous than when identified, and lower social desirability on the web than on paper. The lowest social-desirability scores came from the anonymous-web condition; the highest from the non-anonymous-paper condition. In short, people answered more candidly when they felt unidentifiable and were not face-to-face with an administrator.
This is not a one-off. A meta-analytic review by Timo Gnambs and Kai Kaspar, "Socially Desirable Responding in Web-Based Questionnaires: A Meta-Analytic Review of the Candor Hypothesis" (Assessment, 2017, 24(6), 746–762; DOI: 10.1177/1073191115624547), found that self-administered web surveys tend to elicit more honest reporting of sensitive information than interviewer-administered or paper modes, especially for highly sensitive content. The broader theoretical account is Roger Tourangeau and Ting Yan's "Sensitive questions in surveys" (Psychological Bulletin, 2007, 133(5), 859–883; DOI: 10.1037/0033-2909.133.5.859), which explains that self-administration, the absence of an interviewer, and credible confidentiality all reduce socially desirable responding to sensitive items.
Applied to teaching evaluation, the relevant fear is identifiability: students worry — sometimes realistically — that candid criticism could be traced back to them and affect their grade or relationship with staff. That fear is itself a driver of social-desirability bias, and it interacts with the selection effects discussed in our note on non-response and selection bias.
Why it matters for course evaluation in practice
Three things follow for a quality-assurance office.
First, perceived anonymity is a data-quality lever, not just an ethics checkbox. If students do not believe their responses are confidential, the evaluation systematically under-reports problems — and an institution that reads bland, positive results as "no issues" may simply be measuring its own students' caution. Visible, credible anonymity guarantees (small-cohort suppression, no instructor access to raw responses until grades are submitted) are part of getting valid data, not merely part of compliance.
Second, mode and timing matter. Joinson and the candour literature favour self-administered, screen-based collection over face-to-face or in-class paper, where the instructor presence and peer visibility heighten social pressure. Timing relative to grade release also shapes how safe students feel being candid.
Third, there is a genuine tension with accountability. Total anonymity maximises candour but removes the ability to follow up with an individual, to detect coordinated grade-bombing, or to act on a safeguarding disclosure. It can also lower the cost of abusive comments (see abusive open-text comments and duty of care). The right design is not "maximum anonymity" but credible confidentiality with proportionate safeguards — enough protection that students speak freely, enough structure that the institution can respond responsibly.
Limitations and honest caveats
The evidence supports anonymity-as-candour, but a careful reader should note the boundaries.
Joinson 1999 is a lab study with general psychological measures, not a course-evaluation field experiment. The constructs (self-esteem, social anxiety) are not teaching ratings, so the transfer is by analogy. The mechanism — anonymity reduces self-presentation — is general and well replicated, but the precise size of the effect on teaching evaluations specifically is less firmly established, and we should not over-claim a number.
More candour is not automatically more accuracy. Anonymity can also lower accountability for the respondent, which may increase careless responding, exaggeration, or venting. Candid is not the same as calibrated; the literature on satisficing is the counterweight. The goal is honest and considered feedback, which anonymity alone does not guarantee.
Anonymity can be illusory in small cohorts. In a seminar of eight students, "anonymous" free text plus demographic fields can be effectively identifying. Promising anonymity you cannot deliver is worse than not promising it, both ethically and for trust.
Web-mode candour effects vary by sensitivity and population. Gnambs and Kaspar find the candour advantage is largest for sensitive topics; for innocuous items the difference shrinks. Cultural norms about criticising authority also moderate how freely students will speak regardless of mode (see cross-cultural response styles).
How Koji incorporates this
Koji for Education is designed to capture candid feedback while handling the candour-versus-accountability tension deliberately — described as mitigation and good practice, not a guarantee of perfect honesty.
- Confidential, self-administered conversational collection. Koji evaluations are completed by the student on their own screen, in their own time — the self-administered, no-administrator-present mode the candour literature favours over in-class paper. The AI moderator is a neutral, non-judgemental interlocutor, which lowers the social-presentation pressure a student feels when criticising a person who will grade them.
- Credible anonymity and small-cohort protection. Reporting is built to protect identifiability — aggregating and suppressing results below safe response thresholds (the same logic as our note on how many responses you need) — so the confidentiality promise students need in order to be candid is one the platform can actually keep.
- AI moderation that invites honest detail without exposure. Because the moderator probes conversationally (Can you say more about what did not work?), it draws out the candid specifics that a one-line anonymous box rarely captures, while never asking the student to identify themselves to the instructor.
- Bias-aware reporting and quality scoring help distinguish considered criticism from venting or careless responding, addressing the honest caveat that anonymity can raise noise as well as candour.
- Closing-the-loop action tracking connects candid feedback to institutional response in a way that respects confidentiality, so students see that speaking up leads to change — which, over time, is what sustains honest participation.
Koji does not pretend anonymity makes every answer truthful; the research says candour rises, not that bias vanishes. The design intent is to remove the identifiability fears that suppress honest feedback while keeping the proportionate safeguards an institution needs. Koji core research platform at koji.so applies the same confidential, AI-moderated approach to customer and employee research, where social-desirability bias is a constant threat to honest signal.
A practical checklist for confidentiality and candour
Designing for honest feedback means treating anonymity as a data-quality requirement and being scrupulously honest with students about what you can and cannot protect.
- Promise only the confidentiality you can deliver. In small cohorts, free text plus demographic fields can be identifying; suppress or aggregate results below a safe response threshold rather than claiming an anonymity you cannot guarantee.
- Collect self-administered, not in-class on paper. The candour literature consistently favours screen-based, no-administrator-present completion over face-to-face or in-room paper, where instructor presence and peer visibility heighten social pressure.
- Separate raw responses from the instructor until grades are submitted. The fear that candid criticism could affect a grade is itself a driver of social-desirability bias; visible safeguards against it improve the data, not just the ethics.
- Design proportionate safeguards, not maximal anonymity. Retain the ability to act on safeguarding disclosures and to detect coordinated grade-bombing, and pair anonymity with moderation that filters abuse.
- Close the loop visibly. When students see that candid feedback leads to change, honest participation becomes self-sustaining; when they see nothing happen, caution and disengagement return.
The aim is a system in which a student believes it is both safe and worthwhile to tell the truth — the precondition the social-desirability evidence says you need for valid feedback.
Related Resources
- Online vs. Paper Course Evaluations: Do Lower Response Rates Mean Worse Data?
- Selection Bias in Course Evaluations
- Survey Fatigue and Response Rates
- Abusive Open-Text Comments and the Duty of Care
- How Many Responses Do You Need for a Reliable Course Evaluation?
- Response Styles and Likert Scales: Cross-Cultural Evaluation
References
- Joinson, A. (1999). Social desirability, anonymity, and Internet-based questionnaires. Behavior Research Methods, Instruments, & Computers, 31(3), 433–438. https://doi.org/10.3758/BF03200723
- Gnambs, T., & Kaspar, K. (2017). Socially desirable responding in web-based questionnaires: A meta-analytic review of the candor hypothesis. Assessment, 24(6), 746–762. https://doi.org/10.1177/1073191115624547
- Tourangeau, R., & Yan, T. (2007). Sensitive questions in surveys. Psychological Bulletin, 133(5), 859–883. https://doi.org/10.1037/0033-2909.133.5.859
- Paulhus, D. L. (1984). Two-component models of socially desirable responding. Journal of Personality and Social Psychology, 46(3), 598–609. https://doi.org/10.1037/0022-3514.46.3.598
Related articles
How Many Responses Do You Need for a Reliable Course Evaluation?
Reliability of a class-average rating rises steeply with the number of students who respond. Generalizability theory — from Marsh''s syntheses to Dzakadzie (2026) — shows raters matter more than items, and a class of 12 cannot be evaluated like a class of 120.
Survey Fatigue: Why Over-Surveying Students Quietly Wrecks Your Response Rates
Porter, Whitcomb & Weitzer (2004) showed that administering multiple surveys in one year suppresses later response rates. A research-grounded guide to survey fatigue in course evaluation — what causes it, what the evidence shows, and how to design around it.
Abusive Open-Text Comments in Course Evaluations: The Evidence and the Duty of Care
Lakeman et al. (2022) found that 91% of surveyed Australian academics received non-constructive anonymous comments — insults, threats, remarks on appearance. A research-grounded look at abusive student feedback, its effect on staff wellbeing, and how to moderate open text responsibly.
Response Styles and Likert Scales: Why Cross-Cultural Evaluation Needs More Than Numbers
Acquiescence and extreme response styles vary systematically by culture (Harzing, 2006; Baumgartner & Steenkamp, 2001), which means raw Likert averages are not directly comparable across nationalities in Europe's multinational classrooms. This article explains the evidence and how to evaluate fairly across diverse cohorts.