New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods10 min read

Does a Conversational Course Evaluation Make Students Less Honest? Mode Effects and Social Desirability

Survey mode shapes honesty: interviewer-administered surveys invite more social-desirability bias than self-administered ones. What Tourangeau and Yan (2007) and the mode-effects literature mean for anonymous, AI-moderated conversational course evaluations.

Koji Education Team

Product

Answer (BLUF): Whether a course evaluation is self-administered (an anonymous web form) or interviewer-administered (a human asking questions) changes how honestly students answer. The survey-methods evidence — anchored by Tourangeau and Yan's (2007) review and Richman et al.'s (1999) meta-analysis — is consistent: a visible human interviewer increases social-desirability bias on sensitive items, while self-administration improves honest reporting. The key design question for AI-moderated conversational evaluation is therefore not "interview vs survey" but "is the interviewer a person or an anonymous, non-judgmental computer?" An anonymous text-based AI interview is, on this evidence, closer to self-administration than to a human interviewer — so it can keep the candour of an anonymous form while adding the depth of a conversation.

What the research says

Survey methodologists distinguish two broad administration modes. In interviewer-administered surveys (face-to-face, telephone), a human poses the questions and records answers. In self-administered surveys (paper, web), the respondent answers privately with no human present. Decades of research show these modes do not yield the same answers to sensitive questions.

Roger Tourangeau and Ting Yan (2007), in their authoritative Psychological Bulletin review Sensitive Questions in Surveys, synthesise this literature. They define three features that make a question sensitive: it can be intrusive (inappropriate in ordinary conversation), it can carry a threat of disclosure (the respondent fears consequences if a third party learned the answer), and it can be socially undesirable (the truthful answer admits violating a norm). Their central conclusion is that misreporting on sensitive topics is common but largely situational — it depends on whether the respondent has something embarrassing to report and on design features of the survey. Crucially, self-administration reduces bias: relative to an interviewer, private modes improved the quality of reports about mental-health symptoms, reduced over-reporting of socially approved behaviours (such as attending religious services), and improved reporting of sensitive behaviours.

The mechanism is the felt presence of another judging mind. William Richman and colleagues (1999) made this precise in a meta-analysis of 61 studies and 673 effect sizes comparing computer questionnaires, paper questionnaires, and face-to-face interviews. The overall computer-versus-paper difference was near zero — but there was less social-desirability distortion on computerised instruments than in face-to-face interviews, and distortion shrank when respondents were alone, anonymous, and able to backtrack. In other words, removing the human observer, not the technology itself, is what protects candour.

A third study sharpens the practical implication. Kreuter, Presser and Tourangeau (2008), comparing interviewer-led telephone interviews (CATI), interactive voice response (IVR), and web surveys, found that the more self-administered the mode, the more willing respondents were to disclose sensitive or unflattering information. The web condition elicited more honest reporting of undesirable items than the interviewer condition.

For balance, the effect is not unbounded. Dodou and de Winter (2014), in a meta-analysis pointedly titled Social desirability is the same in offline, online, and paper surveys, found that for many ordinary (non-highly-sensitive) measures the social-desirability gap between modes was small. The honest synthesis: interviewer presence reliably increases social-desirability bias on genuinely sensitive items, but the penalty is modest when the content is not very sensitive.

Why it matters for course evaluation in practice

Course evaluation sits in an interesting middle zone of sensitivity. Rating a lecturer is rarely "intrusive," but it can carry a real threat of disclosure: students worry — rightly or not — that critical feedback could be traced back and affect their grades or relationship with staff. That is exactly the dimension Tourangeau and Yan identify as most responsive to design.

Several practical conclusions follow:

  • Any human in the loop suppresses criticism. Evaluations collected by a tutor in the room, by a programme rep reading questions aloud, or in a focus group with peers present will skew positive. The student is performing for an audience. This is a mode effect, not a measurement of better teaching.
  • Anonymous self-administration is the candour baseline. The widespread shift to anonymous online forms was, in mode-effects terms, the right move: it removed the interviewer and lowered the threat of disclosure. The trade-off has been thin, monosyllabic open-text and falling response rates (see survey fatigue).
  • "Conversational" evaluation raises a legitimate question. If you replace a static form with an interview, do you re-import the interviewer-presence penalty? The literature says: only if the interviewer is a judging human. An anonymous, non-human, text-based interviewer that cannot report you and expresses no approval or disapproval should behave far more like self-administration than like a person in the room.
  • The honesty of any mode still rests on perceived anonymity. No mode is candid if students do not believe their answers are confidential — which is why mode design and the anonymity vs confidentiality question must be solved together.

Limitations & honest caveats

A PhD reader will rightly push back on a clean "AI interview = self-administration" claim.

  • Direct evidence on AI interviewers is still thin. The canonical studies compared humans, paper, telephone, and web forms — not large-language-model interviewers. Whether students psychologically treat a conversational AI as "no one is watching" or as "something is listening and judging" is an empirical question that the classic literature only lets us reason toward, not settle.
  • Anthropomorphism could cut either way. A warm, conversational agent might increase rapport and disclosure, or it might trigger mild evaluation apprehension that a cold radio-button form does not. Richman's moderators (alone, anonymous, able to revise) suggest the design details — visible anonymity, no human notification, freedom to edit — matter more than the conversational surface.
  • Course evaluation is only moderately sensitive. Per Dodou and de Winter (2014), mode penalties are small for non-sensitive content, so the upside of any mode change on candour may be modest in the average case and larger only where students feel genuinely at risk.
  • Sensitivity is heterogeneous. A general "rate the lecturer" item is low-threat; a question about harassment, accessibility failures, or discrimination is high-threat, and that is precisely where self-administration and ironclad anonymity matter most.

These caveats argue for humility in the marketing claim, not for abandoning conversational evaluation. They define the conditions under which it preserves honesty.

How Koji incorporates this

Koji is built so that a conversation does not reintroduce interviewer bias.

  • Anonymous, non-human, non-reporting interviewer. Koji's AI moderator has no social stake in the answer, cannot judge the student, and cannot inform staff who said what. By the Richman (1999) logic — distortion falls when respondents are alone and anonymous — this positions the conversation on the self-administration side of the mode divide, not the human-interviewer side.
  • Probing depth without a person in the room. The classic trade-off was candour (self-administered) versus depth (interviewer-administered). Koji's AI-moderated interview is designed to capture both: it follows up on a vague answer the way a skilled interviewer would, but without the judging human presence that the literature shows suppresses honest criticism.
  • Visible, repeated anonymity assurances and editable answers. Because perceived anonymity and the ability to revise are the moderators that protect candour, Koji surfaces anonymity guarantees in the flow and lets respondents reconsider — directly targeting Tourangeau and Yan's "threat of disclosure" dimension.
  • Sensitivity-aware handling. For higher-threat themes (discrimination, accessibility, wellbeing), Koji is designed to keep the exchange private and non-attributable, where the self-administration advantage is largest. The aim is to mitigate social-desirability bias, not to claim it is eliminated.

The same conversational engine underpins Koji's core research platform at koji.so, where the identical mode logic applies to customer and employee research: an anonymous AI interviewer is designed to elicit the candid answers a human moderator or a focus group would suppress.

A practical rule of thumb

The operational test is simple: ask "could a student plausibly fear that an identifiable human will read this and react?" Wherever the answer is yes — a tutor collecting forms, a programme representative facilitating a session, a named staff member visibly running the interview — expect social-desirability inflation and discount glowing results accordingly. Wherever the channel is genuinely anonymous and no human is notified of who said what, the candour penalty largely dissolves, conversational or not. The format of the questioning matters far less than two things: the presence or absence of a judging human, and the student''s belief in their own anonymity. Design for both and a conversation becomes an asset rather than a liability.

Related Resources

References

  1. Tourangeau, R., & Yan, T. (2007). Sensitive Questions in Surveys. Psychological Bulletin, 133(5), 859–883. https://doi.org/10.1037/0033-2909.133.5.859
  2. Richman, W. L., Kiesler, S., Weisband, S., & Drasgow, F. (1999). A Meta-Analytic Study of Social Desirability Distortion in Computer-Administered Questionnaires, Traditional Questionnaires, and Interviews. Journal of Applied Psychology, 84(5), 754–775. https://doi.org/10.1037/0021-9010.84.5.754
  3. Kreuter, F., Presser, S., & Tourangeau, R. (2008). Social Desirability Bias in CATI, IVR, and Web Surveys: The Effects of Mode and Question Sensitivity. Public Opinion Quarterly, 72(5), 847–865. https://doi.org/10.1093/poq/nfn063
  4. Dodou, D., & de Winter, J. C. F. (2014). Social desirability is the same in offline, online, and paper surveys: A meta-analysis. Computers in Human Behavior, 36, 487–495. https://doi.org/10.1016/j.chb.2014.04.005

Related articles