Are Students Telling You What You Want to Hear? Social Desirability Bias in Course Evaluation
Even anonymous course evaluations are shaped by what students think they ought to say. Here is what the evidence shows about social desirability bias, why anonymity only partly fixes it, and how to design evaluation that gets closer to the truth.
Koji Education Team
Product ·
Bottom line up front: Social desirability bias is the tendency for respondents to answer in ways that present themselves, or the person they are rating, in a socially acceptable light rather than truthfully. In course evaluation it pushes ratings toward the polite and the positive — students soften criticism of an instructor they like personally, withhold negative comments they fear could be traced back to them, or simply give the "expected" answer to get through the form. Anonymity helps but does not eliminate the effect, because the pressure is partly internal: people give socially approved answers even when no one is watching. The practical fix is not a better Likert scale but a different mode of collecting feedback — one that builds enough trust and specificity that honesty becomes the path of least resistance.
What social desirability bias actually is
Social desirability bias is one of the oldest known threats to self-report data. It describes a systematic distortion: respondents over-report attitudes and behaviours they believe are approved of, and under-report those they think are frowned upon. The classic measurement instrument, the Marlowe-Crowne Social Desirability Scale, has been used since 1960 to detect respondents who are "faking good." The Wikipedia overview of social-desirability bias is a useful entry point, but the phenomenon is documented across decades of survey methodology.
In course evaluation, the bias has a specific texture. Students are not usually trying to deceive. They are navigating a social situation: a person they have spent a semester with is asking, in effect, "How did I do?" Most people are reluctant to deliver a blunt negative verdict on a named individual, especially one with power over their grades. So the rating drifts upward, the open-text box gets a vague "Good course, thanks!", and the genuinely useful criticism — the thing that would actually improve teaching — stays unwritten.
Why this is not the same as the other biases you already know
Readers of this blog will recognise that we have covered negativity bias, central tendency bias, and the halo effect. Social desirability is distinct from all three. Negativity bias is about a few harsh voices being over-weighted; social desirability pushes the opposite way, suppressing criticism. Central tendency is about avoiding the extremes of a scale; social desirability is about choosing the approved end of it. The halo effect is a cognitive shortcut; social desirability is a social-impression strategy. They can co-occur — a student who likes an instructor (halo) may also feel extra reluctant to criticise them (social desirability) — but they have different causes and different remedies.
Does anonymity solve it? The honest answer is "only partly"
The intuitive fix is anonymity: if students cannot be identified, surely they will speak freely. The evidence is more nuanced than that hope.
A well-designed study by Scherer, Straub, Schnyder and Schaffner (2013), published in the International Journal of Nursing Education Scholarship, compared 306 anonymous and 309 personalised (named) student evaluation forms in a Bachelor nursing programme. The striking result: there were no significant differences in the quantitative ratings between the anonymous and the identified groups. Students who signed their names did not inflate their scores to avoid offending instructors. The authors concluded that personalised evaluations "do not generate more biased results in terms of social desirability" — provided the process is transparent, results are reported back to students, and there is room for open discussion.
That finding cuts two ways. On one hand, it is reassuring that removing anonymity does not automatically wreck the data. On the other hand, it punctures the comforting assumption that anonymity removes social desirability. If anonymous and named ratings look the same, and we know social desirability operates on named responses, then the anonymous responses are carrying a similar load. The pressure to give the approved answer is partly internalised — a norm people apply to themselves. This matches the broader survey-methodology consensus, summarised in reviews such as this educational-research overview, that anonymity is only minimally effective at removing misreporting on its own.
Critics argue: "Isn't this just speculation? Most students are honest."
It is a fair challenge, and intellectual honesty requires us to state the limits of the evidence. Social desirability bias is hard to measure directly in course evaluation, precisely because the "true" answer is unobservable. Much of what we know is inferred from response patterns, from controlled studies in adjacent domains, and from the Scherer et al. result above. We should resist the temptation to attribute every positive rating to dishonesty — many courses genuinely are good, and ceiling effects have several causes beyond social pressure (we cover those in our piece on why almost every course scores four out of five).
So the claim is deliberately modest. We are not saying student feedback is worthless or that students are liars. We are saying that the quiet, polite suppression of criticism is a real and well-documented mechanism, that it biases data in a predictable direction, and that a one-shot anonymous Likert form does little to counter it. The goal is not cynicism. It is to design feedback that makes honesty easier than politeness.
What actually reduces social desirability bias
The methodological literature points to a consistent set of levers, none of which is "add another rating question":
- Specificity over global judgement. It is socially hard to say "the instructor was bad." It is much easier to describe a concrete moment: "I got lost when we jumped from chapter three to chapter six without the intermediate steps." Specific, behaviour-anchored prompts lower the social stakes.
- Genuine confidentiality, communicated credibly. Students will only believe their feedback is safe if you tell them clearly how it is handled and when instructors see it. The Scherer study suggests trust in the process matters more than the mechanical presence or absence of a name field.
- Demonstrated action — closing the loop. When students see that previous feedback led to change, the exercise stops feeling like a hollow courtesy and starts feeling consequential. That shift changes what they are willing to say. We explore this in depth in the action gap.
- A conversational mode that probes. A static form accepts the first polite answer and moves on. A good interviewer notices the vague "it was fine" and gently asks, "Fine in what way? Was there anything that didn't work for you?"
Where Koji fits
This last lever is where Koji for Education is built differently from a legacy survey tool. Instead of a one-shot form that records whatever a student types first, Koji runs an AI-moderated conversational interview. When a student gives a thin, socially safe answer, the AI follows up — asking for a concrete example, surfacing the thing the student was reluctant to lead with. Because the moderation is standardised, every student gets the same calm, neutral probing, with none of the variability of a human interviewer who might inadvertently signal approval or disapproval.
Koji also leans on the levers the evidence supports. Its six structured question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) let you anchor questions in specifics rather than global verdicts. Its automatic thematic analysis reads across thousands of open-text responses to find the criticism that is there but understated, rather than reducing everything to an average. And its GDPR- and AVG-compliant handling lets you make a credible, specific confidentiality promise — the kind students actually believe. Koji is careful to claim only what it can deliver: it reduces and surfaces social-desirability effects; it does not eliminate them, because no instrument can fully remove a pressure that lives partly inside the respondent.
The same conversational interview engine powers the main Koji platform for general customer and user research, where social desirability is an equally well-known threat to honest feedback — so the techniques are battle-tested well beyond the classroom.
The takeaway for quality-assurance teams
Treat your positive scores with the same scepticism you apply to your negative ones. A 4.6 average is not proof of excellence any more than a single angry comment is proof of failure. Both are filtered through what students felt able to say. If you want feedback that reflects what students actually think — including the criticism that would help you improve — the lever to pull is the mode and trustworthiness of collection, not the wording of the rating scale. That is the case for moving from static surveys to conversational, evidence-aware evaluation.
Koji for Education helps European universities collect course feedback that gets past the polite answer. See how Koji works.