New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods9 min read

Demand Characteristics: When Students Guess What Your Course Evaluation Wants to Hear

Orne''s concept of demand characteristics explains why students who infer what an evaluation is "for" answer to fit it. What the reactivity evidence says, and how to design evaluations that measure experience rather than compliance.

Koji Education Team

Product

In short: A demand characteristic is any cue in an evaluation that tells respondents what answer is expected of them. Michael Orne''s foundational 1962 work showed that people readily infer the "hypothesis" behind a task and adjust their behaviour to fit it — and Nichols and Maner (2008) demonstrated experimentally that a majority of participants shift responses toward a hypothesis once they guess it. In course evaluation, leading item wording, a visible institutional agenda, and instructors coaching students on how to answer all function as demand characteristics that make ratings measure compliance rather than experience. You cannot eliminate them, but neutral wording, protected anonymity, and open probing that lets students set the agenda substantially reduce them.

The problem: students answer the question they think you are really asking

Every course evaluation carries an implicit message about what the institution wants to hear. Students are not passive measuring instruments. They read the form, infer its purpose, notice who is asking and why, and — often without deciding to — shape their answers to that inferred purpose. When that happens, the rating stops being a clean report of the student''s experience and becomes a report of what the student thinks the evaluation is for. This is the phenomenon of demand characteristics, and it is one of the oldest and best-documented threats to any self-report measure.

For quality-assurance offices, demand characteristics are more insidious than random noise. Noise averages out across a cohort; systematic pull toward an expected answer does not. If the instrument, the framing, or the administration nudges every respondent in the same direction, the bias survives aggregation and contaminates exactly the comparisons — instructor to instructor, year to year — that QA cares about most.

What the research says

The concept comes from Michael T. Orne''s 1962 article in American Psychologist, "On the social psychology of the psychological experiment: With particular reference to demand characteristics and their implications." Orne defined demand characteristics as "the totality of cues which convey an experimental hypothesis to the subject," and argued that the cues most powerful in shaping behaviour are those that "convey the purpose of the experiment effectively but not obviously." His central insight was that participants are active, motivated sense-makers: they treat a study as a problem to be solved, form a hypothesis about what is wanted, and — many of them — behave so as to confirm it. Orne called the cooperative version of this the "good subject" role.

For decades this was demonstrated indirectly. Nichols and Maner (2008), in the Journal of General Psychology, tested it head-on. A confederate told participants (N = 100) the study''s supposed hypothesis before they began a laboratory task. Participants then, on average, responded in ways that confirmed the hypothesis they had been given — and the strength of the effect depended on their attitude toward the experimenter and other individual differences. A secondary but practically important finding: standard "suspicion probes" (asking afterwards whether people guessed the purpose) were poor at detecting who actually knew. In other words, the pull is real, it is moderated by rapport and motivation, and you cannot reliably screen it out after the fact.

The mechanism generalises well beyond the psychology lab. Weber and Cook (1972), in Psychological Bulletin, catalogued the "subject roles" people adopt — the cooperative good subject, the apprehensive subject, the negativistic subject — and showed how each biases self-report differently. In experimental economics, Zizzo (2010) formalised "experimenter demand effects" as the change in behaviour caused by cues about what the experimenter considers appropriate, and showed they are strongest precisely when the task is ambiguous and the "right" answer is easy to infer. Course evaluations sit squarely in that danger zone: the questions are evaluative, the stakes for the instructor are visible, and the "socially expected" direction of a nice answer is obvious.

Two clarifications matter for reading this literature correctly. First, demand characteristics are distinct from social desirability. Social desirability is the pull toward answers that make the respondent look good; demand characteristics are the pull toward answers that fit the perceived purpose of the study. They frequently co-occur but have different fixes. Second, they are distinct from the Hawthorne effect and general observer reactivity: Hawthorne is about behaving differently because you are watched; demand characteristics are about behaving differently because you have decoded what is wanted. A student can be entirely unbothered by being observed and still steer their answers toward a hypothesis they have guessed.

Why it matters for course evaluation in practice

Concretely, demand characteristics enter a course-evaluation programme through at least four doors:

  • Leading and loaded item wording. "How much did this innovative teaching approach improve your learning?" presupposes that the approach was innovative and that it improved learning; it broadcasts the expected answer. Agree/disagree formats amplify this because acquiescence and demand pull in the same direction.
  • A visible institutional agenda. If the survey is introduced as evidence for a teaching-award nomination, a re-accreditation submission, or a contested course redesign, students infer the desired result and many will oblige — in either direction, depending on their sympathies.
  • Instructor coaching. When an instructor tells a class "these scores matter for my tenure case, please be generous," they have installed a demand characteristic by hand. Even a well-meant "be honest, it really helps me improve" can signal the hoped-for narrative.
  • Question order and context. Earlier items frame later ones. A block of warm rapport-building questions before the substantive items primes a cooperative "good subject" stance; a battery of complaint-shaped items primes fault-finding.

The consequence is a validity problem, not merely a precision problem. A demand-inflated score does not just have a wide confidence interval — it is centred in the wrong place, and it is centred in the wrong place consistently, which is why it can quietly reverse a fair comparison between two instructors who differ only in how they framed the survey to their students.

Limitations and honest caveats

A PhD reader will rightly push back on several points, and the honest position acknowledges them.

  • Effect sizes are context-dependent and often modest. Orne''s original demonstrations were vivid but came from settings (hypnosis research, sensory-deprivation studies) engineered to make demand salient. In routine, low-stakes, anonymous surveys the pull is usually smaller. Nichols and Maner had to tell participants the hypothesis to produce a clean effect; naturally occurring guessing is noisier.
  • Direction is not fixed. Demand characteristics do not always inflate. A "negativistic" or apprehensive subject, or a student who resents a mandated survey, may push against the perceived expectation. The bias is systematic in the sense of being non-random, but its sign depends on the respondent''s stance.
  • Detection is genuinely hard. Because suspicion probes are weak (Nichols and Maner''s own finding), you generally cannot measure how much demand contaminated a given dataset. Claims that a particular survey was "12% inflated by demand" should be treated with suspicion — the honest statement is that the risk was present and was designed against.
  • Some inferred purpose is legitimate. You want students to understand that the evaluation is about improving teaching. The goal is not to hide the purpose entirely (that raises its own ethical issues) but to avoid signalling a specific desired answer.

How Koji incorporates this

Koji for Education is designed to reduce the cues that let students decode a wanted answer, and to make the parts of a response that are hardest to fake more visible. The mechanisms map directly onto the sources of demand:

  • Neutral, item-specific wording by default. Koji''s question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) are templated toward specific, non-presupposing phrasing — "How clearly were the assessment criteria explained?" rather than "How much did this excellent course help you?" Item-specific scales instead of blanket agree/disagree formats break the acquiescence-plus-demand alignment that inflates ratings.
  • AI-moderated conversational probing that lets the student set the agenda. Rather than marching every respondent through the same leading battery, Koji''s moderated interview opens with broad, non-directive prompts and follows the student''s own concerns. Because the follow-up questions are generated from what the student actually said — not from a fixed hypothesis the student can reverse-engineer — there is less of a single "expected answer" to detect and comply with.
  • Protected anonymity and consistent framing. Demand pressure spikes when students believe a specific person will read their words and wants a specific verdict. Koji supports anonymised collection and a neutral, standardised introduction, so the administration does not itself broadcast a desired result. This also helps with the adjacent social-desirability problem discussed in Koji''s mode-effects and social desirability guidance.
  • Triangulation so no single self-report stands alone. Because demand characteristics are largely undetectable within a self-report, the defensible response is to corroborate. Koji''s thematic analysis of open text, quality scoring, and support for combining conversational evidence with other signals let a QA office ask whether the numbers, the students'' own words, and independent evidence tell the same story — the same triangulation logic behind the Hawthorne and observation-reactivity guidance.

None of this eliminates demand characteristics — Orne''s point was precisely that a motivated respondent can always infer something. Koji is designed to mitigate the effect by lowering the signal-to-guess ratio and by making compliance a less rewarding strategy than candour. Koji''s core research platform at koji.so applies the same AI-moderated interview engine to product and customer research, where demand effects — "tell the nice founder what they want to hear" — are an equally serious threat to validity.

Frequently asked questions

What is the difference between demand characteristics and social desirability bias? Social desirability is the pull toward answers that make the respondent look good (competent, kind, unbiased). Demand characteristics are the pull toward answers that fit the perceived purpose of the evaluation. A student might give an inflated rating because they have guessed the survey is meant to defend a course from cuts (demand), not because a generous rating flatters them (desirability). The two often act together, but neutral wording addresses demand while anonymity primarily addresses desirability.

Can you detect how much demand characteristics affected a set of evaluations? Generally, no — and this is a documented finding, not just a limitation. Nichols and Maner (2008) showed that post-hoc "suspicion probes" are unreliable at identifying who actually inferred the intended answer. Because you cannot measure the contamination directly, the credible strategy is prevention through design (neutral items, protected anonymity, non-directive probing) plus triangulation against independent evidence, rather than a statistical correction after collection.

Do demand characteristics always make ratings higher? No. The direction depends on the respondent''s stance. A cooperative "good subject" tends to confirm the answer they think is wanted, which often inflates scores; but an apprehensive or negativistic student — or one who resents a mandated evaluation — may push against the perceived expectation and deflate them. This is why demand is a validity threat (systematically off-centre) rather than simply a precision threat (widely scattered).

Does telling students the evaluation matters make the problem worse? It can. Announcing that scores affect promotion, awards, or a course''s survival installs a strong demand characteristic and, per Zizzo (2010), demand effects are strongest when the task is evaluative and the "right" answer is easy to infer. You do want students to know the purpose is genuine improvement, but the framing should communicate why their honesty helps without signalling a specific desired verdict.

How is this different from the Hawthorne effect? The Hawthorne effect is reactivity to being observed — behaving differently because you know you are being watched or measured. Demand characteristics are reactivity to an inferred hypothesis — behaving so as to confirm (or defy) what you think the study wants. A student can be indifferent to being surveyed yet still steer answers toward a purpose they have decoded. Koji treats them as related but separate design problems.

Are open-text and conversational formats more or less vulnerable than Likert scales? It depends on how they are run. A leading open prompt ("What did you love about this course?") carries demand just as a leading scale item does. But a non-directive, student-led conversation — where follow-ups are generated from the student''s own answers rather than a fixed script — gives the respondent less of a single expected answer to detect, which is the design Koji uses to reduce the pull.

Related resources

References

  • Orne, M. T. (1962). On the social psychology of the psychological experiment: With particular reference to demand characteristics and their implications. American Psychologist, 17(11), 776–783. https://doi.org/10.1037/h0043424
  • Nichols, A. L., & Maner, J. K. (2008). The good-subject effect: Investigating participant demand characteristics. The Journal of General Psychology, 135(2), 151–166. https://doi.org/10.3200/GENP.135.2.151-166
  • Weber, S. J., & Cook, T. D. (1972). Subject effects in laboratory research: An examination of subject roles, demand characteristics, and valid inference. Psychological Bulletin, 77(4), 273–295. https://doi.org/10.1037/h0032351
  • Zizzo, D. J. (2010). Experimenter demand effects in economic experiments. Experimental Economics, 13(1), 75–98. https://doi.org/10.1007/s10683-009-9230-z
  • Corneille, O., & Lush, P. (2023). Sixty years after Orne''s American Psychologist article: A conceptual framework for subjective experiences elicited by demand characteristics. Personality and Social Psychology Review, 27(1), 83–101. https://doi.org/10.1177/10888683221104368