The Hawthorne Effect and Observation Reactivity in Course Evaluation
Does being observed change teaching and student behaviour enough to bias evaluation evidence? What the research actually shows about the Hawthorne effect and reactivity, its contested size, and how to design evaluation that does not depend on the one week someone is watching.
Koji Education Team
Product
The short answer
The worry is intuitive: if an instructor knows a peer observer is in the room, or students know an evaluation is running, they behave differently — so the evidence reflects the observation, not the ordinary course. This is reactivity, popularly called the Hawthorne effect. The honest research picture is more deflating than the folklore: the original Hawthorne "experiments" do not show what they are famous for showing, and rigorous reviews find reactivity effects that are real but generally small and inconsistent. The practical lesson is not "observation is worthless" but "do not let a single observed session, or an announced one-off survey, carry the weight of your judgement" — sample repeatedly, over time, and through channels students experience as routine.
BLUF: The Hawthorne effect — behaviour changing because people know they are being observed — is real but far smaller and less consistent than its reputation suggests, and its most famous evidence base was largely mythical (Levitt & List, 2011; McCambridge et al., 2014). For course evaluation this means one-off peer observations and heavily announced surveys can be modestly distorted by reactivity, so evidence should be gathered continuously and unobtrusively rather than in a single watched moment. Koji is designed to reduce reactivity by making feedback routine, embedded, and low-salience rather than a one-time event.
What the research says
The term comes from studies at Western Electric''s Hawthorne plant in the 1920s–30s, where worker productivity was said to rise whenever conditions were changed — even when lighting was dimmed — implying that attention itself, not the intervention, drove the effect. That story became a textbook staple. Then Steven Levitt and John List reanalysed the original illumination-experiment data (recovered after being thought lost) in the American Economic Journal: Applied Economics (2011). Their verdict is blunt: the celebrated patterns are, in their words, largely fictional. They do find some evidence consistent with a Hawthorne-type response — output tended to rise when experimental manipulations were made, and there were suggestive day-of-week patterns — but nothing like the dramatic, attention-driven surge of legend. The foundational anecdote, in short, does not support the strong claim built on it.
The most careful synthesis is McCambridge, Witton and Elbourne''s systematic review in the Journal of Clinical Epidemiology (2014). Reviewing 19 purpose-designed studies (including 8 randomised controlled trials), they conclude that research participation effects do occur, but are inconsistent in direction and size and poorly understood — sometimes increasing a behaviour, sometimes not detectable. They argue the vague, catch-all "Hawthorne effect" label should be retired in favour of more precise concepts about which aspect of being studied changes which behaviour and how. In other words: reactivity is not one thing, and treating it as a single large, predictable bias is a mistake.
Direct experimental evidence comes from McCarney and colleagues in BMC Medical Research Methodology (2007), who randomised participants in a dementia-treatment trial to more- versus less-intensive follow-up. The more-observed group showed modestly different outcomes on some measures, offering a clean demonstration that the intensity of being monitored can shift behaviour — but again, the effects were modest and measure-specific rather than sweeping. Across this literature the consistent themes are: reactivity is genuine, usually small, context-dependent, and often fades with habituation as observation becomes routine.
Why it matters for course evaluation in practice
Evaluation relies on several observation-heavy methods, each exposed to reactivity in a different way. Peer and teaching observation is the clearest case: an instructor teaching their single "observed lecture" may over-prepare, adopt behaviours they think the observer wants, and revert afterward — so the observation captures a performance, not the term. Structured observation protocols help standardise what is recorded but do not remove the fact that the observed session is atypical. Student behaviour can shift too: an announced, high-stakes evaluation window can prompt more engaged (or more performative) behaviour than an ordinary week, and the salience of "this counts" can change how students respond.
Reactivity also interacts with sampling. A single observed session is a sample of size one from a noisy process; even without reactivity it is unreliable, and reactivity adds a systematic tilt on top of the noise. The mitigation the evidence supports is therefore double: sample more (multiple sessions, multiple time points) to beat the noise, and sample less obtrusively (routine, embedded, low-announcement collection) to beat the reactivity. The reassuring corollary of the small-and-fading effect sizes is that continuous, normalised feedback largely dissolves the problem — when being asked is ordinary, there is little special "being watched" signal to react to.
Limitations and honest caveats
A rigorous reader should resist two opposite over-claims. The first is the folklore over-claim — that observation massively inflates performance — which the reanalyses (Levitt & List, 2011) directly undercut. The second is the dismissive over-claim — that reactivity is a myth and can be ignored — which the RCT and review evidence (McCarney et al., 2007; McCambridge et al., 2014) do not license, since real if modest effects appear. Beyond that: most of the strongest evidence comes from workplace and clinical settings, not university classrooms, so generalisation to course evaluation is an inference, not a measurement; classroom-specific quantification of reactivity is thin. "Reactivity" also bundles several distinct mechanisms (evaluation apprehension, demand characteristics, novelty, altered effort) that may not behave alike, which is precisely McCambridge and colleagues'' point. Habituation is assumed to attenuate the effect but its speed and completeness vary. And unobtrusive measurement raises its own ethical and validity questions — covert observation is not acceptable, and reducing salience must never mean reducing consent or transparency. The defensible position is calibrated: design so that no single observed moment is decisive, precisely because reactivity is real enough to matter but too small and variable to "correct for" with a formula.
How Koji incorporates this
Koji''s design philosophy — continuous, embedded, low-friction feedback — is directly aligned with what the reactivity evidence recommends: make being asked ordinary, so there is little special observation to react to.
- Routine, low-salience collection instead of a watched event. Koji supports mid-cycle and in-semester feedback rather than one high-stakes, heavily announced window, so no single moment carries decisive weight and habituation reduces reactivity over the term.
- Triangulation across many moments and cohorts. Because evidence accumulates across sessions, cohorts, and time points, a single atypical (over-prepared or performative) episode is diluted rather than decisive — the same "sample more" logic that beats both noise and reactivity.
- Conversational, student-paced interviews. Koji''s AI-moderated format lets students respond in their own time and words rather than performing for an obvious, high-stakes instrument, and consistent AI moderation removes the observer-specific "who is watching and what do they want" cues that drive demand characteristics in peer observation.
- Complement, not replacement, for observation. Koji positions student and self-report evidence alongside observation-based methods (documented in the platform''s guidance on peer observation and structured protocols), so an institution never relies on the single observed lecture that reactivity most distorts.
- Transparent and consented throughout. Reducing salience in Koji means embedding feedback into the normal rhythm of a course, never covert measurement — students always know feedback is being collected; it simply stops being an exceptional event.
Framed honestly, Koji does not "eliminate" reactivity — no method can, and the effect is real if modest. It is designed to mitigate it by dissolving the special-occasion quality of evaluation. Koji''s core research platform at koji.so applies the same continuous, embedded approach to product and customer research, where one-off, announced surveys create the same performance-versus-reality gap.
Frequently asked questions
What is the Hawthorne effect? It is the tendency for people to change their behaviour because they know they are being observed or studied, independent of any actual intervention. The name comes from 1920s–30s studies at the Hawthorne plant, though those studies are now known not to support the strong version of the claim.
Is the Hawthorne effect real or a myth? Both, in a sense. The famous original evidence was largely mythical (Levitt & List, 2011), but rigorous later work finds genuine research-participation effects that are typically small, inconsistent in direction, and context-dependent (McCambridge et al., 2014; McCarney et al., 2007).
How does reactivity bias course evaluation? Chiefly through observation-heavy methods: an instructor may over-prepare for a single observed lecture and revert afterward, and students may behave or respond differently during an announced, high-stakes evaluation window than in an ordinary week.
How can institutions reduce reactivity? Gather evidence continuously and across multiple sessions and time points rather than in one watched moment, keep collection routine and low-salience, and never let a single observation be decisive. Habituation reduces the effect once feedback becomes ordinary.
Does reducing salience mean observing students covertly? No. Reducing salience means embedding feedback into the normal rhythm of a course so it is not an exceptional event — always with consent and transparency. Covert observation is neither ethical nor acceptable.
Related resources
- COPUS and TDOP: structured classroom-observation protocols
- Peer observation vs student evaluations: convergent validity
- Panel conditioning and response drift in repeated evaluations
- Feeling of learning vs actual learning in active classrooms
- Experience sampling and EMA for in-semester feedback
- Mode effects and social desirability in conversational evaluation
References
- Levitt, S. D., & List, J. A. (2011). Was there really a Hawthorne effect at the Hawthorne plant? An analysis of the original illumination experiments. American Economic Journal: Applied Economics, 3(1), 224–238. https://doi.org/10.1257/app.3.1.224
- McCambridge, J., Witton, J., & Elbourne, D. R. (2014). Systematic review of the Hawthorne effect: New concepts are needed to study research participation effects. Journal of Clinical Epidemiology, 67(3), 267–277. https://doi.org/10.1016/j.jclinepi.2013.08.015
- McCarney, R., Warner, J., Iliffe, S., van Haselen, R., Griffin, M., & Fisher, P. (2007). The Hawthorne effect: A randomised, controlled trial. BMC Medical Research Methodology, 7, 30. https://doi.org/10.1186/1471-2288-7-30
Related articles
Does Evaluating Course After Course Change How Students Answer? Panel Conditioning in Repeated Course Evaluations
Students at a European university complete dozens of course evaluations across a degree. Panel-conditioning research shows that the mere act of being surveyed repeatedly can change later answers — a threat to comparing scores across years and cohorts.
Stop Waiting for the End of Term: Experience Sampling for In-the-Moment Course Feedback
Experience-sampling methods capture what students feel and think during a course, not their reconstructed memory of it months later. Here is why in-the-moment data can be more valid than the end-of-term survey — and how to use it responsibly.
Beyond the Student Survey: Structured Classroom-Observation Protocols (COPUS and TDOP) as Evaluation Evidence
What COPUS and the TDOP measure that student surveys and traditional peer visits cannot: low-inference, reliable records of what actually happens in a classroom. Their evidence base, their limits, and where they fit in a triangulated evaluation.
Does a Conversational Course Evaluation Make Students Less Honest? Mode Effects and Social Desirability
Survey mode shapes honesty: interviewer-administered surveys invite more social-desirability bias than self-administered ones. What Tourangeau and Yan (2007) and the mode-effects literature mean for anonymous, AI-moderated conversational course evaluations.