The Dr. Fox Effect: Can a Charismatic Lecturer Fool Your Course Evaluation?
A famous 1973 experiment claimed an actor could earn glowing ratings for a lecture with no content. The truth is more interesting than the myth—and it tells you exactly what a single average score can and cannot see.
Koji for Education
Research & Editorial Team ·
Bottom line: Instructor expressiveness—enthusiasm, wit, vocal energy, charisma—has a large effect on how highly students rate teaching and only a small effect on how much they actually learn. Lecture content shows the reverse pattern. That mismatch, dramatised by the 1973 "Dr. Fox" experiment, is real and well-replicated. But the popular version of the story—that a charming actor "seduced" experts into believing they had learned from an empty lecture—overstates the evidence. The honest lesson is not that ratings are worthless; it is that a single global score cannot separate style from substance, and that separation is exactly what a well-designed evaluation should surface.
The experiment everyone half-remembers
In 1973, Naftulin, Ware, and Donnelly published The Doctor Fox Lecture: A Paradigm of Educational Seduction in the Journal of Medical Education. They hired a professional actor, gave him the invented title "Dr. Myron L. Fox," and coached him to deliver a lecture on "mathematical game theory as applied to physician education" to an audience of psychiatrists, psychologists, and educators. The lecture was deliberately built to be authoritative-sounding but contentless—padded with jargon, contradictions, and meaningless references, delivered with warmth and confidence. Afterwards, the audience rated him highly and several reported that the talk had been stimulating and informative.
The study became one of the most-cited weapons in the case against student evaluations of teaching (SET). If experts could be charmed into praising an empty lecture, the argument went, what hope is there that undergraduate ratings measure teaching quality rather than showmanship?
What the better evidence actually shows
The single most important synthesis came from Abrami, Leventhal, and Perry, whose meta-analysis of the "educational seduction" experiments reached a precise and often-misquoted conclusion. Across the controlled studies, instructor expressiveness had a substantial effect on student ratings but only a small effect on student achievement, while lecture content had a substantial effect on achievement but only a small effect on ratings. In other words, the two things a good evaluation should care about—did students enjoy it, and did students learn—are driven by largely different factors, and a global rating leans heavily toward the first.
That is the defensible core of the Dr. Fox literature, and it survives scrutiny. Expressiveness is not nothing—an engaging lecturer can raise motivation and attention, which can indirectly help learning—but its direct pull on the number students write on a form is far larger than its pull on what they can do afterwards.
The myth, corrected
Here is where intellectual honesty matters, because the original 1973 study was methodologically weak: no control lecture, tiny samples, a rating form never validated, and conclusions drawn from a handful of items. In 2014, Eyal Peer and Elisha Babad re-ran the analysis in the Journal of Educational Psychology under the pointed title The Doctor Fox Research (1973) Re-Revisited: "Educational Seduction" Ruled Out. Using the original Dr. Fox video and the original methodology, they found that the rating effect replicated—people did rate the charismatic speaker favourably—but the seduction did not. When asked directly, viewers did not believe they had learned a great deal. They enjoyed the performance and said so; they did not confuse enjoyment with mastery.
So the accurate statement is narrower and more useful than the legend: charisma reliably inflates satisfaction-type ratings, but students are not so easily fooled into over-reporting their own learning when you ask them the right question. The bias lives in the global "overall, this was an excellent course" item, not in students' entire perception of reality.
But doesn't this just prove student ratings are useless?
This is the strongest counterargument, and it deserves a straight answer. No—the Dr. Fox findings do not show that ratings are meaningless, and treating them that way is its own error.
First, expressiveness is partly a legitimate teaching skill. A lecturer who is audible, structured, and engaging genuinely helps a room full of people stay with a difficult idea. The problem is not that ratings reward clarity and energy; it is that a one-number summary cannot tell an administrator whether a high score reflects clarity and rigour, or charm instead of rigour.
Second, the effect is a reason to change what you collect, not to stop collecting. If a global rating conflates style and substance, the remedy is to ask questions that pull them apart—about specific skills gained, about workload and challenge, about what students can now do—and to triangulate ratings with other evidence (peer observation, learning-outcome data, samples of student work). The Dr. Fox literature is an argument for better instrumentation, not for silence.
What this means for how you evaluate courses
Three practical implications follow directly from the evidence:
- Distrust the single global score for high-stakes decisions. An "overall excellence" mean is the item most contaminated by expressiveness. It is the worst possible number on which to base promotion, tenure, or "course of concern" flags.
- Ask about learning specifically, not just satisfaction. Peer and Babad's correction is the actionable part: students can distinguish "I enjoyed this" from "I learned this" if you ask them separately. Evaluations that only ask for a global impression throw that information away.
- Probe the reason behind the praise. "The lecturer was engaging" and "the lecturer helped me understand a hard concept" can produce the same 5/5, but they mean very different things for course design.
A note on what "engaging" actually buys you
It is worth being precise about the mechanism, because it changes how you read a report. Expressiveness works largely through affect and attention: an animated lecturer raises arousal and liking, and liking bleeds into every judgement a student makes at the end of term. That is why the effect concentrates in global, evaluative items ("overall, an excellent course") and fades on specific, behavioural ones ("I received feedback that changed how I work"). The practical tell is the gap between a course's global score and its concrete, skill-anchored answers. A module that scores 4.8 overall but cannot point to a single specific thing students can now do is showing you the Dr. Fox signature: delight without demonstrable substance. A module that scores 4.2 overall but whose open comments are full of "I finally understand this" is the opposite—modest charisma, real learning. Neither pattern is visible if you only look at the mean of the global item, which is exactly why so many evaluation systems miss it. Reading for the gap, rather than the level, is the single most useful habit an evaluation committee can adopt, and it costs nothing but attention.
Where Koji fits
Koji for Education is built around the third point. Instead of a static form that collects a number and a shrug, Koji runs an AI-moderated conversational interview that can follow "I really liked this course" with "What specifically helped you learn?"—and can tell, from the answer, whether the praise attaches to delivery or to substance. Its automatic thematic analysis of open-text responses surfaces whether a cohort is describing charisma ("energetic," "funny," "great presence") or learning ("I can now build X," "the feedback on my draft changed how I write"), so a teaching-and-learning centre can see the difference at a glance rather than inferring it from a mean.
Because Koji supports six structured question types—including scale, open-ended, and single-choice—you can measure enjoyment and self-reported learning as distinct constructs rather than collapsing them into one global rating. The moderation is standardised and bias-aware: every student is probed with the same calibrated logic, removing the human-interviewer inconsistency that would otherwise add its own noise. And because collection can run mid-cycle, an engaging-but-shallow module can be caught and fixed while the students are still enrolled, rather than diagnosed a term too late.
Teams that also run general customer and user research use the same AI interview engine on the main Koji platform—the education product simply applies it to the evaluation problem.
The Dr. Fox effect is not a reason to stop listening to students. It is a reason to stop asking them a single, style-contaminated question and calling the average "teaching quality." Ask better questions, separate delight from learning, and the charisma stops hiding the signal.
Ready to move beyond the global score? See how Koji for Education evaluates courses.