The Active-Learning Penalty: Why Your Best Teaching Can Score Worst
A landmark Harvard experiment found students learned significantly more in active classrooms but rated their learning—and their instructor—lower. If your course evaluation rewards the feeling of learning over learning itself, it can quietly punish your most effective teachers.
Koji Education Team
Product ·
Bottom line up front: There is robust experimental evidence that students can learn more from a teaching method while rating it—and the instructor—lower. A 2019 study published in the Proceedings of the National Academy of Sciences found students in active-learning physics classes scored about 0.46 standard deviations higher on a test of actual learning, yet reported their feeling of learning roughly 0.56 standard deviations lower than peers who sat through a polished lecture. A course evaluation that captures satisfaction and the feeling of learning, but never separates it from evidence of actual learning, can therefore penalise exactly the teaching your institution most wants to encourage. This is not an argument against student feedback. It is an argument for collecting feedback that can tell the difference.
The experiment that should worry every quality office
Louis Deslauriers and colleagues at Harvard ran a controlled, randomised comparison inside introductory college physics (Deslauriers et al., 2019, PNAS 116(39): 19251–19257). Across two semesters and 149 students, the same instructors taught the same content (statics and fluids) using either active learning—problem-solving, peer discussion, in-class struggle—or a fluent, well-organised passive lecture. The materials were identical; only the mode of engagement differed.
The results were paradoxical and remarkably consistent. On an objective multiple-choice test of what they had learned, the active-learning students outperformed the lecture students by approximately 0.46 standard deviations. But when asked how much they felt they had learned, the same active-learning students rated their learning about 0.56 standard deviations lower. They also rated the instructor as less effective, and a majority said they would have preferred all their physics classes to be taught by lecture. They learned more and believed the opposite.
The authors traced the gap to the fluency of instruction. A smooth, confident lecture feels like understanding. The cognitive effort of active learning—the friction of being asked to work something out before you have been told the answer—feels like confusion. Students mistook the absence of struggle for the presence of learning.
This is a known cognitive effect, not a one-off
The Harvard finding sits on top of decades of cognitive-science research. Robert Bjork's concept of "desirable difficulties" describes how conditions that slow learning down and make it feel harder—spacing, interleaving, retrieval practice—reliably improve long-term retention. The corollary is the "fluency illusion": when material is presented smoothly, learners systematically overestimate how well they have learned it. Re-reading a chapter feels productive; it usually is not. Watching an expert solve a problem feels like competence; it rarely transfers.
For a quality-assurance officer, the implication is uncomfortable. A single end-of-semester question—"How much did you learn in this course?" or "How effective was the teaching?"—does not measure learning. It measures the feeling of learning, and that feeling is biased in a specific, predictable direction: against effortful, evidence-based pedagogy and in favour of charismatic, frictionless delivery. The Dr. Fox effect, where an expressive lecturer delivering vacuous content earns high ratings, is the same coin's other face.
Why this matters more in 2026 than it did in 2019
Two trends sharpen the stakes. First, European institutions are under sustained pressure—from employers, from accreditation bodies, from their own teaching-and-learning centres—to move away from transmissive lecturing toward active, applied, competency-focused teaching. Second, those same institutions still make promotion, probation, and timetabling decisions partly on the basis of student satisfaction scores. If the metric and the mandate point in opposite directions, rational academics will notice. The quiet result is a structural disincentive to adopt the very methods the institution is asking for.
"But doesn't this just prove students are unreliable judges?"
This is the strongest counterargument, and it deserves a fair hearing. If students cannot tell when they are learning, why ask them at all?
The answer is that students are not unreliable about everything—they are unreliable about one specific thing: judging their own learning gains in the moment. They remain the only people who experienced the course from the inside. Students are excellent witnesses to whether instructions were clear, whether feedback arrived in time to be useful, whether the workload was survivable, whether they felt able to ask questions, and whether the course connected to anything they cared about. The error is not in asking students; it is in asking them a question they are cognitively ill-equipped to answer, and then treating their answer as a measure of teaching quality.
There is also a second-order point the Deslauriers team themselves raised: the feeling-of-learning penalty for active learning shrinks when instructors explicitly explain to students why the harder method works. Perception is partly a function of framing. That is something feedback can surface and something teaching can address—but only if your evaluation instrument is rich enough to capture the reasons behind a rating, not just the number.
What a better instrument actually does
If the problem is that a Likert item collapses "I enjoyed this" and "I learned from this" into a single number, the fix is to stop collecting a single number. Three moves help:
- Separate satisfaction from learning, explicitly. Ask about clarity, support, and workload as distinct constructs—and treat self-reported learning as a perception to be interpreted, not a measure to be averaged. Triangulate it against assessment data, not against other surveys.
- Capture the why, not just the what. A score of 3/5 on "I learned a lot" is useless on its own. "I found the in-class problems frustrating because I didn't see the point until the exam" is actionable—and tells you the framing problem, not a teaching-quality problem.
- Collect formatively, mid-cycle. If active-learning friction is depressing perceptions, you want to know in week 4, when you can still explain the method, not in week 12 when the ratings are locked.
This is where modern, AI-native evaluation earns its place. Koji for Education runs AI-moderated conversational interviews rather than static Likert grids. When a student says a course felt confusing or that they "didn't learn much," Koji's moderator can probe—neutrally and consistently—why it felt that way: was the material unclear, or was it the productive struggle of active learning misread as failure? Its automatic thematic analysis surfaces the pattern across hundreds of responses ("students value the problem sets but resent not being told the answer first"), and its support for distinct structured question types (open-ended, scale, single-choice, ranking) lets you keep enjoyment and learning conceptually separate instead of fusing them into one misleading average. Because the moderation is standardised and bias-aware, you avoid the human-interviewer inconsistency that makes qualitative feedback hard to scale—and because it supports mid-cycle collection, you can catch the perception gap while there is still time to act.
Koji does not claim to eliminate the fluency illusion—no instrument can read a student's mind. It mitigates the damage by refusing to reduce a complex perception to a single rankable digit, and by giving teachers the context to interpret a low "feeling of learning" score correctly instead of abandoning effective methods to chase it.
The same conversational interview engine underpins the main Koji platform for general user and customer research, where the identical problem—people are poor judges of their own experience in the moment—shows up constantly.
A note on the strength of the evidence
It is fair to ask how far a single physics study generalises. Two things make the Deslauriers result hard to dismiss. It was a controlled, randomised, within-subjects design—the same students experienced both conditions, which rules out the usual confounds of cohort or instructor differences. And it converges with a much larger literature: meta-analyses of active learning across STEM disciplines find consistent gains in performance, while the cognitive-science work on the fluency illusion and desirable difficulties supplies the mechanism. The specific magnitudes will vary by discipline and context, but the direction of the perception gap is robust and theoretically expected, not a statistical fluke.
The takeaway for quality assurance
If your evaluation framework cannot distinguish a course students enjoyed from a course they learned from, it will, on average, reward fluency over rigour. The evidence that this happens is experimental, replicated, and grounded in cognitive science. The remedy is not to stop listening to students—it is to ask them better questions, interpret their answers in light of what they can and cannot judge, and triangulate the feeling of learning against the fact of it. Teaching that feels hard is often teaching that works. Your metrics should be able to tell.