Students Rate the Classes Where They Learn the Most Lower: The Feeling-of-Learning Gap
A randomised Harvard experiment found students learned more in active classrooms but rated their own learning lower. What the feeling-of-learning gap means for interpreting course-evaluation items that ask how much students learned.
Koji Education Team
Product
In short: A randomised, controlled experiment in Harvard introductory physics (Deslauriers, McCarty, Miller, Callaghan & Kestin, 2019, PNAS) found that students taught with active learning learned significantly more than peers taught the identical content by a polished lecturer — yet rated their own learning, their enjoyment, and the instruction lower. Students mistake the cognitive effort of active learning for ineffective teaching. Any course-evaluation item that asks "how much did you learn?" or "how effective was the instruction?" can therefore penalise exactly the evidence-based methods that improve outcomes. The remedy is not to discard student feedback, but to read perceived-learning items as measures of experience, not attainment, and to probe the feeling behind the rating.
What the research says
Louis Deslauriers and colleagues ran one of the cleanest experiments in the higher-education literature on the relationship between how learning feels and how much learning happens. In two large-enrolment introductory physics courses at Harvard, students were randomly assigned, topic by topic, to one of two conditions delivering identical content and identical handouts: an active-learning version built on in-class problem solving and small-group work, and a passive version delivered as a continuous lecture by an experienced, highly rated instructor. The same instructor taught both, and made no attempt to persuade students that either format was superior. After each session, students completed a short test of actual learning and a survey capturing their perceived learning, enjoyment, and engagement.
The result was a near-textbook dissociation. Students in the active condition scored higher on the test of actual learning — the expected direction given decades of evidence that active learning outperforms lecturing. But on the perception survey they reported feeling that they had learned less, found the active sessions less enjoyable, and judged the instruction less effective. In other words, the format that produced more learning produced a worse subjective experience. The authors trace the mechanism to the increased cognitive effort and the absence of the smooth, fluent narrative a skilled lecturer provides: effortful struggle is misread as failure to learn, while a polished lecture produces a comforting — but misleading — sense of mastery.
Crucially, the study also points to a partial fix. When the instructor briefly explained why active learning feels harder but works better, the gap between feeling and reality narrowed. Perception is malleable; it is not a fixed verdict on teaching quality.
This finding does not stand alone. Kirschner and van Merriënboer (2013), in their widely cited critique "Do Learners Really Know Best? Urban Legends in Education" (Educational Psychologist), argue that learners are unreliable judges of the instructional conditions that best serve their own learning — a direct challenge to treating student perception as a proxy for instructional quality. Robert Bjork's programme on desirable difficulties (Bjork & Bjork, 2011) supplies the underlying cognitive account: manipulations such as spacing, interleaving, and retrieval practice depress the subjective fluency of learning while improving durable retention and transfer. The learner's metacognitive read-out — "this feels easy, so I must be learning" — is precisely backwards. Deslauriers et al. is the field demonstration of what Bjork's laboratory work predicts.
Why it matters for course evaluation in practice
Most institutional evaluation instruments still contain at least one item of the form "I learned a great deal in this course" or "the teaching methods helped me learn," typically answered on a five-point agreement scale and then averaged into a headline number that feeds appraisal, promotion, and programme review. The Deslauriers finding shows that such items can be negatively correlated with the thing they purport to measure. An instructor who replaces comfortable lecturing with cognitively demanding active learning may see their perceived-learning scores fall because their students are learning more.
The practical risk is a perverse incentive. If perceived-learning and satisfaction items drive consequential decisions, rational instructors will optimise for fluency and comfort — clear, entertaining, low-friction lectures that feel productive — rather than for the effortful methods that improve attainment. This is the SET analogue of teaching to the test: teaching to the feeling. Quality-assurance offices that want their evaluation system to reward good pedagogy, not punish it, need to design and interpret these items with the feeling-of-learning gap firmly in view.
None of this means student feedback is worthless. Student experience is a legitimate object of measurement in its own right: a course that feels confusing, demoralising, or pointless is a real problem, even if some discomfort is productive. The error is the category mistake of treating "how much I felt I learned" as evidence of "how much I learned." Those are different constructs, and this study shows they can move in opposite directions.
Limitations and honest caveats
A critical reader should hold several caveats in mind before generalising. First, discipline and setting: the experiment was conducted in introductory physics at a highly selective institution, with students unaccustomed to active learning. Novelty effects — the discomfort of an unfamiliar format — may inflate the perception gap and could shrink as students acclimatise across a semester. Second, dose and duration: the manipulation operated at the level of individual topics, not a full term of summative evaluation, so the translation to end-of-semester SET is an inference, not a direct measurement. Third, the study measured perceived learning and instruction quality on a survey, which is closely related to, but not identical with, the multi-item institutional SET instruments used for personnel decisions. Fourth, the very fact that a short explanatory intervention narrowed the gap suggests the effect is context-dependent and partly remediable, not an immovable law. Finally, replication across disciplines, course levels, and cultures is still accumulating; the prudent reading is that the feeling-of-learning gap is real and well-grounded in cognitive theory, but its magnitude is moderated by familiarity, framing, and subject matter.
Acknowledging these limits actually strengthens the practical conclusion. Even at its most conservative, the evidence establishes that perceived learning cannot be assumed to track actual learning — which is all that is required to change how the item is interpreted.
How Koji incorporates this
Koji for Education is built on the premise that a single Likert number rarely tells you why a student answered the way they did — and the feeling-of-learning gap is the clearest case of why that matters. Koji is designed to mitigate the gap in several concrete ways:
- Conversational probing that separates effort from ineffectiveness. When a student rates perceived learning low or signals frustration, Koji's AI-moderated interview follows up — what specifically felt unproductive? This routinely surfaces the distinction between "this was uncomfortable and effortful" and "I genuinely could not learn from this," a distinction a bare scale erases. The same AI-moderated interview engine powers Koji's core research platform at koji.so for product and customer research.
- Construct separation in question design. Koji supports structured item types (scale, single_choice, open_ended) that let you measure affect, perceived learning, and observable behaviours (attendance, time-on-task, help-seeking) as distinct items rather than collapsing them into one satisfaction score — so a dip in comfort does not silently contaminate an inference about quality.
- Bias-aware reporting. Koji's analysis is designed to flag perceived-learning items as experience measures and to recommend triangulation against actual outcome evidence, rather than presenting a perceived-learning mean as if it were an attainment statistic.
- Formative, mid-cycle collection. Because the gap narrows when students understand why a method feels hard, Koji supports mid-semester formative cycles that give instructors the chance to explain their pedagogy and re-measure — closing the loop within the term rather than discovering the problem after grades are in.
- Triangulation across cohorts and evidence types. Koji is designed to sit alongside, not replace, learning-outcome data, so a programme can ask whether low perceived-learning scores coincide with high or low assessed attainment before drawing conclusions.
Koji is designed to mitigate the feeling-of-learning confound; it does not eliminate it, and no instrument can make student perception a perfect proxy for learning. What it can do is stop a perception item from being silently misread as an attainment item.
Related Resources
- The Fluency Illusion: Why a Polished Lecturer Earns Higher Course Evaluations Without Teaching More
- Does the Professor Students Rate Highest Teach Them Most? Value-Added Evidence vs Student Ratings
- Do Student Evaluations Measure Learning? The Uttl Meta-Analysis Revisited
- Why "I Learned a Lot" Can't Be Compared Across Courses: Reference Bias
- Response-Shift Bias: Why Self-Reported Learning Gains in Course Evaluations Can Mislead
References
- Deslauriers, L., McCarty, L. S., Miller, K., Callaghan, K., & Kestin, G. (2019). Measuring actual learning versus feeling of learning in response to being actively engaged in the classroom. Proceedings of the National Academy of Sciences, 116(39), 19251–19257. https://doi.org/10.1073/pnas.1821936116
- Kirschner, P. A., & van Merriënboer, J. J. G. (2013). Do learners really know best? Urban legends in education. Educational Psychologist, 48(3), 169–183. https://doi.org/10.1080/00461520.2013.804395
- Bjork, E. L., & Bjork, R. A. (2011). Making things hard on yourself, but in a good way: Creating desirable difficulties to enhance learning. In M. A. Gernsbacher et al. (Eds.), Psychology and the Real World (pp. 56–64). Worth Publishers.
Related articles
Do Cookies, Treats, and Mood Bias Course Evaluations?
Two controlled studies show that giving students chocolate or cookies before an evaluation measurably raises teaching scores. What the affect heuristic means for the validity of course evaluations — and how to design around it.
Does the Professor Students Rate Highest Teach Them Most? Value-Added Evidence vs Student Ratings
The strongest causal evidence — random-assignment studies that track follow-on course performance — shows the instructors students rate highest are often not the ones whose students learn most over time. What Carrell & West (2010) and Braga et al. (2014) found, and what it means for course evaluation.
The Fluency Illusion: Why a Polished Lecturer Earns Higher Course Evaluations Without Teaching More
A research-grounded look at the lecture fluency illusion — why a smooth, confident presentation inflates student ratings and perceived learning while leaving actual learning unchanged — and what it means for interpreting course evaluations.
Do Student Evaluations Measure Learning? The Uttl Meta-Analysis Revisited
A research-grounded reading of the Uttl, White & Gonzalez (2017) meta-analysis on the SET–learning relationship, with implications for European course evaluation and quality-assurance policy.