New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods9 min read

The Fluency Illusion: Why a Polished Lecturer Earns Higher Course Evaluations Without Teaching More

A research-grounded look at the lecture fluency illusion — why a smooth, confident presentation inflates student ratings and perceived learning while leaving actual learning unchanged — and what it means for interpreting course evaluations.

Koji Education Team

Product

In brief

A confident, fluent delivery makes students feel they have learned more and rate the instructor more highly — but controlled experiments show it does not increase how much they actually learn. The gap between perceived and measured learning means that a course evaluation built only on satisfaction and self-reported learning can reward presentation polish over instructional substance. Reading evaluations well, and designing the questions that produce them, requires separating the experience of a lecture from its effect.

What the research says

The cleanest demonstration comes from Carpenter, Wilford, Kornell, and Mullaney (2013). Participants watched a short video of the same instructor explaining a science concept (how the body absorbs nutrients) in one of two styles. In the fluent condition the instructor stood upright, maintained eye contact, and spoke fluidly without notes; in the disfluent condition the same person slumped, looked away, and spoke haltingly while reading. Students who saw the fluent lecture predicted they had learned substantially more and rated the instructor as better prepared and more effective. Yet when both groups were tested, actual performance did not differ — fluency moved perceptions without moving learning. The authors titled the paper "Appearances can be deceiving" for good reason.

This is not a one-off. Toftness and colleagues (2018) replicated the effect with a full-length (~31-minute) recorded lecture, manipulating gesture, vocal variety, and movement versus a monotone reading from notes. Fluent-instructor viewers again rated teaching effectiveness higher and predicted greater learning, while immediate and one-day-delayed test scores were statistically indistinguishable between conditions — an "illusion of learning." Carpenter, Northern, Tauber, and Toftness (2020), publishing in the Journal of Experimental Psychology: Applied, extended this across two further experiments and added information about instructor experience. Fluency again inflated judgments of learning and produced more favourable instructor evaluations without raising test scores, and telling students the instructor was "experienced" or "inexperienced" did not undo the effect: students anchored on how the lecture looked and sounded, not on what it taught.

The pattern aligns with a larger and more consequential finding from Deslauriers, McCarty, Miller, Callaghan, and Kestin (2019) in PNAS. In a randomized crossover in introductory physics, the same students experienced both active learning and a polished passive lecture on matched topics. Students scored roughly 0.46 standard deviations higher on tests after active learning — yet reported about 0.56 standard deviations lower feeling of learning, and rated the smoother passive lecture more favourably. The effortful, less fluent experience taught more but felt worse. Step back to the question of whether ratings track learning at all, and Uttl, White, and Gonzalez (2017) provide the sobering meta-analytic answer: once small-sample and publication-bias artefacts are removed, multisection student-evaluation-of-teaching ratings explain at most about 1% of the variance in student learning, and are essentially unrelated to it.

Together these studies converge on a single uncomfortable point: the very experience that makes a lecture feel effective — disfluency removed, confidence projected, cognitive effort minimized — is partly orthogonal to whether students learn.

Why it matters for course evaluation in practice

Most institutional evaluation instruments lean heavily on two kinds of item: global satisfaction ("Overall, this was an excellent course/teacher") and self-reported learning ("I learned a great deal in this course"). The fluency research shows both are vulnerable to a systematic distortion. An instructor who polishes delivery — rehearsed transitions, no visible hesitation, a confident manner — will likely score higher on these items even if the underlying instruction is unchanged, and an instructor who introduces effortful, evidence-based methods (retrieval practice, problem-solving, productive struggle) may score lower precisely because those methods feel harder.

For a quality-assurance office, three practical consequences follow. First, a high satisfaction score is not evidence of learning, and a dip after a pedagogical redesign may signal the redesign is working, not failing. Second, comparisons across instructors on global items partly rank presentation style, which disadvantages exactly the colleagues a teaching-and-learning centre is trying to encourage. Third, the open-ended comments and the reasons behind a rating carry more diagnostic value than the number itself — a student who says "it felt slow and I had to think hard" is describing a feature, not a defect. None of this means evaluations are worthless; it means they measure the student experience, which is worth knowing, and must be triangulated with other evidence before it is read as a measure of teaching effectiveness.

Limitations and honest caveats

A critical reader should hold several caveats in view. Most fluency experiments use single, short, video-recorded lectures with student or participant samples, not semester-long courses; the effect on a cumulative end-of-term evaluation, formed over many sessions and assessments, may be smaller or partially corrected by experience. The Deslauriers crossover is powerful but was conducted in one discipline at one selective institution, and the "feeling of learning" gap may narrow as students acclimate to active methods. Effect sizes for the perceived-versus-actual dissociation vary across replications, and a few studies find fluency exerts small positive effects on some learning measures under particular conditions. The meta-analytic claim that ratings barely correlate with learning rests on multisection designs that themselves have constraints (common exams, comparable cohorts). Finally, "feeling of learning" is not worthless: student affect, confidence, and willingness to engage matter for persistence and motivation, even when they decouple from a single test score. The honest synthesis is not "evaluations are invalid" but "satisfaction and self-reported learning are measures of experience that can move independently of learning, so they must be interpreted as such."

How Koji incorporates this

Koji for Education is designed to mitigate the fluency illusion at the point of measurement, not to pretend it away. Several mechanisms are relevant:

  • Probing beyond the number. Koji's AI-moderated conversational interviews do not stop at a Likert rating. When a student gives a high or low overall score, the interviewer follows up to surface why — asking what specifically helped or hindered understanding. This separates "I enjoyed the delivery" from "I can now solve these problems," information a single global item collapses.
  • Behaviourally specific, low-inference questions. Alongside global items, Koji supports structured open_ended, scale, single_choice, multiple_choice, ranking, and yes_no questions that can target concrete teaching behaviours and learning activities (Was retrieval practice used? Did you have to apply concepts to new problems?) rather than only global impressions, reducing the weight carried by presentation polish.
  • Automatic thematic analysis of open text. Koji codes open responses into themes, so a quality office can see whether "felt hard / had to think" clusters appear — often a marker of effortful, effective instruction — instead of reading a low satisfaction mean as straightforward bad teaching.
  • Bias-aware reporting. Reports are framed to discourage treating a satisfaction or self-reported-learning score as a direct measure of learning, and to encourage triangulation with assessment data and peer review.
  • Formative, mid-cycle collection. Because Koji supports mid-semester feedback, instructors who adopt effortful methods can capture and explain the temporary "this feels harder" reaction before it crystallizes into a lower end-of-term rating.

The aim is framed honestly: these features are designed to mitigate the fluency illusion by enriching what a rating means, not to eliminate a bias rooted in human metacognition. Koji's core research platform at koji.so applies the same AI-moderated interview engine to product and customer research, where the gap between what users say they want and what actually drives behaviour is a close cousin of the perceived-versus-actual-learning gap described here.

Reading a course evaluation in light of the fluency illusion

The practical upshot is a small set of interpretive habits. When a score on a global or self-reported-learning item moves, ask first whether anything about delivery changed — a new room, a more rehearsed term, a switch from improvised to scripted lectures — before concluding that learning changed. When an instructor introduces effortful methods and ratings soften, read the open-ended comments for effort markers ("had to think", "felt slow", "wished she would just tell us the answer"); these frequently accompany the methods that the evidence links to better performance. Conversely, an unusually high satisfaction score deserves the same scrutiny as a low one: it may reflect a fluent, low-demand experience rather than a demanding, high-yield one.

For instrument designers, the lesson is to stop treating "I learned a lot" as a proxy for learning and to pair it with items that ask about specific, effortful activities — retrieval practice, application to novel problems, productive struggle — whose presence is a better-validated signal than a student's global feeling. Pairing a perception item with a behaviour item lets a quality office triangulate the two and notice when they diverge, which is exactly the diagnostic moment the fluency research predicts. The number is not meaningless; it simply answers a different question ("how did this feel?") than the one a personnel committee usually wants answered ("how much did students learn?").

Related resources

References

  • Carpenter, S. K., Wilford, M. M., Kornell, N., & Mullaney, K. M. (2013). Appearances can be deceiving: instructor fluency increases perceptions of learning without increasing actual learning. Psychonomic Bulletin & Review, 20(6), 1350–1356. https://doi.org/10.3758/s13423-013-0442-z
  • Toftness, A. R., Carpenter, S. K., Geller, J., Lauber, S., Johnson, M., & Armstrong, P. I. (2018). Instructor fluency leads to higher confidence in learning, but not better learning. Metacognition and Learning, 13(1), 1–14. https://doi.org/10.1007/s11409-017-9175-0
  • Carpenter, S. K., Northern, P. E., Tauber, S. K., & Toftness, A. R. (2020). Effects of lecture fluency and instructor experience on students' judgments of learning, test scores, and evaluations of instructors. Journal of Experimental Psychology: Applied, 26(1), 26–39. https://doi.org/10.1037/xap0000234
  • Deslauriers, L., McCarty, L. S., Miller, K., Callaghan, K., & Kestin, G. (2019). Measuring actual learning versus feeling of learning in response to being actively engaged in the classroom. PNAS, 116(39), 19251–19257. https://doi.org/10.1073/pnas.1821936116
  • Uttl, B., White, C. A., & Gonzalez, D. W. (2017). Meta-analysis of faculty's teaching effectiveness: student evaluation of teaching ratings and student learning are not related. Studies in Educational Evaluation, 54, 22–42. https://doi.org/10.1016/j.stueduc.2016.08.007