New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods11 min read

Desirable Difficulties: Why the Teaching That Improves Learning Often Lowers Satisfaction

Spacing, interleaving, and retrieval practice are among the best-evidenced ways to make learning durable — and they make a course feel harder and less smooth in the moment. Bjork & Bjork's (2011) desirable-difficulties principle explains why end-of-term satisfaction ratings systematically penalise the most effective teaching.

Koji Education Team

Product

In short: Some of the most effective teaching techniques — spacing practice over time, interleaving topics, and using tests to learn rather than just to grade — deliberately make learning feel harder and slower in the moment while making it far more durable. Bjork and Bjork (2011) call these desirable difficulties. Because students judge how much they are learning by how fluent and easy it feels, an end-of-term satisfaction survey systematically under-rates instructors who use these techniques. A quality-assurance process that treats raw satisfaction as a measure of teaching effectiveness will therefore penalise exactly the practices it should reward. The fix: measure specific practices, use delayed and outcome-based evidence, and explain to students why the course felt hard.

What the research says

Elizabeth and Robert Bjork''s (2011) chapter "Making things hard on yourself, but in a good way" crystallised decades of memory research into a deceptively simple idea. Conditions of learning that make performance improve rapidly and feel easy often produce fragile knowledge that fades, whereas conditions that introduce difficulty — as long as the learner can overcome it — produce durable, flexible, transferable learning. They named the second category desirable difficulties. The canonical examples are well replicated:

  • Spacing study over time rather than massing ("cramming") it. Massed practice feels efficient and produces strong immediate performance, but spaced practice yields markedly better long-term retention.
  • Interleaving different problem types or topics rather than blocking them. Blocked practice feels smoother and boosts performance during practice; interleaving feels harder and slower but improves later discrimination and transfer (Rohrer & Taylor, 2007).
  • Retrieval practice — using low-stakes tests to pull information out rather than re-reading to put it in. Roediger and Karpicke (2006), in Psychological Science, showed that students who repeatedly tested themselves retained substantially more a week later than students who repeatedly re-read — even though the re-readers felt more confident and predicted they would do better. This dissociation between felt fluency and actual retention is the heart of the matter.
  • Generation and varying the conditions of practice, which similarly depress immediate ease while strengthening durable learning.

The mechanism that makes this an evaluation problem is metacognitive: learners systematically mistake current fluency for learning. Bjork, Dunlosky, and Kornell (2013), reviewing self-regulated learning in the Annual Review of Psychology, document that students routinely misjudge which strategies work, favouring re-reading and massing because those feel productive, and undervaluing spacing, interleaving, and testing because those feel like struggle. The learner''s in-the-moment sense of "I''ve got this" is generated precisely by the conditions that fail to produce lasting learning.

This connects directly to a finding already documented in the course-evaluation literature. DeSlauriers and colleagues (2019, PNAS) ran a controlled comparison in which students in active-learning classes learned more (measured by test performance) yet rated their learning and the instruction lower than students in a fluent passive lecture — the feeling-of-learning gap. Desirable difficulties are the broader learning-science principle behind that specific result: any technique that trades in-the-moment ease for durable learning will tend to feel worse to the student while working better. It also helps explain why large-scale value-added studies — such as Carrell and West''s analysis and the Bocconi natural experiment — repeatedly find that instructors who produce better long-term outcomes are sometimes rated lower by their students.

Why it matters for course evaluation in practice

The practical implication is stark: a raw end-of-term satisfaction rating is negatively biased against desirable difficulties. An instructor who spaces and interleaves content, replaces re-reading with frequent low-stakes quizzing, and forces students to generate answers before revealing them is deliberately increasing in-the-moment difficulty. Students experience more struggle, less fluency, and a lower sense of "I''m learning this easily" — and, if the survey asks "How much did you learn?" or "How satisfied were you?", they will tend to score that course lower than a smoother, less effective one. Using such scores in personnel or ranking decisions creates a direct incentive to remove desirable difficulties: to teach in the fluent, crammable, satisfying way that produces good evaluations and poor retention. This is a textbook case of the difficulty-and-workload confound working against good teaching, and it compounds the broader question of whether student ratings measure learning at all.

Three design responses follow.

First, measure the practice, not the feeling. Instead of relying on global satisfaction, include behaviourally specific items that record whether desirable difficulties were used: "The course used frequent low-stakes quizzes or retrieval activities," "Topics were revisited and mixed over the term rather than covered once and dropped," "I had to attempt problems before being shown the solution." These are neutral, observable low-inference items — they let QA staff see the pedagogy rather than infer it from a satisfaction score that penalises it.

Second, separate satisfaction from effectiveness, and triangulate with delayed evidence. Because the cost of desirable difficulties is paid immediately and the benefit arrives later, end-of-term timing is the worst possible moment to judge effectiveness. Where feasible, complement the end-of-term survey with delayed measures — performance in later courses that build on the material, or alumni feedback — which reverse the bias by capturing durable learning.

Third, close the metacognitive gap with students. Part of the low rating is simple misattribution: students blame the instructor for the discomfort of effortful learning. Interventions that explain why a course uses spacing, interleaving, and retrieval — and that show students their own improved retention — can reduce the satisfaction penalty without abandoning the technique. An evaluation programme can support this by asking students to reflect on effort and durability rather than only on ease, nudging the metacognition the research says is miscalibrated. This also relieves some of the cognitive-load confusion between productive struggle and poor design, and complements an ICAP-style focus on what students actually did.

Limitations and honest caveats

Several honest qualifications matter. First, difficulties are only desirable when the learner has the background and support to overcome them; the same technique becomes an undesirable difficulty for an underprepared student, producing frustration and worse learning. So a low satisfaction score is not automatically a badge of good teaching — sometimes a course is simply too hard, poorly scaffolded, or badly designed, and the discomfort is extraneous rather than germane. Distinguishing productive struggle from destructive struggle is exactly the judgement an evaluation must support, not short-circuit.

Second, most of the foundational evidence comes from controlled memory experiments on relatively well-defined material; the size of the retention benefit in a messy, semester-long, multi-topic course with unequal prior knowledge is harder to pin down, and effect sizes vary with content and implementation. Interleaving, in particular, helps most when the to-be-discriminated categories are genuinely confusable.

Third, the argument is not a licence to ignore satisfaction. Student experience matters in its own right — for motivation, wellbeing, persistence, and retention in the programme — and a course that produces durable learning at the cost of alienating half the cohort has a real problem. The claim is narrower: satisfaction should not be equated with effectiveness, and a rating dip that coincides with well-implemented desirable difficulties should be interpreted, not punished.

Finally, the metacognitive-illusion evidence is robust in aggregate but varies across individuals; some students accurately perceive the value of effortful learning. Reading every low rating as a metacognitive error would be as wrong as reading every low rating as bad teaching. The point is to hold both possibilities open and gather the evidence — practice items, outcomes, delayed measures — that lets you tell them apart.

How Koji incorporates this

Koji is designed to keep the satisfaction penalty for good teaching from silently corrupting quality decisions.

  • Practice-specific items alongside satisfaction. Koji''s structured question types (scale, yes_no, single_choice, multiple_choice) let institutions record whether retrieval practice, spacing, and interleaving were used — so a report distinguishes "students felt it was hard" from "the instructor used evidence-based difficult practices," rather than collapsing them into one score.
  • Conversational probing of struggle. Because Koji runs an AI-moderated conversational interview, when a student reports that a course was hard or frustrating, the follow-up probe distinguishes productive difficulty ("I had to work for it but it stuck") from destructive difficulty ("I was lost with no support") — the desirable-versus-undesirable distinction that a single satisfaction number erases.
  • Effort-and-durability reflection. Koji can ask students to reflect on how well they expect to retain the material, not just how easy it felt, gently surfacing the metacognitive gap the research documents and giving QA a richer signal than momentary fluency.
  • Triangulation and closing the loop. Koji supports collection across cohorts and over time and tracks actions taken, so a satisfaction dip tied to desirable difficulties can be read against later outcomes instead of triggering a reflexive correction. The same AI-moderated interview engine powers Koji''s core research platform at koji.so, where separating what users felt in the moment from what actually worked is the identical methodological problem.

As always, Koji is designed to mitigate the fluency illusion in feedback, not to eliminate it; interpreting a low rating still requires the QA team to weigh the practice evidence and outcomes together.

Related Resources

References

  • Bjork, E. L., & Bjork, R. A. (2011). Making things hard on yourself, but in a good way: Creating desirable difficulties to enhance learning. In M. A. Gernsbacher, R. W. Pew, L. M. Hough, & J. R. Pomerantz (Eds.), Psychology and the Real World: Essays Illustrating Fundamental Contributions to Society (pp. 56–64). Worth Publishers.
  • Roediger, H. L., & Karpicke, J. D. (2006). Test-Enhanced Learning: Taking Memory Tests Improves Long-Term Retention. Psychological Science, 17(3), 249–255. https://doi.org/10.1111/j.1467-9280.2006.01693.x
  • Rohrer, D., & Taylor, K. (2007). The shuffling of mathematics problems improves learning. Instructional Science, 35(6), 481–498. https://doi.org/10.1007/s11251-007-9015-8
  • Bjork, R. A., Dunlosky, J., & Kornell, N. (2013). Self-Regulated Learning: Beliefs, Techniques, and Illusions. Annual Review of Psychology, 64, 417–444. https://doi.org/10.1146/annurev-psych-113011-143823
  • Deslauriers, L., McCarty, L. S., Miller, K., Callaghan, K., & Kestin, G. (2019). Measuring actual learning versus feeling of learning in response to being actively engaged in the classroom. Proceedings of the National Academy of Sciences, 116(39), 19251–19257. https://doi.org/10.1073/pnas.1821936116

Related articles

research-methods

Does Course Difficulty and Workload Lower Student Evaluations? What Centra Found

A research-grounded look at whether harder, heavier courses are punished in student evaluations of teaching. Centra (2003) and Marsh & Roche (2000) show the relationship is non-linear and weaker than faculty fear — with practical implications for fair course evaluation.

research-methods

Does the Professor Students Rate Highest Teach Them Most? Value-Added Evidence vs Student Ratings

The strongest causal evidence — random-assignment studies that track follow-on course performance — shows the instructors students rate highest are often not the ones whose students learn most over time. What Carrell & West (2010) and Braga et al. (2014) found, and what it means for course evaluation.

research-methods

Students Rate the Classes Where They Learn the Most Lower: The Feeling-of-Learning Gap

A randomised Harvard experiment found students learned more in active classrooms but rated their own learning lower. What the feeling-of-learning gap means for interpreting course-evaluation items that ask how much students learned.

research-methods

The Teachers Who Help You Most Get Rated Worst: What the Bocconi Natural Experiment Proved

Braga, Paccagnella and Pellizzari (2014) used near-random assignment of students to professors at Bocconi University to show that the teachers who most improved students' performance in later courses received the lowest student evaluations. Here is the design, the effect sizes, the caveats, and what it means for European quality assurance.