What Cognitive Load Theory Says Your Course Evaluation Should — and Shouldn't — Ask
Cognitive load theory (Sweller, van Merriënboer & Paas, 2019) distinguishes the unavoidable difficulty of content from difficulty caused by poor design. That distinction changes what a course evaluation should measure: not overall 'difficulty', but the design choices that impose or remove extraneous load.
Koji Education Team
Product
In short: A course evaluation should not ask students to rate a course''s "difficulty" as a proxy for teaching quality. Cognitive load theory (CLT) shows that difficulty has three distinct sources — the intrinsic complexity of the material, the extraneous load created by poor instructional design, and the germane effort that builds understanding. Only extraneous load is a design fault. Well-targeted evaluation items ask about the specific, designer-controllable practices that impose or relieve extraneous load — sequencing, worked examples, split attention, redundancy — rather than lumping them into a single difficulty score that confounds a badly designed easy course with a well-designed hard one.
What the research says
Cognitive load theory begins from a fact about human cognition: working memory is severely limited when it processes novel information, holding only a handful of interacting elements at once, whereas long-term memory is effectively unlimited and stores organised schemas. John Sweller''s foundational work (1988) argued that instruction succeeds or fails largely by how it manages the load on that narrow working-memory channel. In their 20-year retrospective, Sweller, van Merriënboer, and Paas (2019, Educational Psychology Review) consolidate decades of experiments into a mature account with three load types:
- Intrinsic load is set by the element interactivity of the material — how many things must be held in mind simultaneously to understand it. Solving a multi-step equation has high intrinsic load; memorising an isolated vocabulary word has low intrinsic load. Intrinsic load is a property of the content relative to the learner''s expertise; it cannot be wished away, only sequenced and segmented.
- Extraneous load is imposed by the way material is presented — and it is the designer''s responsibility. Splitting related information across a diagram and a separate caption (the split-attention effect), narrating text that is also shown on screen word-for-word (the redundancy effect), or withholding worked examples from novices all consume working memory without contributing to learning.
- Germane load is the productive effort learners invest in constructing schemas. In the current formulation it is not a third independent source but the working-memory resources devoted to dealing with intrinsic load; good design frees capacity so that more of it can be germane.
The instructional prescription follows directly: reduce extraneous load, manage intrinsic load through sequencing and segmenting, and leave room for germane processing. A large experimental literature supports specific techniques — the worked-example effect, the modality effect, the split-attention and redundancy effects, and the expertise-reversal effect, in which supports that help novices (e.g., detailed worked examples) become redundant or harmful for experts. Kirschner, Sweller, and Clark (2006), in a widely cited and much-debated Educational Psychologist article, extended the argument to warn that minimally guided instruction ignores working-memory limits and tends to increase extraneous load for novices. Load itself can be measured — Paas''s subjective mental-effort rating is the standard self-report instrument — but the measurement only becomes interpretable once you know which kind of load a rating reflects.
Why it matters for course evaluation in practice
The single most common way course evaluations misuse the idea of difficulty is to treat it as a quality signal in either direction — as if "too hard" means bad teaching and "easy" means good teaching, or the reverse. CLT shows why both readings are wrong. Two courses can earn identical "this course was difficult" ratings for opposite reasons: one because the content has genuinely high intrinsic load (an advanced statistics module), the other because the design imposed needless extraneous load (disorganised slides, no worked examples, information split across incompatible sources). A raw difficulty item cannot tell these apart, so it cannot support any defensible action. This is closely related to the course-difficulty-and-workload confound that already distorts global ratings.
The productive move is to decompose difficulty into its designer-controllable component and ask about that directly. Extraneous-load sources translate cleanly into evaluation items a student can actually answer from experience:
- Sequencing and segmenting (manages intrinsic load): "New ideas were introduced in a sensible order, building on what came before."
- Split attention: "When a diagram was explained, the labels and explanation were easy to connect without flipping between sources."
- Worked examples: "Before I was asked to solve problems on my own, I saw fully worked examples."
- Redundancy: "Slides and narration reinforced each other rather than repeating identical text word-for-word."
These are low-inference teaching-behaviour items: they ask about observable design choices, not a global judgement of difficulty or clarity. They also connect to the deeper lesson of the feeling-of-learning gap — students'' sense of effort is not a direct readout of teaching quality, so an evaluation must ask about the mechanisms of good design rather than the felt sensation of ease. And because the expertise-reversal effect means the "right" amount of support depends on prior knowledge, evaluation data should be interpreted alongside where students sit on the novice–expert continuum, a point that overlaps with the case against one-size-fits-all learning-styles items.
There is one more implication. If perceived mental effort is collected — and it can be a useful signal — it must be interpreted, not scored. High effort in a well-designed course may be germane and desirable; high effort in a poorly designed one is extraneous and wasteful. A number without that interpretive frame invites exactly the wrong conclusion.
Limitations and honest caveats
A careful reader should hold several reservations. First, measuring the three load types separately is genuinely hard. The dominant instrument — a single subjective mental-effort rating — does not, by itself, distinguish intrinsic from extraneous from germane load; efforts to build differentiated self-report scales are promising but not settled. So an evaluation cannot simply "measure extraneous load"; it can only ask about the design features known to cause it.
Second, CLT''s strongest evidence comes from controlled experiments on well-structured tasks in mathematics, science, and technical domains. Its applicability to ill-structured, discussion-based, or creative disciplines is more contested, and importing its prescriptions wholesale into, say, a seminar in critical theory would be a category error.
Third, the theory has evolved and been criticised — the germane-load construct in particular was reformulated precisely because early versions risked being unfalsifiable, and the Kirschner–Sweller–Clark critique of inquiry learning drew substantial rebuttal. Treating CLT as settled dogma would misrepresent an active research programme.
Finally, the expertise-reversal effect means there is no universally optimal level of support, so evaluation items about worked examples or guidance must be read relative to the learners'' expertise rather than as absolutes. None of this undermines the core, defensible claim for evaluation design: difficulty is not one thing, and only its extraneous component is a design fault worth asking students to flag.
How Koji incorporates this
Koji is built to separate the difficulty that reflects genuine intellectual demand from the difficulty that reflects fixable design — the distinction CLT insists on.
- Decomposed, behaviourally specific items. Rather than a single "How difficult was this course?" scale, Koji''s structured question types (scale, single_choice, yes_no, open_ended) let institutions author separate items targeting the extraneous-load sources — sequencing, worked examples, split attention, redundancy — so a report can show which design mechanism, if any, was the problem.
- Conversational disambiguation of "it was hard." Because Koji runs an AI-moderated conversational interview, when a student rates a course as difficult the follow-up probe asks what made it hard: the ideas themselves, or how they were presented. That distinction — intrinsic versus extraneous — is exactly what a static form loses, and it is designed to mitigate the difficulty-confound rather than merely record it.
- Effort interpreted, not just scored. Where an institution collects perceived mental effort, Koji''s thematic analysis pairs the effort signal with open-text explanations, so high effort can be read as germane ("challenging but I understood why") or extraneous ("I spent an hour just figuring out where to look").
- Expertise-aware reporting. Koji can triangulate feedback across cohorts and prior-knowledge levels, so results are read with the expertise-reversal effect in mind rather than as a single verdict. The same AI-moderated interview engine powers Koji''s core research platform at koji.so, where distinguishing "hard because it''s complex" from "hard because it''s badly designed" is equally central to usability research.
Koji is designed to mitigate the difficulty confound, not to measure cognitive load directly; the interpretive judgement still rests with the QA team.
Related Resources
- Does Course Difficulty and Workload Lower Student Evaluations? What Centra Found
- Stop Asking "Was the Lecturer Clear?": The Case for Low-Inference Teaching-Behaviour Items
- Teacher Clarity Predicts Learning Better Than Charisma: What Course Evaluations Should Measure
- Students Rate the Classes Where They Learn the Most Lower: The Feeling-of-Learning Gap
- Should Course Evaluations Ask Whether Teaching Matched a Student''s Learning Style?
- The Fluency Illusion: Why a Polished Lecturer Earns Higher Course Evaluations Without Teaching More
References
- Sweller, J., van Merriënboer, J. J. G., & Paas, F. (2019). Cognitive Architecture and Instructional Design: 20 Years Later. Educational Psychology Review, 31(2), 261–292. https://doi.org/10.1007/s10648-019-09465-5
- Sweller, J. (1988). Cognitive load during problem solving: Effects on learning. Cognitive Science, 12(2), 257–285. https://doi.org/10.1207/s15516709cog1202_4
- Kirschner, P. A., Sweller, J., & Clark, R. E. (2006). Why Minimal Guidance During Instruction Does Not Work. Educational Psychologist, 41(2), 75–86. https://doi.org/10.1207/s15326985ep4102_1
- Paas, F., Tuovinen, J. E., Tabbers, H., & Van Gerven, P. W. M. (2003). Cognitive Load Measurement as a Means to Advance Cognitive Load Theory. Educational Psychologist, 38(1), 63–71. https://doi.org/10.1207/S15326985EP3801_8
Related articles
Should Course Evaluations Ask Whether Teaching Matched a Student's Learning Style?
Learning styles are one of the most durable myths in education, and Pashler et al. (2008) found no credible evidence for tailoring instruction to them. Here is why a course-evaluation item that asks students whether teaching 'suited their learning style' quietly measures a debunked construct — and what to ask instead.
Does Course Difficulty and Workload Lower Student Evaluations? What Centra Found
A research-grounded look at whether harder, heavier courses are punished in student evaluations of teaching. Centra (2003) and Marsh & Roche (2000) show the relationship is non-linear and weaker than faculty fear — with practical implications for fair course evaluation.
Stop Asking "Was the Lecturer Clear?": The Case for Low-Inference Teaching-Behaviour Items
High-inference items like "Is the instructor clear?" tell teachers nothing about what to change. Murray's low-inference behaviour research shows how concrete, observable items make evaluations diagnostic.
The Fluency Illusion: Why a Polished Lecturer Earns Higher Course Evaluations Without Teaching More
A research-grounded look at the lecture fluency illusion — why a smooth, confident presentation inflates student ratings and perceived learning while leaving actual learning unchanged — and what it means for interpreting course evaluations.