Evaluating Active Learning: The ICAP Framework as a Course-Evaluation Lens
Most course evaluations ask whether a course was 'engaging' — a word that conflates enjoyment with learning. Chi and Wylie's (2014) ICAP framework replaces it with an observable ladder of cognitive engagement (Passive, Active, Constructive, Interactive), giving evaluation items that measure what students actually did.
Koji Education Team
Product
In short: Stop asking students whether a course was "engaging." The word conflates enjoyment and attention with the cognitive work that actually drives learning. Chi and Wylie''s (2014) ICAP framework defines four modes of engagement by observable behaviour — Passive, Active, Constructive, Interactive — and predicts that learning increases up that ladder (I > C > A > P). For course evaluation, that means asking what students did (Did you generate explanations? Discuss and build on peers'' ideas?) rather than how engaged they felt. Behaviour-based items are both more valid and more actionable than a global engagement rating, and they help separate genuine cognitive engagement from mere entertainment.
What the research says
The ICAP framework, introduced by Michelene Chi and Ruth Wylie (2014) in Educational Psychologist and building on Chi''s earlier (2009) differentiation of overt learning activities, is a theory of cognitive engagement grounded in observable behaviour. Instead of treating engagement as a single feeling, ICAP sorts learning activities into four modes according to what the learner overtly does:
- Passive — receiving information without any overt activity: listening to a lecture, watching a video, reading without annotating. The learner may be attentive, but no observable knowledge-transforming behaviour is occurring.
- Active — physically manipulating material: highlighting, copying notes verbatim, repeating steps, pausing and replaying a video. There is overt doing, but it does not go beyond the given information.
- Constructive — generating output that goes beyond what was presented: self-explaining, drawing a concept map, posing questions, comparing and contrasting, working a novel problem. The learner produces new ideas or inferences.
- Interactive — constructive activity carried out in dialogue, where partners build on each other''s contributions: genuine discussion, peer instruction, collaborative problem-solving in which each turn extends the other''s reasoning.
The central ICAP hypothesis is that these modes form an ordered predictor of learning: Interactive > Constructive > Active > Passive. The theory''s mechanism is cognitive — moving up the ladder engages progressively more knowledge-change processes, from storing (passive) to integrating (active) to inferring (constructive) to co-inferring (interactive). Chi and Wylie marshal evidence from laboratory and classroom studies supporting the ordering, and later empirical tests — for example Menekse, Stump, Krause, and Chi (2013) in engineering courses — found the predicted differences, with constructive and interactive activities outperforming active and passive ones on learning measures.
ICAP does not stand alone. It coheres with the largest body of evidence on active learning: Freeman and colleagues'' (2014) meta-analysis in PNAS, covering 225 studies across STEM, found that active-learning approaches raised examination performance by roughly half a standard deviation and substantially reduced failure rates relative to traditional lecturing. ICAP supplies the mechanism and gradient behind that headline: not all "active learning" is equal, and the constructive and interactive varieties are where the largest gains sit. What matters is what learners cognitively do — not whether the room looked busy.
Why it matters for course evaluation in practice
Almost every course-evaluation instrument contains a version of "This course was engaging" or "The instructor kept me interested." ICAP exposes why that item is weak. "Engaged" is a feeling word, and feelings of engagement track enjoyment, novelty, and instructor charisma at least as much as they track cognitive work. A charismatic lecturer can make a wholly passive class feel engaged — the same mechanism behind the Dr Fox effect and the fluency illusion, where students report high engagement and learning from a polished delivery that taught them little. An engagement rating therefore risks rewarding performance over pedagogy.
ICAP offers a repair: ask about behaviour, not feeling. Students are far more reliable reporters of what they did than of how engaged they were, and the four modes convert directly into concrete, answerable items:
- Passive vs Active: "In this course I regularly did more than listen and read — I worked problems, annotated, or applied the material during class."
- Constructive: "The course asked me to generate my own explanations, examples, questions, or diagrams, not just reproduce what was presented."
- Interactive: "I discussed ideas with peers in ways that built on each other''s thinking, not just divided up tasks."
These are low-inference behavioural items: they describe observable activity rather than a global judgement, so they are less contaminated by halo and charisma. Crucially, they also give the actionable information a global engagement score never can. A course that scores high on "engaging" but low on the constructive and interactive items is a diagnostic finding: students enjoyed a fundamentally passive experience. That is precisely the kind of insight that lets feedback actually improve teaching — it names the missing pedagogical move (add self-explanation prompts, peer instruction, generative tasks) rather than leaving the instructor to guess.
ICAP also complements two constructs already central to good evaluation. It sits naturally alongside the feeling-of-learning gap: constructive and interactive work often feels harder and less smooth than a fluent lecture, so an ICAP-informed instrument expects — and can contextualise — lower comfort ratings in the most effective classes. And for online and blended courses, ICAP pairs with the Community of Inquiry framework, giving a behavioural read on whether "interaction" in a discussion forum was genuinely co-constructive or merely parallel posting.
Limitations and honest caveats
Several caveats keep the framework honest. First, the ICAP ordering is a general tendency, not a guarantee. Interactive activity only outperforms constructive activity when the dialogue is genuinely co-constructive; two students taking turns without building on each other can be less effective than one student self-explaining well. Poorly designed group work can score "interactive" on a checklist while delivering passive learning. An evaluation item must therefore probe the quality of interaction, not just its presence.
Second, mode is not the only thing that matters. A brilliantly clear passive lecture can lay essential groundwork that later constructive work builds on; ICAP is a claim about engagement''s contribution to learning, not a mandate to abolish direct instruction. Reading an evaluation as "more interactive is always better regardless of content or sequence" would misapply the theory and ignores the cognitive-load reality that novices often need substantial guidance first.
Third, the framework relies on overt behaviour as a proxy for covert cognition. A student can self-explain silently (constructive but not observable) or perform a constructive-looking task mechanically. Self-report items inherit this slippage: a student may check "I generated my own explanations" for an activity they completed without real cognitive investment.
Fourth, much of the supporting evidence comes from STEM and well-structured tasks; the ordering''s applicability to studio, clinical, or humanities learning is plausible but less thoroughly tested. And self-reported behaviour, like all self-report, is subject to memory and social-desirability effects. None of this undermines the core, practical advantage for evaluation: a behavioural ladder is more valid and more actionable than a feeling of engagement.
How Koji incorporates this
Koji is designed to move evaluation from "did it feel engaging" to "what cognitive work did students actually do" — the shift ICAP demands.
- Behaviour-anchored engagement items. Instead of a single engagement scale, Koji''s structured question types (scale, single_choice, multiple_choice, yes_no) let institutions author separate items for the active, constructive, and interactive modes, producing an engagement profile rather than one contaminated number.
- Conversational probing of interaction quality. Because Koji runs an AI-moderated conversational interview, when a student reports discussion or group work, the follow-up probe asks whether peers built on each other''s ideas or merely split the task — surfacing the constructive-vs-parallel distinction that a checkbox misses. This is designed to mitigate the "interactive-in-name-only" problem the framework warns about.
- Separating entertainment from cognition. Koji''s thematic analysis of open text can distinguish comments about enjoyment and charisma from comments describing generative work, helping QA staff see when a high engagement feeling masks a passive experience — the Dr Fox pattern.
- Actionable, mode-specific reporting. Results map to specific pedagogical moves (add self-explanation prompts, peer instruction), so closing-the-loop actions are concrete. Koji''s core research platform at koji.so uses the same interview engine for product and customer research, where distinguishing what users did from what they say they felt is the same methodological discipline.
Koji is designed to mitigate the vagueness of engagement ratings, not to observe cognition directly; the behavioural signal still depends on honest self-report and thoughtful item design.
Related Resources
- Stop Asking "Was the Lecturer Clear?": The Case for Low-Inference Teaching-Behaviour Items
- Students Rate the Classes Where They Learn the Most Lower: The Feeling-of-Learning Gap
- Evaluating Online and Blended Teaching: The Community of Inquiry Framework
- The Fluency Illusion: Why a Polished Lecturer Earns Higher Course Evaluations Without Teaching More
- Do Student Evaluations Actually Improve Teaching? The Feedback-Intervention Evidence
- What Cognitive Load Theory Says Your Course Evaluation Should — and Shouldn''t — Ask
References
- Chi, M. T. H., & Wylie, R. (2014). The ICAP Framework: Linking Cognitive Engagement to Active Learning Outcomes. Educational Psychologist, 49(4), 219–243. https://doi.org/10.1080/00461520.2014.965823
- Chi, M. T. H. (2009). Active-Constructive-Interactive: A Conceptual Framework for Differentiating Learning Activities. Topics in Cognitive Science, 1(1), 73–105. https://doi.org/10.1111/j.1756-8765.2008.01005.x
- Freeman, S., Eddy, S. L., McDonough, M., Smith, M. K., Okoroafor, N., Jordt, H., & Wenderoth, M. P. (2014). Active learning increases student performance in science, engineering, and mathematics. Proceedings of the National Academy of Sciences, 111(23), 8410–8415. https://doi.org/10.1073/pnas.1319030111
- Menekse, M., Stump, G. S., Krause, S., & Chi, M. T. H. (2013). Differentiated Overt Learning Activities for Effective Instruction in Engineering Classrooms. Journal of Engineering Education, 102(3), 346–374. https://doi.org/10.1002/jee.20021
Related articles
Do Student Evaluations Actually Improve Teaching? The Feedback-Intervention Evidence
Student ratings handed back as a bare number barely change teaching (d ≈ 0.2). Paired with structured consultation, the same data produces moderate, durable improvement (d ≈ 0.6–0.7). What the intervention meta-analyses tell quality teams about closing the loop.
What Cognitive Load Theory Says Your Course Evaluation Should — and Shouldn't — Ask
Cognitive load theory (Sweller, van Merriënboer & Paas, 2019) distinguishes the unavoidable difficulty of content from difficulty caused by poor design. That distinction changes what a course evaluation should measure: not overall 'difficulty', but the design choices that impose or remove extraneous load.
Stop Asking "Was the Lecturer Clear?": The Case for Low-Inference Teaching-Behaviour Items
High-inference items like "Is the instructor clear?" tell teachers nothing about what to change. Murray's low-inference behaviour research shows how concrete, observable items make evaluations diagnostic.
The Fluency Illusion: Why a Polished Lecturer Earns Higher Course Evaluations Without Teaching More
A research-grounded look at the lecture fluency illusion — why a smooth, confident presentation inflates student ratings and perceived learning while leaving actual learning unchanged — and what it means for interpreting course evaluations.