The Instrument That Actually Measures Learning: Concept Inventories, Normalized Gain, and What Course Evaluation Can't See
Your course evaluation asks students how much they think they learned. A concept inventory measures whether they actually did — and the two often disagree. What the normalized-gain research means for quality assurance.
Koji Education Team
Product · July 10, 2026
Bottom line up front: A course evaluation cannot tell you whether students learned anything. It can tell you whether they felt they learned, enjoyed the module, or rated the lecturer — and decades of evidence show those perceptions correlate only loosely, and sometimes negatively, with actual knowledge gain. There is a different class of instrument that does measure learning directly: the concept inventory, a validated pre/post test scored with Richard Hake's normalized gain. It will not replace course evaluation, and it does not measure the same thing. But understanding what it measures — and what your Likert-scale survey structurally cannot — is essential to any honest claim that a programme is improving teaching. This essay explains the method, the landmark evidence, its limits, and how a well-designed evaluation should sit beside direct measurement rather than pretend to be it.
The problem: satisfaction is not learning
The core confusion in course evaluation is category error. A student rating of "I learned a great deal in this course" is a perception of learning, filtered through enjoyment, effort, confidence, and the lecturer's fluency. It is not a measurement of learning. And the gap between the two is not small.
The most striking recent demonstration comes from Louis Deslauriers and colleagues, writing in the Proceedings of the National Academy of Sciences (2019, vol. 116, no. 39, pp. 19251–19257). They randomly assigned students in introductory physics to identical content taught either by active learning or by a polished passive lecture. Students in the active classrooms learned measurably more — yet rated their feeling of learning lower than peers in the passive lecture. The authors' conclusion is a warning shot for quality assurance: evaluating teaching on students' perception of learning can systematically favour inferior, passive pedagogy, because a fluent lecture feels like understanding even when less of it sticks. We treat this dynamic at length in The Active-Learning Penalty: Why Your Best Teaching Can Score Worst, and it is the same self-assessment fallibility explored in Can Students Tell You How Much They Learned? Self-Assessment Validity and the Dunning-Kruger Problem.
If perception of learning can move in the opposite direction to real learning, then no amount of refining your satisfaction survey will make it a measure of learning. You need a different instrument.
What a concept inventory is
A concept inventory is a rigorously validated test of conceptual understanding in a specific domain. The archetype is the Force Concept Inventory (FCI), developed by David Hestenes, Malcolm Wells, and Gregg Swackhamer (The Physics Teacher, 1992). It is a set of multiple-choice questions on Newtonian mechanics whose wrong answers are carefully engineered to match the common misconceptions students actually hold. A student who has memorised formulae but never restructured their intuitive physics will reliably pick the seductive distractor. Because the instrument targets deep conceptual change rather than recall, and because it is given both before and after instruction, it measures what the teaching changed — not what the student walked in already knowing.
Concept inventories now exist far beyond physics — in chemistry, biology, statistics, geoscience, engineering, and computer science — though physics remains the most mature. The design discipline is the point: distractors validated against real student thinking, items screened for reliability, and a construct defined tightly enough that a score means something specific.
Normalized gain: the metric that made courses comparable
Giving a pre-test and a post-test is only half the method. The question is how to score the change. A raw gain (post minus pre) unfairly penalises classes that started with well-prepared students, because there is less room to improve — a ceiling effect. Richard Hake's answer, in one of the most-cited papers in physics education research, was the normalized gain, defined as:
g = (post-test % − pre-test %) ÷ (100% − pre-test %)
In plain terms: of the improvement that was available to a class, what fraction did it actually achieve? A class averaging 40% on the pre-test and 70% on the post-test captured 30 of a possible 60 points — a normalized gain of 0.50 — regardless of how high or low they started.
Hake's 1998 study ("Interactive-engagement versus traditional methods," American Journal of Physics, vol. 66, no. 1, pp. 64–74) applied this to pre/post data from 62 introductory physics courses enrolling 6,542 students. The result reframed the field. The 14 traditional courses (N = 2,084) that made little use of interactive engagement achieved an average normalized gain of just 0.23 ± 0.04. The 48 interactive-engagement courses (N = 4,458) achieved 0.48 ± 0.14 — almost two standard deviations higher. The teaching method, not the students' starting ability, drove the difference. Crucially, this signal is invisible to a satisfaction survey; indeed, given the Deslauriers finding, the higher-gain courses might well have scored worse on "I felt I learned a lot."
That is the whole argument in miniature. Direct measurement found a large, reproducible effect of pedagogy on learning that student ratings would have missed or inverted.
But doesn't this just replace one flawed number with another?
The strongest objections to concept inventories and normalized gain are real, and any honest advocate has to concede them.
Critics argue, first, that normalized gain has statistical problems. It is not perfectly independent of pre-test score across the full range, it can behave badly for very high pre-test classes, and some researchers (notably Jennifer Coletta and colleagues, and various measurement critiques) prefer effect sizes, Rasch-modelled measures, or covariate-adjusted scores. This is a fair critique of the metric, and modern physics education research increasingly reports gain alongside effect sizes rather than instead of them. It is a reason to use the tool carefully, not a reason to return to satisfaction as a learning proxy.
Second, concept inventories are narrow by design. An FCI measures Newtonian mechanics concepts — not laboratory skill, not problem-solving transfer, not the ability to write, collaborate, or think like a professional. A course is far more than its concept inventory. Reducing "did they learn?" to a single validated test risks the very construct-underrepresentation that we warn about in Construct-Irrelevant Variance: The Validity Problem Underneath Course-Evaluation Bias. Many disciplines have no validated inventory at all, and building one is expensive and slow.
Third, and most practically, direct measurement is heavier. It requires a validated instrument, two testing occasions, and analysis capacity that most teaching-and-learning offices do not have for every module. You cannot run an FCI on a third-year seminar in comparative literature.
These are genuine limits. They establish that concept inventories are a complement, not a universal replacement — which is exactly the point. The mistake is not using concept inventories too little; it is asking course evaluation to do a job it was never built for and then treating the answer as evidence of learning.
Triangulation, not substitution
The defensible position is the one accreditation frameworks increasingly demand: direct measures of learning and indirect measures of experience answer different questions and should be read together. We lay out that distinction in Direct vs Indirect Measures: Why Your Course Evaluation Is Not Evidence of Learning, and the broader logic of combining evidence sources in Triangulation in Teaching Evaluation: Why Student Ratings Are Necessary but Not Sufficient.
A concept inventory tells you whether the concepts landed. A course evaluation tells you how the experience felt, where students struggled, and why — the qualitative texture that a pre/post score cannot supply. A gain of 0.25 is a flag; it is the evaluation that tells you students found the pace impossible, the assessment misaligned, or the prerequisite assumed but never taught. Read together, they are diagnosis plus explanation. Read alone, each is half-blind.
Where Koji fits
Koji does not administer concept inventories, and it would be dishonest to imply it measures learning gain directly — it does not. What it does is make the indirect half of the triangle far more diagnostic than a Likert grid, so that when a direct measure flags a problem, the evaluation actually explains it.
Concretely: Koji's AI-moderated conversational interviews probe beyond a satisfaction number to surface the mechanisms behind a weak learning gain — which concepts felt confusing, where the workload broke down, whether assessment tested what was taught. Its six structured question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) let you ask targeted self-report items — for example, confidence against specific learning outcomes, aligned to the constructive-alignment logic in Does Your Course Evaluation Ask Whether Students Actually Met the Learning Outcomes? — while its automatic thematic analysis turns thousands of open-text comments into the "why" behind a number. Bias-aware, standardized AI moderation keeps that probing consistent across cohorts, and programme-level reporting lets you place experiential evidence next to whatever direct measures your discipline supports.
The design principle Koji shares with the concept-inventory tradition is intellectual honesty about what a number means: a satisfaction average is not a learning measure, a normalized gain is not a satisfaction measure, and a serious quality process needs both, labelled correctly. Teams doing this outside teaching — measuring what customers actually understood versus what they felt about an onboarding flow — hit the same perception-versus-reality gap, which is why the same conversational engine powers the main Koji platform for product and UX research.
The takeaway
If your quality process claims that teaching improved, ask which instrument supports the claim. If the answer is a rising satisfaction score, you have evidence about perception, not learning — and the Deslauriers result shows those can diverge. Concept inventories scored with normalized gain are the discipline of measuring learning directly, and Hake's 6,542-student result shows how much that discipline can reveal. Use them where they exist; use a genuinely diagnostic evaluation to explain what they find; and never let a Likert average stand in for evidence that students actually learned.