Can Students Tell You How Much They Learned? Self-Assessment Validity and the Dunning-Kruger Problem
Many course evaluations ask students to rate how much they learned. The evidence on whether students can accurately judge their own learning is humbling — and it has direct consequences for how you should interpret, and design, course feedback.
Koji Education Team
Product ·
Bottom line up front: A great many course-evaluation instruments ask some version of "How much did you learn in this course?" The implicit assumption is that students can accurately report their own learning. The evidence says they cannot, at least not well. Across meta-analyses, the correlation between people's self-evaluations of their ability and their actual measured performance is moderate at best (around r = .29), and self-assessments of learning correlate far more strongly with how satisfied and motivated students felt than with how much they actually learned. The least competent students tend to be the most overconfident — the Dunning-Kruger pattern. This does not make self-report useless, but it means a "perceived learning" item is measuring a real thing (the student's experience) that is not the same thing as learning. Treating it as a proxy for learning is a validity error, and it is one course evaluation makes constantly.
The question hiding inside your evaluation form
"I learned a great deal in this course — strongly agree to strongly disagree." Versions of this item appear on evaluation forms worldwide, and the resulting "perceived learning" score is often reported as if it were a learning outcome. It is not. It is a self-assessment of learning, and self-assessment is a measurement instrument with its own well-studied — and limited — validity.
What the evidence actually shows
Three strands of research converge on the same uncomfortable conclusion.
Self-evaluations of ability correlate only moderately with reality. Zell and Krizan's 2014 metasynthesis, "Do People Have Insight Into Their Abilities?", pooled 22 meta-analyses across domains including academic ability and found a mean correlation between self-evaluations and objective performance of just r = .29 (SD = .11). Insight was better when judgements were specific and tasks were objective and familiar — and worse otherwise. A general "how much did you learn?" item is exactly the broad, hard-to-anchor judgement where insight is weakest.
Self-assessed learning tracks affect, not cognition. Sitzmann and colleagues' 2010 meta-analysis in the Academy of Management Learning & Education, "Self-Assessment of Knowledge: A Cognitive Learning or Affective Measure?", found that self-assessment correlated most strongly with motivation (r = .59) and satisfaction (r = .51) — affective outcomes — but only moderately with actual cognitive learning (r = .34). Their sobering conclusion: even under conditions that optimised the link (students practised self-assessing and got feedback on their accuracy), self-assessment still tracked how students felt more than what they knew. When a course-evaluation form asks about perceived learning, it is largely measuring satisfaction wearing a learning costume.
The least competent are often the most overconfident. The Dunning-Kruger pattern shows up repeatedly in education. A 2024 study of first-semester medical students in BMC Medical Education found that 35.5% overestimated and 46.0% underestimated their performance, with a strong negative correlation (ρ = -0.590) between actual score and the accuracy of self-assessment — lower performers overestimated most. Higher achievers, by contrast, tended to underestimate. The direction of error is not random; it is tied to competence. A broader meta-analytic review of self-assessment scoring accuracy reaches a similar verdict: students tend to overestimate their work, and the overlap between student self-scores and expert scores is limited — though accuracy improves measurably when the task and the marking criteria are objective and when students assess after completing the work rather than predicting beforehand. The lesson is not that self-assessment is hopeless, but that its accuracy is conditional, and the broad "how much did you learn?" item meets almost none of those conditions.
Why this matters for course evaluation specifically
If self-assessed learning mostly measures satisfaction, then a "perceived learning" item is largely redundant with the overall-satisfaction item sitting two rows above it — and both are vulnerable to the same biases. Worse, it actively misleads when teaching that improves learning reduces the feeling of learning. That is the well-documented active-learning penalty: in controlled studies, students in active-learning conditions learn more but feel like they learned less, because the productive struggle is uncomfortable. An evaluation that trusts perceived learning will systematically punish exactly the teaching that works.
This is fundamentally a validity problem, in the sense we discuss in are course evaluations valid?. A perceived-learning item can be perfectly reliable — students answer it consistently — while having poor validity as a measure of actual learning. Reliability without validity is precision aimed at the wrong target.
Critics argue: "Students' perceptions are exactly what we want to measure"
This is the strongest counterargument and it deserves a real answer, because it is partly right. There is a defensible position that course evaluation should measure the student experience — perceived clarity, perceived support, perceived workload — and should not pretend to measure learning at all. On that view, asking what students perceived is not a flaw; it is the whole point. Student perceptions matter in their own right: they predict engagement, persistence, and word of mouth, and they capture aspects of teaching that test scores miss.
We agree with this — with one critical condition: label it honestly. The error is not asking about perceptions; the error is calling a perception a learning outcome and then using it to make decisions about teaching quality, promotion, or programme design as though it measured what was learned. Measure perceptions as perceptions. Measure learning, where you need to, with instruments built for it — assessment data, learning-gain measures, the kind of approach explored in measuring learning gain, not satisfaction. The two are both valuable and they are not interchangeable. And because student ratings are necessary but not sufficient evidence of teaching quality, they should be triangulated with other sources, as we argue in triangulation in teaching evaluation.
Where Koji fits
If the goal is to capture student perception well — honestly labelled as such — then the instrument should be good at eliciting rich, specific, self-aware reflection rather than a single noisy number. This is where Koji for Education helps. Rather than a flat "rate how much you learned" item, Koji's AI-moderated conversational interview can ask students to describe specific things they can now do that they could not before, where they got stuck, and what helped — questions that are anchored, concrete, and far more diagnostic than a global self-rating. The Zell-Krizan finding is instructive here: self-insight is better when judgements are specific and task-anchored, which is exactly the kind of judgement a good conversational probe elicits.
Koji's automatic thematic analysis then aggregates these reflections into patterns a programme team can act on, and its six structured question types let you keep a clean separation between perception questions (handled conversationally) and any genuine knowledge-check items, so you never quietly relabel one as the other. Its quality scoring flags thin or contradictory self-reports for appropriate weighting. The honest framing matters: Koji cannot make a student's self-assessment of learning accurate — no instrument can repeal Dunning-Kruger — but it can elicit perception data that is specific, self-aware, and correctly labelled, which is far more useful than a misattributed "perceived learning" average.
The same conversational engine powers the main Koji platform for user and customer research, where the gap between what people say they do and what they actually do is the central methodological challenge — so probing past surface self-report is exactly what the engine is built for.
The takeaway for programme directors
Look at your evaluation form and find every item that asks students to judge their own learning. Then decide, honestly, what you are doing with the answer. If you are treating it as evidence of how much students learned, stop — the research says it is mostly a satisfaction measure, and it is biased against your most effective teaching. If you are treating it as evidence of how students experienced the course, keep it, label it clearly, anchor it in specifics, and corroborate learning itself with assessment data. Students can tell you a great deal about their experience. What they cannot reliably tell you, on a five-point scale, is how much they learned.
Koji for Education helps you capture rich, honestly-labelled student perception — and keep it separate from claims about learning. See how Koji works.