Should Course Evaluations Ask Whether a Course Fostered a Growth Mindset? What the Meta-Analytic Evidence Says
Growth-mindset interventions show weak average effects in two large meta-analyses. Here is what that means for whether — and how — course evaluations should ask about mindset.
Koji Education Team
Product
In brief
Two large meta-analyses (Sisk et al., 2018) found that a student's growth mindset explains only about 1% of the variance in academic achievement, and that mindset interventions produce small average effects — meaningful mainly for lower-achieving and disadvantaged students. The practical implication for course evaluation is cautionary: do not treat "fostered a growth mindset" as a proxy for teaching quality or learning gains. There is, however, a narrower and defensible use — asking whether the learning environment signalled that ability is developable, framed as one facet of climate rather than an outcome measure.
What the research says
The idea that students hold either a "fixed" theory of intelligence (ability is stable) or a "growth" theory (ability is developable), and that the growth view drives achievement, comes from Carol Dweck's programme of work, popularised in Mindset: The New Psychology of Success (Dweck, 2006). The claim became one of the most widely adopted ideas in education, spawning school-wide interventions and, inevitably, survey items.
The most rigorous synthesis of the evidence is Sisk, Burgoyne, Sun, Butler, and Macnamara (2018), published in Psychological Science. The authors ran two separate meta-analyses. The first examined the correlation between mindset and academic achievement across 273 effect sizes from 365,915 students; the second examined the causal effect of mindset interventions across 43 studies with 57,155 students.
Both results were sobering for strong-form mindset claims. The average correlation between mindset and achievement was very weak — mindset accounted for roughly 1% of the variance in academic outcomes. The average intervention effect was also small (a standardised mean difference near d = 0.08 overall). Crucially, the effects were moderated: interventions were more useful for students who were academically at risk or low in socioeconomic status, and the mindset–achievement association was stronger for children and adolescents than for adults.
This does not mean mindset is nothing. The largest and best-designed intervention study to date — Yeager et al. (2019), Nature, the National Study of Learning Mindsets — delivered a short (under one hour) online growth-mindset intervention to a nationally representative sample of over 12,000 U.S. ninth-graders. It found a real but targeted effect: improved grades among lower-achieving students and increased enrolment in advanced mathematics, but only where the school's peer norms supported the mindset message. The headline is not "mindset transforms achievement" but "a light-touch mindset message helps some students in some contexts."
Read together, the evidence supports a modest, conditional conclusion: mindset is a genuine construct with small, context-dependent effects, not a master variable that separates good courses from bad ones.
Why it matters for course evaluation in practice
Course evaluations routinely borrow constructs from the learning-sciences literature and turn them into Likert items — "This course helped me believe I can improve with effort." The mindset evidence is a warning about doing this uncritically, for three reasons.
1. A weak-effect construct makes a poor quality signal. If mindset explains ~1% of achievement variance, then an item measuring perceived mindset cannot carry much weight as evidence that a course "worked." Aggregating a low-signal item into a summative teaching score adds noise, not information, and risks penalising instructors for a variable largely outside their control.
2. The moderation matters more than the main effect. The interesting finding is for whom mindset messaging helps — at-risk and lower-SES students in supportive peer contexts. A single course-average mindset score hides exactly this variation. If you ask about mindset at all, the analysis must be disaggregated by student subgroup, or it will mislead. This is a general lesson: course-average means routinely conceal subgroup reversals (see our note on Simpson's Paradox).
3. Perceived climate is not the same as caused change. A student reporting that a course "made me feel my ability can grow" is describing a perception of the learning environment — legitimately part of course climate, and adjacent to belonging and self-determination constructs. It is not evidence that the course changed the student's implicit theory of intelligence, still less that it changed achievement. Keeping this distinction explicit prevents overclaiming.
The defensible use, then, is narrow: treat mindset-supportive climate as one descriptive facet of the learning environment — how the course framed struggle, error, and improvement — reported alongside belonging and motivation, never as a standalone outcome or a component of a high-stakes teaching score. This mirrors the logic behind treating self-regulated learning as an evaluation lens rather than a verdict.
Limitations and honest caveats
A critical reader should hold several caveats in view.
- Meta-analytic averages hide heterogeneity. Sisk et al. themselves report substantial variability; the "1%" is an average, not a universal ceiling. Some well-targeted interventions produce larger effects, as Yeager et al. (2019) show. Averaging across heterogeneous designs can understate what a well-implemented programme achieves for the right students.
- The debate is genuinely contested. Proponents argue that many studies in the meta-analyses used weak or unfaithful implementations of mindset interventions, diluting the true effect. This is a live methodological dispute about intervention fidelity, not a settled fact — and readers should treat both the strong and the deflationary claims with caution.
- Measurement of mindset is itself imperfect. Self-report mindset scales are susceptible to social desirability and to students answering how they think they should think. An evaluation item inherits all these measurement problems.
- Correlational climate items cannot establish causation. Even a well-worded item captures perception at one moment; it cannot tell you the course caused any change without a pre/post or comparison design.
- Cultural generalizability is limited. Much of the mindset evidence is U.S.-based; its transfer to diverse European higher-education contexts is not guaranteed and should be treated as an open question.
Naming these limits is not a reason to ignore mindset — it is the reason to use it modestly and descriptively rather than as a headline metric.
How Koji incorporates this
Koji for Education is built so that a weak-effect construct like mindset can be explored without being smuggled into a summative score. Concretely:
- Conversational probing instead of a lone Likert item. Rather than reducing mindset to one agree/disagree statement, Koji's AI-moderated interview can ask an
open_endedquestion — "When you got stuck in this course, what did the teaching signal about whether you could improve?" — and follow up adaptively. This captures the climate the mindset literature actually points to, rather than a self-labelled "theory of intelligence." - Structured question types kept in their lane. Mindset-climate questions are configured as descriptive
scaleoropen_endeditems tagged to the learning-environment section, never mixed into the overall teaching-quality index. Koji's reporting keeps climate facets visually and statistically separate from summative ratings, so a low-signal construct does not contaminate a high-stakes number. - Bias-aware, disaggregated reporting. Because the mindset effect lives in the moderators, Koji's analytics are designed to report results by cohort segment where sample sizes and privacy thresholds permit, surfacing subgroup patterns instead of a single course mean that hides them.
- Automatic thematic analysis of open text. Koji clusters free-text responses into themes, so "the instructor treated mistakes as part of learning" surfaces as a recurring qualitative signal — evidence about environment, framed as description, not as a claim that achievement changed.
- Honest framing by design. Koji is designed to mitigate the overclaiming risk by labelling climate constructs as perceptions of the learning environment; it does not assert, and does not let a report assert, that a course "changed students' mindset" or improved achievement through mindset.
Beyond the classroom, Koji's core research platform at koji.so applies the same AI-moderated interview engine to product and customer research — useful when an institution wants to study staff or stakeholder beliefs with the same rigour it applies to students.
Related resources
- Self-Regulated Learning and Metacognition as an Evaluation Lens
- Sense of Belonging as a Course-Evaluation Construct
- Achievement Emotions and Control-Value Theory in Course Evaluation
- Desirable Difficulties: Why Better Teaching Can Lower Satisfaction
- Feeling of Learning vs Actual Learning in Active Classrooms
- Simpson's Paradox in Course-Evaluation Data
References
- Sisk, V. F., Burgoyne, A. P., Sun, J., Butler, J. L., & Macnamara, B. N. (2018). To what extent and under which circumstances are growth mind-sets important to academic achievement? Two meta-analyses. Psychological Science, 29(4), 549–571. https://doi.org/10.1177/0956797617739704
- Yeager, D. S., Hanselman, P., Walton, G. M., Murray, J. S., Crosnoe, R., Muller, C., et al. (2019). A national experiment reveals where a growth mindset improves achievement. Nature, 573, 364–369. https://doi.org/10.1038/s41586-019-1466-y
- Dweck, C. S. (2006). Mindset: The New Psychology of Success. Random House.
Related articles
When Combining Sections Reverses the Result: Simpson's Paradox in Course-Evaluation Data
Aggregating course-evaluation scores across sections, cohorts, or years can reverse the very conclusion you are trying to draw. What Simpson's paradox is, how it appears in evaluation data, and how to report so the reversal cannot bite you.
Your Course Evaluation Measures Satisfaction, Not Emotion. Pekrun Says That's a Problem
Standard course evaluations ask whether students were satisfied. Pekrun's control-value theory and the Achievement Emotions Questionnaire show that discrete emotions — enjoyment, boredom, anxiety, hope, hopelessness — drive learning and are absent from almost every institutional survey. Here is why that gap matters and how to close it.
Desirable Difficulties: Why the Teaching That Improves Learning Often Lowers Satisfaction
Spacing, interleaving, and retrieval practice are among the best-evidenced ways to make learning durable — and they make a course feel harder and less smooth in the moment. Bjork & Bjork's (2011) desirable-difficulties principle explains why end-of-term satisfaction ratings systematically penalise the most effective teaching.
Does Your Course Build Self-Regulated Learners? Metacognition and SRL as an Evaluation Lens
Self-regulated learning — the cycle of planning, monitoring and reflecting — predicts academic achievement, yet standard course evaluations never ask whether a course developed it. Here is the SRL evidence and how to turn it into evaluation questions.