Did Students Leave More Confident They Can Succeed? Academic Self-Efficacy as a Course-Evaluation Outcome
Academic self-efficacy is among the strongest psychological correlates of student achievement. Here is what the evidence says, why it belongs in course evaluation, its limits, and how to measure it without kidding yourself.
Koji Education Team
Product
In brief
Academic self-efficacy — a student's belief in their capacity to organise and execute the actions needed to succeed at academic tasks — is one of the most robust psychological predictors of university performance, which makes "did this course build students' confidence that they can do this kind of work?" a defensible evaluation question. In Richardson, Abraham and Bond's (2012) meta-analysis of 50 correlates of university GPA, performance self-efficacy emerged as the single strongest correlate; Honicke and Broadbent's (2016) systematic review of 59 studies found a moderate self-efficacy–performance relationship that is mediated by effort regulation and deep processing. The catch: self-efficacy is a belief, it is reciprocally entangled with prior success, and a course can inflate confidence without building competence. Measured carefully and triangulated, it is a valuable outcome signal; measured naively, it flatters.
What the research says
Self-efficacy is Albert Bandura's (1997) construct: not global self-esteem, but a task- and domain-specific judgement of one's ability to succeed at a particular class of activities. In education it is typically operationalised as academic self-efficacy — confidence in one's ability to master coursework, prepare for exams, and meet the demands of a subject. Bandura's theory specifies four sources: mastery experiences (past success, the strongest source), vicarious experience (seeing similar others succeed), verbal persuasion (credible encouragement), and physiological/affective states. Teaching acts on all four.
The evidence that this belief matters for outcomes is unusually strong for a psychological construct. Richardson, Abraham and Bond (2012), in a systematic review and meta-analysis published in Psychological Bulletin, synthesised 13 years of research and 1,105 independent correlations across 50 measures of university students' academic performance. Performance self-efficacy showed the largest correlation with GPA of any of the 50 constructs examined — ahead of high-school GPA and standardised test scores — and academic self-efficacy showed a medium-sized correlation, alongside grade goal and effort regulation. This is the finding that licenses treating self-efficacy as an outcome worth evaluating: it is not a peripheral nicety, it is among the best psychological predictors of who succeeds.
Honicke and Broadbent (2016), in Educational Research Review, ran a focused systematic review of 59 studies published 2003–2015 on academic self-efficacy and university performance. They reported a moderate overall correlation and, importantly, identified mechanisms: the self-efficacy–performance link is mediated by effort regulation, deep processing strategies, and goal orientation. In other words, believing you can succeed translates into achievement partly by changing how hard and how well you study. They also flagged a reciprocal relationship — performance feeds back into self-efficacy — and called for designs that untangle direction.
Two clarifications keep the construct honest. First, specificity matters: domain-matched self-efficacy ("I can succeed in statistics") predicts far better than diffuse self-esteem, and mis-specified, over-general items weaken the signal. Second, calibration matters: the productive target is accurate confidence, not maximal confidence. Overconfidence detached from competence predicts poor self-regulation, a caution Bandura himself and the metacognition literature both raise.
Why it matters for course evaluation in practice
Standard evaluations ask whether students were satisfied and whether the instructor was organised. Neither captures a durable, achievement-relevant change the course may have produced: whether students walk out believing — accurately — that they can do this kind of work and tackle the next course in the sequence. Because self-efficacy is domain-specific and built largely through mastery experiences, it is directly sensitive to pedagogy a programme can change: whether assessments were scaffolded so students accumulated genuine successes, whether feedback was specific enough to attribute success to effort and strategy, and whether the course modelled competence via worked examples and near-peer exemplars.
A self-efficacy item set gives quality-assurance staff a forward-looking, sequence-aware signal. A gateway module whose students leave with collapsed confidence in the discipline is a retention and progression risk even if its satisfaction scores are fine — self-efficacy is a documented predictor of persistence, not just grades. Conversely, a demanding course that raises accurate self-efficacy is doing something a satisfaction score can miss entirely, because effortful courses often depress momentary satisfaction while building competence and confidence. Measuring self-efficacy alongside satisfaction helps a programme distinguish "hard and demoralising" from "hard and empowering".
The cleanest design is a retrospective pre/post or an explicit before-and-after item ("At the start of this course, how confident were you that you could…?" / "Now?"), which captures change rather than a level confounded by who enrolled.
Limitations and honest caveats
It is a belief, and beliefs can be inflated. The central risk is that a course boosts confidence without building competence — a "feel-good" module that leaves students sure they understand material they cannot actually apply. Self-reported self-efficacy is vulnerable to exactly the overconfidence the calibration literature warns about, so a rising self-efficacy score is not self-validating; it must be checked against demonstrated performance.
Reciprocal causation and selection. Because performance drives self-efficacy as much as the reverse (Honicke & Broadbent, 2016), a cross-sectional post-course score partly reflects the grades students expect, not the teaching. High-achieving cohorts report high self-efficacy regardless of instruction, confounding cross-course comparison. Retrospective pre/post mitigates but does not eliminate this, and retrospective ratings carry their own response-shift and memory biases.
Construct and measurement specificity. Self-efficacy predicts well only when items are matched to the domain and to concrete tasks. Borrowing a generic confidence item, or conflating self-efficacy with self-esteem, outcome expectancy, or satisfaction, degrades validity. Effect sizes are moderate, not deterministic — many other factors drive achievement — so self-efficacy should be read as one informative correlate, not a sufficient statistic.
Generalisability. The anchor meta-analyses draw heavily on Western university samples; response styles, the social desirability of expressing confidence, and disciplinary norms vary cross-culturally, which matters for the multinational European context. As with any single construct, self-efficacy earns its place in an evaluation only as part of a triangulated picture.
How Koji incorporates this
Koji is designed to measure self-efficacy as a change and to guard against confusing confidence with competence.
- Domain-specific, calibrated items. Koji's
scaleandsingle_choicequestions let an evaluation carry properly specified, task-anchored self-efficacy items ("How confident are you that you could solve a problem like the ones in this course unaided?") rather than a generic confidence slider, preserving the specificity the research shows is essential. - Retrospective pre/post by design. Koji supports paired before/after framing so a programme captures self-efficacy gain attributable to the course, addressing the reciprocal-causation and selection confounds better than a single post-course level.
- Conversational probing for calibration. A high self-efficacy rating followed by an AI-moderated
open_endedprobe ("Describe a moment this term when you realised you could handle something you couldn't before") surfaces whether confidence is anchored in a concrete mastery experience or is free-floating — the calibration check a Likert number cannot perform. Koji's AI moderator is designed to ask for evidence, not just a number. - Thematic analysis mapped to sources of efficacy. Koji's automatic thematic analysis can cluster open-text responses by Bandura's sources — mastery, modelling, encouragement, anxiety — telling a course team which lever moved, and translating a summary score into an actionable design finding.
- Triangulation and honest framing. Koji's reporting is built to present self-efficacy alongside performance and satisfaction, so a confidence rise unaccompanied by demonstrated competence is visible rather than hidden. Koji is designed to mitigate, not eliminate, the inflation risk.
Koji's core research platform at koji.so applies the same AI-moderated interview engine to product and customer research, where perceived-competence and confidence measures play a parallel role in adoption studies.
Related Resources
- Growth Mindset as a Course-Evaluation Lens: What the Evidence Says
- Self-Regulated Learning and Metacognition as a Course-Evaluation Lens
- Self-Determination Theory: Autonomy, Competence, Relatedness in Course Evaluation
- Sense of Belonging as a Course-Evaluation Construct
- Reference Bias in Self-Reported Learning on Course Evaluations
- Response-Shift Bias and the Retrospective Pretest
Which way does the arrow point?
The reciprocal-causation problem is not merely a caveat; it has been studied directly. Talsma, Schüz, Schwarzer and Norris (2018), in a meta-analytic cross-lagged panel analysis in Learning and Individual Differences, modelled self-efficacy and performance across two time points in many samples. They found effects running in both directions, with the path from prior performance to later self-efficacy at least as strong as the reverse. The practical implication for evaluation is precise: a post-course self-efficacy level is partly a residue of grades students have already earned, so an institution serious about crediting a confidence gain to teaching must measure change against a baseline and compare across courses cautiously. This is also why Bandura's emphasis on mastery experiences as the dominant source is so actionable: a course builds durable self-efficacy chiefly by engineering genuine, progressively harder successes, not by encouragement alone. A good evaluation therefore probes not only the level of confidence but its source — was it earned through accomplishment the course arranged, or merely asserted?
References
- Richardson, M., Abraham, C., & Bond, R. (2012). Psychological correlates of university students' academic performance: A systematic review and meta-analysis. Psychological Bulletin, 138(2), 353–387. https://doi.org/10.1037/a0026838
- Honicke, T., & Broadbent, J. (2016). The influence of academic self-efficacy on academic performance: A systematic review. Educational Research Review, 17, 63–84. https://doi.org/10.1016/j.edurev.2015.11.002
- Bandura, A. (1997). Self-Efficacy: The Exercise of Control. W. H. Freeman.
- Talsma, K., Schüz, B., Schwarzer, R., & Norris, K. (2018). I believe, therefore I achieve (and vice versa): A meta-analytic cross-lagged panel analysis of self-efficacy and academic performance. Learning and Individual Differences, 61, 136–150. https://doi.org/10.1016/j.lindif.2017.11.015
Related articles
Why "I Learned a Lot" Can't Be Compared Across Courses: Reference Bias
Self-rated learning items are the backbone of most course evaluations, yet reference bias means students judge themselves against different implicit standards. We review the evidence that this distorts cross-group comparisons and what it means for benchmarking courses and programmes.
Response-Shift Bias: Why Self-Reported Learning Gains in Course Evaluations Can Mislead
When you ask students how much they improved, the course itself has changed the yardstick they use to answer. Response-shift bias, and the retrospective pre-test that corrects it, explained for evaluation committees.
Does Your Course Support Autonomy, Competence, and Relatedness? Self-Determination Theory as an Evaluation Lens
Self-Determination Theory says motivation depends on three basic needs — autonomy, competence, and relatedness. Here is how to evaluate a course by whether it feeds or starves those needs, instead of only whether students were satisfied.
Does Your Course Build Self-Regulated Learners? Metacognition and SRL as an Evaluation Lens
Self-regulated learning — the cycle of planning, monitoring and reflecting — predicts academic achievement, yet standard course evaluations never ask whether a course developed it. Here is the SRL evidence and how to turn it into evaluation questions.