Did Students Learn Something They Can Use Elsewhere? Transfer of Learning as a Course-Evaluation Lens
The ultimate test of a course is not whether students liked it or passed the exam, but whether the learning survives outside the classroom. Barnett and Ceci's taxonomy of transfer — and the sobering evidence that far transfer is hard — gives course evaluation a demanding, honest outcome to aim at.
Koji Education Team
Product
In brief
The deepest question a course evaluation can ask is not "were you satisfied?" or even "did you pass?" but "can you use what you learned somewhere other than the exam?" That is transfer of learning, and Barnett & Ceci (2002) gave the field a taxonomy for it: transfer ranges from near (a highly similar new context) to far (a context that differs in knowledge domain, physical setting, time, or modality). The evidence is sobering — far transfer is genuinely hard and often fails to appear — so treating transfer as an evaluation outcome forces a course to make an honest, high-bar claim rather than a satisfaction claim. You can capture early transfer signals in an evaluation through self-reported application, transfer-oriented items, and open-text examples, provided you stay honest about their limits.
What the research says
For a century "transfer" was argued about without a shared vocabulary, so studies talked past each other. Barnett, S. M., & Ceci, S. J. (2002), When and where do we apply what we learn? A taxonomy for far transfer (Psychological Bulletin, 128, 612–637), fixed this by specifying nine dimensions along which transfer can vary, grouped into two categories:
- Content — what transfers: the specific skill or knowledge, the performance change (speed, accuracy), and the memory demands.
- Context — when and where it transfers: knowledge domain, physical context, temporal context, functional context, social context, and modality.
Near transfer holds when the new situation is close to the original on these dimensions; far transfer requires application across contexts that differ substantially. Their central message: whether transfer "occurs" is unanswerable in the abstract — you must specify how far, on which dimensions.
The corroborating evidence is a useful corrective against over-claiming. Sala & Gobet (2017), Does Far Transfer Exist? Negative Evidence From Chess, Music, and Working Memory Training (Current Directions in Psychological Science, 26, 515–520), summarise meta-analyses showing that the apparent cognitive benefits of chess, music, and working-memory training shrink toward zero as study design improves (e.g. with active control groups). The general lesson — far transfer is rare and easily overstated — long predates them: Thorndike & Woodworth (1901) dismantled the "mental discipline" idea that studying Latin trains the mind generally, and Detterman (1993) argued that significant spontaneous far transfer is the exception, not the rule.
The more optimistic strand identifies the conditions under which transfer is likelier. Perkins & Salomon (1992) distinguished low-road transfer (automatic, from heavy practice in varied contexts) from high-road transfer (deliberate, via mindful abstraction of a principle), and argued transfer must be taught for — through varied examples, explicit principle-extraction, and practice at recognising when a principle applies. Transfer, in other words, is a design outcome, not a lucky by-product.
Why it matters for course evaluation in practice
Transfer reframes what a course is for and therefore what a good evaluation should look for:
- It is the outcome employers, accreditors, and later courses actually care about. A programme-level quality claim ("our graduates can apply X") is a transfer claim. End-of-term satisfaction is silent on it, and even exam performance often only demonstrates near transfer to a very similar problem format.
- It exposes teaching that optimises for the test. A course can produce excellent exam scores through practice on near-identical problems while producing no far transfer at all — the well-known gap between assessment performance and durable, flexible competence. A transfer lens catches this where a satisfaction or grade lens does not.
- It sets a realistic, honest bar. Because far transfer is hard, a transfer-oriented evaluation should ask about specified, moderate transfer ("Have you used this technique in another module / a placement / a real task?") rather than grandiose "learn to think" claims the evidence does not support.
- It links to how the course was taught. If Perkins & Salomon are right that transfer must be designed for, then transfer items and questions about the teaching that supports transfer (varied practice, principle-extraction, authentic tasks) belong in the same instrument — connecting evaluation to constructive alignment and the ICAP active-learning framework.
Limitations and honest caveats
A PhD reader will rightly press on several points:
- Self-reported transfer is weak evidence. Asking students whether they have applied their learning elsewhere is subject to recall error, social desirability, and the difficulty of noticing one's own transfer. It is a leading indicator worth collecting, not proof; behavioural or performance evidence in the new context is the gold standard, and it is expensive.
- Timing undercuts the end-of-term survey. Real far transfer often only becomes observable months or years later, in a later course or a job. A survey administered in the last week of term is measuring, at best, intention and near transfer — a genuine ceiling on what any single course evaluation can claim.
- Far transfer may be genuinely rare. The negative evidence (Sala & Gobet; Detterman) means an honest instrument should expect modest, domain-close transfer and be sceptical of large "general skills" gains. Absence of far transfer is not necessarily a teaching failure.
- Attribution is hard. Even when a student transfers a skill, isolating this course as the cause — rather than prior knowledge, a co-requisite, or a placement — is a confounding problem no self-report resolves.
- Construct breadth. "Transfer" spans nine dimensions; a short evaluation can probe only a few, so results must be reported as evidence about specific transfer, not transfer in general.
How Koji incorporates this
Koji is an AI-native course-evaluation platform, and a transfer lens plays to its strength at going beyond the Likert number:
- Transfer-specific structured items. Koji supports
yes_noandsingle_choiceitems that ask whether and where a skill has been applied (another module, a placement, a personal project), andscaleitems for confidence in applying it in an unfamiliar context — operationalising Barnett & Ceci's near-versus-far distinction rather than a vague "I learned a lot". - Open-text examples, analysed at scale. The strongest self-report transfer evidence is a concrete example ("I used the sampling logic from this course to design my thesis survey"). Koji's AI-moderated conversational interview probes for that specific instance — "Can you describe a situation outside this course where you used it?" — and its automatic thematic analysis clusters the examples by how far the transfer context sits from the original, giving a distribution of near-to-far applications instead of an average.
- Delayed and mid-cycle collection. Because transfer surfaces over time, Koji supports follow-up collection after the course ends — a delayed pulse to a cohort in a later term — so the platform can capture transfer evidence closer to when it actually occurs, not only in the final week.
- Honest framing in reporting. Koji's bias-aware reporting is designed to label self-reported transfer as a leading indicator with stated limits, so a programme does not overclaim far transfer on the strength of a survey — consistent with the evidence that far transfer is hard.
- Connecting outcome to design. Paired with items on varied practice and principle-extraction, Koji lets a programme see whether the teaching conditions Perkins & Salomon associate with transfer were present, turning a transfer finding into an actionable design signal via closing-the-loop action tracking.
Koji's core research platform at koji.so applies the same AI-moderated interview engine to product and customer research, where the parallel question — do users apply a feature in real workflows, not just in onboarding — is the commercial cousin of far transfer.
Designing a transfer-oriented evaluation, dimension by dimension
Barnett & Ceci's taxonomy is not just descriptive — it is a checklist for writing better items. Instead of a vague "I can apply what I learned", specify the transfer distance you are asking about:
- Knowledge domain. "Have you used a method from this course in a different subject (a project, another module)?" probes cross-domain transfer, the hardest and most valuable kind.
- Physical and functional context. "Have you applied it outside a coursework setting — in a placement, a job, or a personal task?" separates classroom-bound skill from portable skill.
- Temporal context. A follow-up pulse weeks or a term later asks about durability, which an end-of-term item cannot. Even a single delayed question meaningfully raises the evidential bar.
- Modality. "Could you explain the principle to someone else, or apply it in a new format (a presentation, a real dataset)?" tests whether the learning is tied to the exam format it was practised in.
Pair these with two design-condition items — whether the course used varied practice contexts and whether it made the underlying principle explicit — because Perkins & Salomon's low-road and high-road routes predict exactly which teaching produces transfer. Reported together, they let a programme say not just "did transfer happen?" but "did we teach in the way that makes transfer likely?" — a far more actionable finding than a satisfaction average, and one that maps directly onto programme-level quality evidence.
Related resources
- Constructive alignment as an evaluation lens
- The ICAP framework for evaluating active learning
- Desirable difficulties and the satisfaction-learning gap
- Deep and surface approaches to learning (R-SPQ-2F)
- Self-regulated learning and metacognition as a lens
- Institution-level reporting for quality audits
References
- Barnett, S. M., & Ceci, S. J. (2002). When and where do we apply what we learn? A taxonomy for far transfer. Psychological Bulletin, 128(4), 612–637. https://doi.org/10.1037/0033-2909.128.4.612
- Sala, G., & Gobet, F. (2017). Does far transfer exist? Negative evidence from chess, music, and working memory training. Current Directions in Psychological Science, 26(6), 515–520. https://doi.org/10.1177/0963721417712760
- Perkins, D. N., & Salomon, G. (1992). Transfer of learning. In International Encyclopedia of Education (2nd ed.). Pergamon Press.
- Thorndike, E. L., & Woodworth, R. S. (1901). The influence of improvement in one mental function upon the efficiency of other functions. Psychological Review, 8(3), 247–261. https://doi.org/10.1037/h0074898
- Detterman, D. K. (1993). The case for the prosecution: Transfer as an epiphenomenon. In D. K. Detterman & R. J. Sternberg (Eds.), Transfer on Trial: Intelligence, Cognition, and Instruction (pp. 1–24). Ablex.
Related articles
Institution-Level Evaluation Reporting for Quality Audits: Building Longitudinal, Cross-Programme Evidence
A buyer's guide to turning course-evaluation data into institution-level evidence for ESG-aligned quality audits and institutional review — mapping ESG Part 1 and Part 2 expectations to concrete, longitudinal, cross-programme outputs.
Is Your Course Pushing Students Toward Deep or Surface Learning? The R-SPQ-2F as an Evaluation Lens
Most course evaluations ask whether students liked the teaching. The deep/surface approaches tradition asks a more consequential question: did the course lead students to engage meaningfully or just memorise to pass? The R-SPQ-2F instrument makes that measurable.
Evaluating Active Learning: The ICAP Framework as a Course-Evaluation Lens
Most course evaluations ask whether a course was 'engaging' — a word that conflates enjoyment with learning. Chi and Wylie's (2014) ICAP framework replaces it with an observable ladder of cognitive engagement (Passive, Active, Constructive, Interactive), giving evaluation items that measure what students actually did.
Desirable Difficulties: Why the Teaching That Improves Learning Often Lowers Satisfaction
Spacing, interleaving, and retrieval practice are among the best-evidenced ways to make learning durable — and they make a course feel harder and less smooth in the moment. Bjork & Bjork's (2011) desirable-difficulties principle explains why end-of-term satisfaction ratings systematically penalise the most effective teaching.