Where Did Students Get Stuck? Threshold Concepts as a Course-Evaluation Lens
Meyer and Land's threshold concepts reframe course evaluation from satisfaction to whether students crossed the transformative, troublesome ideas a discipline turns on — and got unstuck from liminality.
Koji Education Team
Product
In brief
Threshold concepts are the small number of ideas in any discipline that, once grasped, transform how a student sees the whole subject — and that students get stuck on for weeks in an uncomfortable "liminal" state before they cross. Meyer and Land (2003, 2005) argue these transformative, troublesome, irreversible ideas are where the real learning drama happens. Using them as a course-evaluation lens reframes the question from "were you satisfied?" to "where did you get stuck, and did you get through?" — but threshold concepts are notoriously hard to identify reliably, so the lens sharpens your questions rather than delivering a validated scale.
What the research says
Meyer and Land introduced the idea of a threshold concept in a 2003 report from the UK Economic and Social Research Council's Enhancing Teaching-Learning Environments project, later developed in Higher Education (2005). A threshold concept is a portal: opportunity cost in economics, the limit in calculus, deconstruction in literary theory, or "signification" in communication studies. It has, in their account, several defining features. It is transformative — grasping it shifts how the learner perceives the whole subject and often themselves. It is integrative — it exposes the hidden interrelatedness of things that previously looked separate. It is typically irreversible — once you think like an economist about opportunity cost you cannot easily un-see it. It is often bounded, marking the frontier of one conceptual space against another. And it is characteristically troublesome, in the sense Perkins (1999) gave the term: the knowledge is counter-intuitive, alien, or conceptually difficult, and it resists the learner.
The most important contribution of the 2005 paper is the notion of liminality. Borrowing from anthropology, Meyer and Land describe learners as spending extended time in an in-between, "stuck" state, oscillating between a naïve prior understanding and the fuller understanding their tutors expect. In this liminal space students may adopt mimicry — reproducing the language of the concept without owning it — and they experience the frustration, anxiety and identity-disruption that troublesome knowledge provokes. The crossing is not smooth; it is a threshold precisely because it is a place of difficulty, and getting stuck there is normal, not a sign of a bad student.
Cousin (2006) gives the accessible introduction that popularised the framework across disciplines, and stresses that teaching for threshold concepts means tolerating and supporting the mess of liminality rather than smoothing it away. Barradell (2013) provides the necessary counterweight: a review of how threshold concepts are actually identified, concluding that the process is theoretically contested and methodologically fraught. There is no agreed procedure; researchers have used student interviews, analysis of exam scripts, observation, and expert consensus, and different methods surface different candidate concepts. This matters enormously for anyone wanting to evaluate against threshold concepts, because the lens is only as good as the discipline's ability to name its own thresholds — and that naming is difficult, expert-dependent work.
Why it matters for course evaluation in practice
Most course evaluations measure a diffuse satisfaction that is only loosely coupled to learning — a recurring theme in the evidence, from the feeling-of-learning gap to the case for desirable difficulties. The threshold-concepts lens offers a different target. Instead of asking whether the course was clear and enjoyable, it asks whether students crossed the specific conceptual thresholds the course exists to help them cross, and where exactly they got stuck.
That reframing has three practical payoffs. First, it locates difficulty. A programme team that knows opportunity cost is a threshold for first-year economists can ask directly about that idea, and a cluster of students marooned in its liminal space is a precise, actionable signal — far more useful than a low mean on "the pace was about right." Second, it normalises struggle. Because liminality predicts that good teaching of a hard idea will feel effortful and even frustrating, a dip in satisfaction around the threshold is expected and should not be read as failure — which directly aligns evaluation with how learning actually works rather than against it. Third, it distinguishes mimicry from mastery: an evaluation designed around thresholds probes whether students can use the concept in an unfamiliar setting, not merely restate it, which connects to the deep-versus-surface distinction in how students approach learning. A programme that surfaces where its thresholds sit can redesign the curriculum around them rather than sprinkling generic study support.
Limitations and honest caveats
The framework is generative but should be handled with a critical eye.
Identification is the Achilles' heel. Barradell (2013) is explicit that there is no reliable, agreed method for determining what a discipline's threshold concepts are. Two expert panels in the same field can nominate different concepts, and student-derived and expert-derived lists frequently diverge. An evaluation built on a mis-identified "threshold" measures the wrong thing confidently. The prudent move is to treat candidate thresholds as hypotheses to be tested against student data, not as settled facts.
It is a lens, not a validated instrument. Unlike a psychometric scale with published reliability and validity, "threshold concepts" is a conceptual framework. There is no standard questionnaire, no norm data, and no established way to score "how liminal" a cohort is. Anyone importing it into evaluation is doing qualitative, interpretive work and should report it as such rather than dressing it in false quantitative precision.
Some critics doubt the construct itself. A strand of the literature questions whether "threshold concept" is definable sharply enough to be operationalised, or whether the five features are post-hoc labels applied to any hard idea. The concept of irreversibility, in particular, is difficult to evidence. A balanced evaluation acknowledges that threshold concepts organise attention usefully without claiming they are a precise, falsifiable measurement model.
It risks over-attribution. Not every struggling student is in liminal transit toward a threshold; some are simply under-supported, unwell, or short on prior knowledge. Reading every difficulty as productive liminality can excuse genuinely poor teaching. The lens must sit alongside, not replace, the ordinary diagnostics of workload, clarity and support.
How Koji incorporates this
Koji's design is well matched to a threshold-concepts approach precisely because that approach is conversational and diagnostic rather than a fixed scale. Instead of a Likert battery, a programme can configure Koji's AI-moderated interview to probe named candidate thresholds — "Tell me about a moment in this course where an idea finally clicked, or where you felt stuck for a while" — and the moderator's follow-up probes are designed to distinguish genuine crossing from mimicry by asking students to apply the idea, not restate it. The open-text responses are then run through automatic thematic analysis, which can cluster the places where a cohort reports getting stuck, turning scattered anecdotes about liminality into a mapped, programme-level signal of where the thresholds bite.
Because Koji supports structured question types alongside open text, a team can pair a direct scale item on a specific concept with the conversational probe, and triangulate the two. The platform is designed to report these as diagnostic patterns — "a substantial group described being stuck on marginal analysis in weeks 5–7" — rather than as a satisfaction score, which is exactly the reframing the framework calls for. None of this is presented as measuring threshold-crossing definitively; it is designed to help a programme locate and investigate its own thresholds, consistent with Barradell's warning that identification is hard and provisional. The same AI-moderated interview engine underpins Koji's core research platform at koji.so, where product teams use identical "where did it click / where did you get stuck" probing to find the conceptual sticking points in onboarding and documentation.
Frequently asked questions
What exactly is a threshold concept?
A threshold concept is a discipline's transformative gateway idea: once a student truly grasps it, their view of the whole subject shifts, and the change is usually irreversible. Meyer and Land characterise these concepts as transformative, integrative, often bounded, and typically troublesome — counter-intuitive and hard to accept. Opportunity cost in economics and the limit in calculus are stock examples. They are few in number but pivotal, which is why they are worth evaluating against.
What is liminality and why does it matter for evaluation?
Liminality is the extended "stuck" state learners occupy while crossing a threshold, oscillating between an old, naïve understanding and the fuller one their tutors expect. It matters because it predicts that good teaching of a hard idea will often feel frustrating and lower short-term satisfaction. An evaluation informed by liminality reads a dip around a difficult concept as expected transit, not as a teaching failure, and looks for whether students eventually crossed.
How is this different from just measuring satisfaction?
Satisfaction is diffuse and only loosely tied to learning; the threshold lens targets specific conceptual crossings the course exists to enable. Rather than "was the course clear and enjoyable?", it asks "did you get through opportunity cost, and where did you get stuck?" That produces a precise, actionable map of conceptual bottlenecks instead of a single mean that hides where the real difficulty lived.
Can I build a validated questionnaire around threshold concepts?
Not in the psychometric sense. Threshold concepts are a conceptual framework, not a scale with published reliability and norms, and Barradell (2013) shows that even identifying a discipline's thresholds is contested. The honest approach is qualitative and interpretive: treat candidate thresholds as hypotheses, probe them with open questions, and report the findings as diagnostic patterns rather than validated scores.
How do I identify my course's threshold concepts?
There is no single agreed method. Common approaches combine expert nomination by experienced teachers, analysis of where students repeatedly fail or get confused (exam scripts, assessment data), and student interviews about moments of being stuck or of sudden insight. Because these methods can disagree, triangulate at least two and treat the resulting list as provisional, revisiting it as evidence accumulates.
Doesn't this risk excusing bad teaching as "productive struggle"?
Yes, and that is a real danger. Not every stuck student is in productive liminal transit — some are simply under-supported or short on prerequisites. The lens should sit alongside ordinary diagnostics of workload, clarity and support, not replace them. Use it to interpret difficulty around known thresholds, while still treating widespread confusion elsewhere as a signal that something in the teaching needs fixing.
References
- Meyer, J. H. F., & Land, R. (2003). Threshold concepts and troublesome knowledge: Linkages to ways of thinking and practising within the disciplines. In C. Rust (Ed.), Improving Student Learning: Theory and Practice — Ten Years On (pp. 412–424). Oxford: Oxford Centre for Staff and Learning Development.
- Meyer, J. H. F., & Land, R. (2005). Threshold concepts and troublesome knowledge (2): Epistemological considerations and a conceptual framework for teaching and learning. Higher Education, 49(3), 373–388. https://doi.org/10.1007/s10734-004-6779-5
- Perkins, D. (1999). The many faces of constructivism. Educational Leadership, 57(3), 6–11.
- Cousin, G. (2006). An introduction to threshold concepts. Planet, 17(1), 4–5. https://doi.org/10.11120/plan.2006.00170004
- Barradell, S. (2013). The identification of threshold concepts: A review of theoretical complexities and methodological challenges. Higher Education, 65(2), 265–276. https://doi.org/10.1007/s10734-012-9542-3
Related resources
- Desirable Difficulties: Why the Teaching That Improves Learning Often Lowers Satisfaction
- Students Rate the Classes Where They Learn the Most Lower: The Feeling-of-Learning Gap
- Is Your Course Actually Aligned? Using Biggs's Constructive Alignment as an Evaluation Lens
- Is Your Course Pushing Students Toward Deep or Surface Learning? The R-SPQ-2F as an Evaluation Lens
- What Cognitive Load Theory Says Your Course Evaluation Should — and Shouldn't — Ask
- The Qualitatively Different Ways Students Experience Your Course: Phenomenography
Related articles
Students Rate the Classes Where They Learn the Most Lower: The Feeling-of-Learning Gap
A randomised Harvard experiment found students learned more in active classrooms but rated their own learning lower. What the feeling-of-learning gap means for interpreting course-evaluation items that ask how much students learned.
Is Your Course Pushing Students Toward Deep or Surface Learning? The R-SPQ-2F as an Evaluation Lens
Most course evaluations ask whether students liked the teaching. The deep/surface approaches tradition asks a more consequential question: did the course lead students to engage meaningfully or just memorise to pass? The R-SPQ-2F instrument makes that measurable.
What Cognitive Load Theory Says Your Course Evaluation Should — and Shouldn't — Ask
Cognitive load theory (Sweller, van Merriënboer & Paas, 2019) distinguishes the unavoidable difficulty of content from difficulty caused by poor design. That distinction changes what a course evaluation should measure: not overall 'difficulty', but the design choices that impose or remove extraneous load.
Desirable Difficulties: Why the Teaching That Improves Learning Often Lowers Satisfaction
Spacing, interleaving, and retrieval practice are among the best-evidenced ways to make learning durable — and they make a course feel harder and less smooth in the moment. Bjork & Bjork's (2011) desirable-difficulties principle explains why end-of-term satisfaction ratings systematically penalise the most effective teaching.