Is Your Course Pushing Students Toward Deep or Surface Learning? The R-SPQ-2F as an Evaluation Lens
Most course evaluations ask whether students liked the teaching. The deep/surface approaches tradition asks a more consequential question: did the course lead students to engage meaningfully or just memorise to pass? The R-SPQ-2F instrument makes that measurable.
Koji Education Team
Product
In short
Whether students adopt a deep approach (seeking meaning, connecting ideas) or a surface approach (memorising to meet minimum requirements) is not a fixed trait — it is strongly shaped by how a course is taught and assessed, and it can be measured. The Revised Two-Factor Study Process Questionnaire (R-SPQ-2F) by Biggs, Kember and Leung (2001) captures this with 20 items. For a quality-assurance office, adding an approaches-to-learning lens answers a question ordinary satisfaction surveys cannot: not "did students like it?" but "did the course lead students to engage in the kind of learning it claims to develop?" That is closer to the outcome universities actually care about.
What the research says
The distinction originates with Ference Marton and Roger Säljö (1976), whose study "On qualitative differences in learning" in the British Journal of Educational Psychology asked Swedish university students to read academic prose and then examined both what they took from it and how they went about reading it. They identified two qualitatively different processes: deep-level processing, in which the learner engages with the meaning and argument of the text, and surface-level processing, in which the learner focuses on the text itself and on reproducing it. Crucially, the same student could be induced toward either process by the demands they anticipated — the approach was a response to context, not a stable characteristic.
This grew into the influential approaches to learning tradition. John Biggs and colleagues operationalised it in the Study Process Questionnaire and then refined it into the R-SPQ-2F (Biggs, Kember & Leung, 2001), a 20-item instrument with two main scales — Deep Approach and Surface Approach — each split into motive (why a student studies as they do) and strategy (how they study). Deep motive is intrinsic interest and the intention to understand; surface motive is fear of failure and the intention to do the minimum. In the original validation, the scales showed acceptable reliability (Cronbach's alpha of roughly 0.73 for Deep and 0.64 for Surface) and confirmatory factor analysis supported the intended two-factor structure. Biggs explicitly presented the R-SPQ-2F as a tool teachers can use to evaluate the effect of their own teaching environment on how students learn.
The evaluation-relevant claim is not merely that these approaches exist, but that teaching and assessment shape them. Trigwell, Prosser and Waterhouse (1999), in Higher Education, found that in classes where teachers reported a more student-focused, conceptual-change approach to teaching, students were more likely to report a deep approach and less likely to report a surface approach — direct evidence that the learning approach is a function of the environment. This links to Biggs's wider concept of constructive alignment (Biggs, 1996): when learning outcomes, teaching activities and assessment are aligned to reward understanding rather than reproduction, students are steered toward deep approaches; when assessment rewards memorisation, even able students rationally adopt surface strategies. The approaches-to-learning framework therefore reframes course evaluation from measuring satisfaction to measuring whether the course elicited the learning behaviour it was designed to elicit.
This connects directly to a theme running through the course-evaluation evidence base: student enjoyment and student learning can diverge. The feeling-of-learning gap and the wider SET-and-learning meta-analytic evidence both show that ratings of teaching correlate weakly with actual learning. Measuring approaches to learning is one principled way to look past liking toward learning-relevant behaviour.
Why it matters for course evaluation in practice
A conventional course evaluation is a satisfaction instrument. It asks whether the lecturer was clear, whether the materials were useful, whether the student would recommend the module. These are legitimate questions, but they do not tell a programme director whether the course achieved its deeper educational aim: getting students to think like historians, engineers or clinicians rather than to memorise and forget. An approaches-to-learning lens adds exactly that missing dimension:
- It aligns evaluation with mission. Most programme specifications promise to develop critical thinking, synthesis and independent learning — all deep-approach behaviours. Measuring the deep/surface balance tells you whether the taught reality matches the promise, which is precisely the kind of evidence outcome-focused accreditation frameworks value.
- It diagnoses assessment problems that satisfaction surveys hide. A module can score highly on satisfaction while quietly driving surface learning, because cramming for a well-signposted exam feels comfortable. A rise in surface-approach scores is an early warning that assessment is rewarding reproduction — a diagnosis a satisfaction item will never surface.
- It complements self-reported learning gains. Instruments like SALG and the Course Experience Questionnaire already move beyond satisfaction; the R-SPQ-2F is complementary, focusing specifically on the approach students took rather than the gains they claim.
- It supports before-and-after evaluation of teaching change. Because approaches respond to the environment, the R-SPQ-2F can be administered before and after a redesign to test whether an intervention (say, shifting from an exam to a project) actually moved students toward deeper engagement.
In short, the framework helps a QA office evaluate the right thing — learning-relevant behaviour — rather than a proxy (liking) that the evidence shows is only loosely connected to learning.
Limitations and honest caveats
The approaches-to-learning tradition is influential but genuinely contested, and a credible evaluation office should present it with its caveats:
- The two-factor structure is not universally replicated. While Biggs, Kember and Leung reported good fit, subsequent studies have found the factor structure of the R-SPQ-2F less clean in some populations, and cross-cultural and disciplinary validation studies report varying reliability — the Surface scale in particular sometimes shows lower internal consistency. The instrument is not a settled, context-free measure.
- Self-report of "approach" is itself fallible. The R-SPQ-2F asks students to describe how they study, which is subject to the same social-desirability and self-insight limits as any self-report; students may over-report deep motives.
- "Deep good, surface bad" is an oversimplification. Some memorisation is a legitimate and necessary part of many disciplines (anatomy, languages, law). A high surface score is not automatically a failure, and the framework has been criticised for implying a value hierarchy that does not always hold.
- Approach is not the same as achievement. A deep approach is associated with higher-quality learning outcomes but does not guarantee them, and the correlations reported in the literature are moderate, not deterministic.
- What Marton and Säljö "really said" is debated. Later scholarship (including reflections on the original Göteborg work) notes that "levels of processing" in a reading experiment and "approaches to learning" across a whole degree are not identical constructs, and the popularised version has drifted from the careful original. Claims should respect that nuance.
These caveats argue for using approaches-to-learning as one lens among several, triangulated with satisfaction, learning-gains and outcome data — not as a single new number to rank courses by.
How Koji incorporates this
Koji for Education is well suited to bringing an approaches-to-learning lens into routine evaluation, because it is built around structured constructs and probing conversation rather than a single satisfaction score:
- Validated multi-item constructs, not one global rating. Koji supports scale items that can carry a validated construct such as the deep/surface motive and strategy subscales, so a QA team can embed an approaches-to-learning short form alongside their standard questions rather than reducing everything to "overall satisfaction".
- AI-moderated probing beyond the Likert number. A student who marks "I only study what's on the exam" can be gently probed by Koji's AI moderator about why — surfacing whether the assessment design is driving surface strategies. This conversational depth is designed to distinguish a genuine surface approach from a mis-checked box, addressing the self-report limitation above.
- Automatic thematic analysis links approach to cause. Koji's thematic analysis of open text can connect reported surface strategies to specific course features (assessment format, workload, signposting), giving a programme team an actionable diagnosis rather than an abstract score.
- Before-and-after and cohort triangulation. Because Koji administers evaluations consistently and segments by cohort, an approaches measure can be tracked across a teaching redesign or compared across parallel cohorts, supporting the causal reading the framework invites — while Koji's reporting keeps the measure alongside satisfaction and learning-gains evidence rather than in isolation.
- Honest framing. Koji is designed to help teams measure and interpret the deep/surface balance, not to certify a course as producing "deep learning" — consistent with the research message that approach is a contextual, imperfectly measured behaviour, not a verdict.
Universities that also run learning-behaviour or engagement research beyond individual courses can apply the same construct-driven, AI-moderated approach through Koji's core research platform at koji.so.
Related resources
- Beyond the Lecturer: What the Course Experience Questionnaire (CEQ) Measures
- SALG: Can Students Reliably Report Their Own Learning Gains?
- Students Rate the Classes Where They Learn the Most Lower: The Feeling-of-Learning Gap
- Teacher Clarity Predicts Learning Better Than Charisma
- Do Student Evaluations Measure Learning? The Uttl Meta-Analysis Revisited
References
- Marton, F., & Säljö, R. (1976). On qualitative differences in learning: I—Outcome and process. British Journal of Educational Psychology, 46(1), 4–11. https://doi.org/10.1111/j.2044-8279.1976.tb02980.x
- Biggs, J., Kember, D., & Leung, D. Y. P. (2001). The revised two-factor Study Process Questionnaire: R-SPQ-2F. British Journal of Educational Psychology, 71(1), 133–149. https://doi.org/10.1348/000709901158433
- Trigwell, K., Prosser, M., & Waterhouse, F. (1999). Relations between teachers' approaches to teaching and students' approaches to learning. Higher Education, 37(1), 57–70. https://doi.org/10.1023/A:1003548313194
- Biggs, J. (1996). Enhancing teaching through constructive alignment. Higher Education, 32(3), 347–364. https://doi.org/10.1007/BF00138871
Related articles
Do Student Evaluations Measure Learning? The Uttl Meta-Analysis Revisited
A research-grounded reading of the Uttl, White & Gonzalez (2017) meta-analysis on the SET–learning relationship, with implications for European course evaluation and quality-assurance policy.
Why "I Learned a Lot" Can't Be Compared Across Courses: Reference Bias
Self-rated learning items are the backbone of most course evaluations, yet reference bias means students judge themselves against different implicit standards. We review the evidence that this distorts cross-group comparisons and what it means for benchmarking courses and programmes.
Beyond the Lecturer: What the Course Experience Questionnaire (CEQ) Measures and Why It Predicts Learning
Ramsden's Course Experience Questionnaire reframed evaluation around the learning environment, not the instructor's personality. We unpack what the CEQ measures, the evidence that its scales predict deep learning and outcomes, its limitations, and how Koji operationalises learning-environment evaluation.
Students Rate the Classes Where They Learn the Most Lower: The Feeling-of-Learning Gap
A randomised Harvard experiment found students learned more in active classrooms but rated their own learning lower. What the feeling-of-learning gap means for interpreting course-evaluation items that ask how much students learned.