New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods9 min read

The Qualitatively Different Ways Students Experience Your Course: Phenomenography

A standard survey tells you how many students rated "feedback" a 3. Phenomenography, the higher-education research tradition founded by Ference Marton, instead maps the qualitatively different ways students experience a phenomenon — and explains why the same number means different things.

Koji Education Team

Product

In brief

Phenomenography is a qualitative research approach, developed inside higher education by Ference Marton and the Gothenburg group, that maps the small number of qualitatively different ways a group of people experience or understand a particular phenomenon — say "feedback", "group work", or "a difficult course". Its output is not a satisfaction score or a frequency count but an outcome space: a limited, logically related set of categories of description capturing the variation in how students make sense of the same thing. For course evaluation, phenomenography answers a question Likert scales cannot: not how much students liked the feedback, but in what fundamentally different ways they experienced it — which is exactly what you need before you can design good questions or act on a puzzling score.

What the research says

Marton (1981), writing in Instructional Science, named phenomenography and drew its defining distinction between two perspectives. From the first-order perspective a researcher describes the world directly ("how good is the feedback?"); from the second-order perspective the researcher describes people's experience of the world ("what are the different ways students experience the feedback?"). Phenomenography deliberately takes the second-order stance. Its aim is to find and systematise the qualitatively different ways in which people conceive of significant aspects of reality — and, crucially, to show that for almost any phenomenon there is only a limited number of such ways, usually three to six, that can be arranged into a logically structured outcome space.

The tradition grew from Marton and Säljö's (1976) landmark study "On qualitative differences in learning", in the British Journal of Educational Psychology. Asking students to read an academic text and then probing how they went about it, the authors identified two qualitatively distinct approaches: a deep approach oriented to the author's meaning and argument, and a surface approach oriented to the text as so many signs to be memorised. That deep/surface distinction — one of the most influential ideas in higher-education research — is itself a phenomenographic outcome space, and it shaped instruments such as the R-SPQ-2F.

Åkerlind (2005), in Higher Education Research & Development, gives the most-cited account of how the analysis is actually done. Data are open, semi-structured interviews with a purposive sample; the researcher reads whole transcripts before fragmenting them, works iteratively and comparatively across the pooled set of meanings rather than case by case, and brackets their own preconceptions so the categories emerge from the students' accounts. The categories are validated less by inter-rater agreement in the psychometric sense than by their communicability — whether other researchers, shown the categories and the data, find them recognisable and well-evidenced.

Why it matters for course evaluation in practice

Almost every course-evaluation instrument is first-order and frequency-driven: it fixes the constructs in advance ("the feedback was helpful: 1–5") and reports how the answers distribute. That is efficient for monitoring, but it hides two things a phenomenographic reading exposes.

First, the same rating conceals different experiences. Two students who both rate feedback a 3 may be experiencing completely different phenomena — one means "prompt but generic", another means "detailed but too late to use". Aggregating them into a mean treats these as the same signal. Phenomenography names the distinct conceptions so you can act on the right one.

Second, surveys miss conceptions the designer never imagined. Because the constructs are fixed, a standard form cannot surface a way of experiencing the course that its authors did not anticipate. Phenomenography, being open and second-order, is designed precisely to discover that variation — which then becomes the raw material for better, more valid questionnaire items. A practical workflow is therefore phenomenography first, scale later: interview a purposive sample to map the outcome space of how students experience a target aspect of the course, then build closed items that reflect the categories students actually use. This is the qualitative counterpart to the psychometric work covered elsewhere in this knowledge base, and it complements rather than replaces the numbers.

A concrete example makes the contrast vivid. Ask "how did you experience the group project?" on a survey and you might get a mean of 3.4 that satisfies a reporting requirement and tells you almost nothing. Interview the same students phenomenographically and you may recover four qualitatively distinct conceptions: the project experienced as a scheduling and logistics problem, as an opportunity to divide labour and minimise effort, as genuine intellectual collaboration, and as an unfair obligation to carry weaker peers. Each conception implies a different intervention — clearer scheduling, redesigned interdependence, better team formation, fairer individual accountability — and none of them is visible in the single number 3.4. That is the practical payoff of taking the second-order stance seriously.

Limitations and honest caveats

Phenomenography is a demanding method with real boundaries, and a critical reader should hold it to them. It is not about frequency or representativeness. The outcome space tells you what qualitatively different conceptions exist, not how many students hold each; you cannot benchmark departments or track a KPI with it, and using small phenomenographic samples to make prevalence claims is a category error. Analysis is interpretive and researcher-dependent. Despite bracketing, category formation involves judgement, and critics have long argued that different analysts can produce different outcome spaces from the same transcripts; the field's answer — communicability and transparent evidencing rather than a reliability coefficient — is a genuine, contested weakness. The deep/surface heritage has itself been criticised: approaches to learning are now understood as substantially context- and task-dependent rather than fixed student traits, so a conception surfaced in one course should not be read as a stable property of the student. It is slow and skill-intensive: good phenomenographic interviewing and analysis take trained researchers and time that routine end-of-term cycles rarely allow. Finally, because it is second-order, phenomenography describes experience, not learning outcomes — a student's way of experiencing a course is not evidence of what they learned. Use it to enrich and interpret evaluation, not to certify quality on its own.

How Koji incorporates this

Koji cannot — and should not claim to — replace a trained phenomenographer, but its architecture is unusually well suited to the second-order stance the method requires, and it lowers the cost of the interview-and-analyse workflow that usually makes phenomenography impractical at scale:

  • Natively second-order interviews. Koji's AI-moderated conversational interviews probe a student's own way of experiencing a topic — following up on meaning ("what did you mean by feedback being too late?") rather than forcing the answer onto a fixed scale. That is exactly the open, experience-focused data phenomenography analyses.
  • Surfacing variation, not just sentiment. Koji's automatic thematic analysis of open text can be oriented to map the range of distinct conceptions students express about a target aspect of the course, giving a QA analyst a first-pass structure of the variation to refine into categories of description.
  • From conceptions to instrument. Once the outcome space is validated by your team, Koji's structured question types (open_ended, single_choice, ranking) let you operationalise the discovered conceptions into closed items for the next cycle — the phenomenography-first, scale-later workflow, in one platform.
  • Honest division of labour. Koji organises and quality-scores the raw variation; forming and validating a rigorous, communicable outcome space remains expert human work. We frame the tool as accelerating phenomenographic analysis, not automating away the interpretation.

The same conversational engine underpins Koji's core research platform at koji.so, where product teams run essentially phenomenographic "jobs-to-be-done" interviews to map the different ways customers experience a product.

Frequently asked questions

How is phenomenography different from thematic analysis? Thematic analysis codes what is talked about across a dataset. Phenomenography specifically maps the qualitatively different ways of experiencing or understanding one phenomenon and arranges them into a logically related outcome space, taking a strict second-order stance (the student's experience, not the researcher's account of reality).

How many students do I need to interview? Phenomenographic studies typically use purposive samples of roughly 15–30 participants — enough to capture the range of variation, since the goal is to find the limited set of distinct conceptions, not to estimate how common each one is.

Can I turn a phenomenographic outcome space into survey questions? Yes, and this is one of its most practical uses. Mapping how students actually experience a topic gives you valid, student-grounded categories that you can convert into closed items for large-scale administration in the next evaluation cycle.

Does a phenomenographic finding tell me which department is best? No. The method is not about frequency or representativeness, so it cannot benchmark units or produce a KPI. It tells you what qualitatively different experiences exist, which you then interpret alongside quantitative evaluation.

Is the deep/surface distinction a fixed student type? No. Current evidence treats approaches to learning as strongly context- and task-dependent rather than stable traits, so a conception surfaced in one course should not be labelled as a permanent property of the student.

Related resources

References