Codebook, Coding-Reliability, or Reflexive? Which Thematic Analysis Your Open-Text Course Feedback Actually Needs
When you analyse open-text student comments, you are choosing a paradigm — usually without knowing it. Braun and Clarke distinguish coding-reliability, codebook and reflexive thematic analysis, and they are not interchangeable. Which one your course feedback needs depends on what the analysis is for.
Koji Education Team
Product · July 31, 2026
Bottom line up front: Every time someone "themes" a pile of open-text course comments, they make a methodological choice — and most institutions make it by accident. Virginia Braun and Victoria Clarke, whose 2006 paper Using thematic analysis in psychology is among the most-cited academic articles ever written (over 220,000 Google Scholar citations by late 2024), argue that thematic analysis is not one method but a spectrum: coding-reliability TA, codebook TA, and reflexive TA (Braun & Clarke, 2021, Counselling and Psychotherapy Research). They rest on incompatible assumptions about what a "theme" even is. Choosing the wrong one for your purpose is why so much student-feedback analysis feels either shallow or unaccountable. Here is how to choose deliberately.
Three thematic analyses, not one
Coding-reliability TA treats coding as something to be done accurately. Two or more coders apply a pre-agreed codebook, and their agreement is quantified — typically with Cohen's kappa or a similar statistic. Themes are understood as patterns that reliable coders can consistently identify. This is the positivist-leaning end of the spectrum, and it is exactly the paradigm behind our own guidance on inter-coder reliability and Krippendorff's alpha. It prizes defensibility and comparability.
Codebook TA (approaches such as framework analysis and template analysis) sits in the middle. It uses a structured codebook — often mapped to research questions or, in a university, to evaluation domains — but treats coding as interpretive rather than as a reliability test. It suits applied, policy-facing work where you need structure and nuance.
Reflexive TA is Braun and Clarke's own approach and the opposite pole. Here, coding is inherently subjective and interpretive; themes are analytic outputs the researcher actively constructs at the intersection of the data, their theoretical stance and their skill — not objects that "emerge" from the data waiting to be found. Crucially, Braun and Clarke explicitly discourage reflexive-TA users from pursuing "accurate" coding, seeking consensus between coders, or reporting Cohen's kappa — because, in their framework, inter-rater reliability is not a coherent quality goal for interpretive work (Braun & Clarke, 2019/2021).
Read that last point carefully, because it means the two most common instincts in a quality office — "let's get a second coder to check agreement" and "let's do a rich interpretive read" — belong to different paradigms and cannot simply be bolted together.
Why this matters for course feedback specifically
The reason a sentiment score is not insight is that it collapses interpretation into a number. But the fix is not automatically "do rich reflexive analysis" — because the purpose of the analysis should decide the paradigm. Course feedback is analysed for at least two very different reasons:
-
Accountability and comparison. A quality committee flags courses of concern, compares programmes, and must defend decisions that affect people. Here you need consistency, an audit trail and defensibility — the strengths of coding-reliability or codebook TA. A committee cannot act on "one analyst's rich interpretation" if a second analyst would have read it differently.
-
Understanding and improvement. A teaching team wants to grasp why students experienced a redesign as they did. Here, forcing consensus can actively destroy value: the interesting reading is often the non-obvious one, and reliability statistics reward the surface reading two coders can most easily agree on. This is where reflexive TA earns its keep, and where the depth that focus-group-style methods offer over surveys lives.
The mistake institutions make is applying one paradigm to both jobs — usually demanding reliability everywhere, which flattens the improvement work, or doing loose interpretive reads everywhere, which leaves the accountability work indefensible.
But isn't reliability what makes analysis trustworthy?
This is the strongest objection, and it comes from a good place: if coding is "just subjective", how can anyone trust it? Two honest responses.
First, Braun and Clarke's argument is not that anything goes. It is that agreement is the wrong proxy for truth in interpretive work. Two coders converging on a code does not make the interpretation correct; it often just means both defaulted to the most obvious reading. Reflexive TA replaces the reliability criterion with a different, still-demanding quality standard: a documented, reflexive, internally coherent analysis with a transparent trail from data to theme, and researcher subjectivity treated as a resource to be examined rather than a bias to be eliminated. That is accountability of a different kind, not its absence.
Second — and this is the practical resolution — you do not have to pick one paradigm for your whole institution. Match the method to the stake. For a promotion case or a course-of-concern flag, where a decision must be defended, use codebook or coding-reliability TA and be honest about its shallower interpretive ceiling. For a teaching team's mid-cycle sense-making, use reflexive TA and judge it by reflexive-TA standards. What you must not do is smuggle one paradigm's authority into the other — quoting a kappa to make a rich interpretive claim sound objective, or presenting a loose thematic read as if it were audited.
The universal safeguard across all three paradigms is provenance: every theme should be traceable to the specific comments that produced it. That is the real trust mechanism, and it is why an AI theme you cannot trace back to a quote should not be trusted, regardless of which thematic analysis you claim to be doing.
Where AI changes the calculus — and where Koji fits
Large language models scramble this picture in a productive way. Historically, coding-reliability TA was expensive precisely because it required multiple trained human coders; reflexive depth was expensive because it required a skilled analyst reading everything. Whether AI can reliably analyse thousands of open-text comments is therefore a paradigm question, not just a capacity question.
Here is the honest positioning. A well-instrumented LLM applying a consistent scheme across thousands of comments behaves most like coding-reliability or codebook TA done at scale: it delivers the consistency and auditability that accountability work needs, without the cost of two human coders. What it does not do is replace the reflexive interpretive layer — the human insight into why a theme matters in a specific programme's context. Pretending otherwise would be exactly the paradigm-smuggling error above.
Koji for Education is built around that division of labour. Its AI-moderated conversational interviews gather richer open-text data in the first place — probing beyond a first answer across six structured question types — and its automatic thematic analysis produces consistent, quality-scored themes with the underlying student quotes attached to every one. That gives a quality committee the defensible, provenance-backed consistency that coding-reliability work demands, and it gives a teaching team a clean, quote-linked starting point for the reflexive interpretation only a human can add — instead of spending their scarce time hand-coding. In other words, Koji does the coding-reliability backbone so your people can do the reflexive thinking. It also standardises the moderation itself, removing the human-interviewer inconsistency that would otherwise contaminate the data before analysis even begins.
Teams that run qualitative research beyond the classroom will find the main Koji platform uses the same AI interview and thematic-analysis engine, so the paradigm discipline you build for course feedback carries into user and customer research too. And if the goal is genuinely to raise the quality of what students write, that is its own project — see student feedback literacy.
Stop "theming" open-text feedback on autopilot. Decide, deliberately, whether this analysis is for accountability or for understanding — then choose the thematic analysis that fits, and make every theme traceable to a quote.
A two-track policy you can actually run
The cleanest institutional response is to stop treating "thematic analysis" as one thing and write a two-track policy. Track one, accountability: any analysis that feeds a decision affecting a person — a course-of-concern flag, a promotion case, a programme review — uses codebook or coding-reliability TA, with a documented codebook, consistent application and every claim traceable to quotes. Its interpretive ceiling is lower, and that is an acceptable price for defensibility. Track two, understanding: any analysis whose purpose is a teaching team's own sense-making uses reflexive TA, judged by reflexive standards — coherence, transparency and reflexivity about the analyst's standpoint — not by kappa. Braun and Clarke's 2019 and 2021 refinements exist precisely because so many researchers claimed to do reflexive TA while importing reliability machinery that contradicts it. The same discipline protects a university: name which track you are on, apply its quality criteria, and never borrow the other track's authority. Provenance — theme-to-quote traceability — is the one requirement both tracks share, and the one your tooling should enforce automatically.
Want thematic analysis that is both defensible and deep? See how Koji analyses open-text feedback.