The Framework Method for Open-Text Course Feedback: A Structured Alternative to Thematic Coding
When a quality committee — not a lone qualitative researcher — has to make sense of thousands of student comments, the Framework Method offers a transparent, auditable matrix-based approach. What Gale et al. (2013) actually proposed, and how to use it.
Koji Education Team
Product
Answer in brief. The Framework Method (Gale, Heath, Cameron, Rashid & Redwood 2013) is a systematic, matrix-based approach to analysing qualitative data that is especially suited to multidisciplinary teams and applied, policy-oriented questions — exactly the situation a quality-assurance committee faces when it must turn thousands of open-text course comments into decisions. Unlike open-ended thematic analysis, it produces a visible framework matrix (cases as rows, codes as columns, summarised data in the cells) that makes the path from raw comment to conclusion auditable. It is not a shortcut and not a substitute for interpretive skill, but for accountable, committee-based evaluation it is often a better fit than either informal reading or a black-box algorithm.
What the research says
Open-text comments are the richest part of a course evaluation and the least well used. Most institutions either read them impressionistically — which invites cherry-picking and negativity bias — or run them through automated topic modelling or sentiment analysis that is fast but hard to defend to a sceptical committee. The Framework Method sits deliberately between the two.
The anchor is Gale, N. K., Heath, G., Cameron, E., Rashid, S., & Redwood, S. (2013), "Using the framework method for the analysis of qualitative data in multi-disciplinary health research," BMC Medical Research Methodology, 13, 117. The method itself originated with Ritchie and Spencer (1994) at the UK''s National Centre for Social Research, developed explicitly for applied policy research — settings where findings must inform a decision on a timetable, and where the analysis team is not composed solely of experienced qualitative researchers. Gale et al. codified it into seven stages:
- Transcription — for course evaluations, the "transcripts" are the open-text responses, already in text form.
- Familiarisation — reading a sample of comments to get a feel for range and tone.
- Coding — labelling segments of text with what they are about (open coding), line by line at first.
- Developing a working analytical framework — agreeing a set of codes and categories after several coders have independently coded a subset and compared, reconciling differences.
- Applying the framework — indexing the remaining data using the agreed codes.
- Charting into the framework matrix — the defining step: a spreadsheet-like matrix with one row per case (e.g., per respondent or per cohort) and one column per code, with summarised data in each cell.
- Interpretation — reading across and down the matrix to identify themes, contrasts and explanations, typically in team discussion.
The paper''s central claim is not that this is the "truest" qualitative method — Gale et al. are careful to say it is one approach among many — but that its visibility and structure make it uniquely suitable when analysis must be shared across a team, defended to non-specialists, and audited later. The framework matrix is the artefact that makes each conclusion traceable back to the specific summarised evidence that supports it.
For context, the dominant alternative in the literature is Braun and Clarke''s (2006) reflexive thematic analysis, which is more flexible and interpretively rich but deliberately resists rigid procedure and inter-coder "reliability" framing. The Framework Method, by contrast, embraces multiple coders, explicit reconciliation, and a fixed matrix — a trade of interpretive fluidity for accountability and reproducibility. Neither is "better"; they answer different institutional needs. For committee-driven QA, the accountability of Framework usually wins.
Why it matters for course evaluation in practice
Quality assurance is a collective, accountable activity, not a solo research project, and that changes what a good analysis method looks like.
- It survives scrutiny. When a programme director disputes a conclusion — "students didn''t really complain about assessment" — the framework matrix lets you point to the exact cells and summarised comments that support the theme. Impressionistic reading cannot do this.
- It distributes work without losing consistency. Several administrators can index comments against an agreed framework, and the reconciliation step in stage 4 keeps them aligned. This scales to the volume real institutions face.
- It is naturally comparative. Because cases are rows, you can read across cohorts, campuses, or years and see where a theme is concentrated — the same structure that makes it easy to feed into inter-rater reliability checks.
- It closes the loop with evidence. The matrix is exactly the kind of transparent evidence trail that accreditation panels and self-evaluation reports reward, because it shows how student voice was analysed, not just what was concluded.
Limitations and honest caveats
A PhD reader should not mistake structure for rigour, and the method''s own proponents are the first to say so.
- The matrix can flatten meaning. Summarising each comment into a cell risks stripping context and nuance — the very things qualitative data exists to preserve. Gale et al. warn that charting must retain enough of the original to avoid decontextualisation.
- It is only as good as the framework. A poorly developed coding framework — built by one person, or fixed too early — bakes its blind spots into every subsequent step. The stage-4 reconciliation across multiple coders is not optional window-dressing; skipping it is where the method fails.
- It leans deductive. Because a working framework is agreed relatively early, the method can under-detect genuinely unexpected themes compared with a more inductive, reflexive approach. Analysts should keep an "other/emergent" code and revisit the framework.
- Reliability is contested in qualitative circles. Braun and Clarke and others argue that inter-coder agreement is a positivist import that mismeasures interpretive work. Institutions should present Framework results as transparent and auditable, not as "objective."
- Volume still bites. With tens of thousands of comments, manual charting is infeasible; the method then needs computational assistance, which reintroduces the validity questions of automated coding.
How Koji incorporates this
Koji for Education is built to produce the structured, case-indexed material the Framework Method needs, and to automate the mechanical stages without hiding the evidence trail.
- Case-structured data by design. Every Koji interview is already tied to a respondent and cohort, so the "cases as rows" structure of a framework matrix exists natively — you are not reconstructing it from a pile of anonymous comments.
- Automatic thematic indexing as a first pass. Koji''s thematic analysis proposes codes and clusters comments against them, doing the heavy lifting of stages 3–5 while leaving a human committee to develop and reconcile the working framework in stage 4 — the step the research says matters most.
- Traceability, not a black box. Koji is designed to keep each theme linked back to the specific verbatim responses that support it, so the output resembles a framework matrix (theme × cohort with underlying quotes) rather than an unaccountable score. This directly answers the "survives scrutiny" requirement.
- AI-moderated probing enriches the cells. Because Koji''s conversational interviews follow up on vague answers, the raw material charted into each cell is fuller and less ambiguous than a one-line survey comment — mitigating the decontextualisation risk.
- Human-in-the-loop framing. Koji positions automated coding as a first pass to be reviewed and reconciled by staff, consistent with the method''s insistence that the framework be team-developed, not machine-imposed. It is designed to support Framework-style analysis, not to replace the interpretive judgement stage 7 requires.
Koji''s core research platform at koji.so applies the same case-structured, thematically indexed engine to product and customer research, where framework-style analysis of open-ended interviews is a standard way to make qualitative findings defensible to stakeholders.
Related resources
- How Reliable Is Your Coding of Open-Text Feedback? Inter-Rater Reliability and Thematic Analysis
- How Many Open-Text Comments Are Enough? Thematic Saturation in Course Evaluations
- Topic Modeling for Open-Text Course Evaluations: LDA, STM and BERTopic Explained
- Can a Large Language Model Code Your Open-Text Course Feedback?
- Beyond "Positive or Negative": Aspect-Based Sentiment Analysis of Open-Text Course Feedback
- Why One Cruel Comment Outweighs Twenty Kind Ones: Negativity Bias in Reading Course Evaluations
Putting the Framework Method into practice
A workable course-evaluation implementation does not need to be elaborate. Start by pulling a stratified sample of comments across programmes and score bands and reading it cold to build familiarity. Have two or three staff independently open-code the same subset, then meet to reconcile a working framework — this is the single most important step, and the point at which disagreements about what a comment "means" are surfaced and resolved rather than buried. Keep the framework small (typically 15–25 codes grouped into four to six categories) and always retain an explicit "emergent/other" code so that genuinely new themes are not forced into pre-existing boxes. Only then index the full corpus and chart summaries into the matrix, preserving a short verbatim exemplar in each populated cell so the summary can be checked against its source. Finally, interpret in a group: read down each column to characterise a theme, and across each row to see how an individual cohort experienced the course. Documenting these steps — sample, coders, framework version, reconciliation notes — is itself the audit trail that makes the analysis defensible to a programme committee or an external reviewer.
References
- Gale, N. K., Heath, G., Cameron, E., Rashid, S., & Redwood, S. (2013). Using the framework method for the analysis of qualitative data in multi-disciplinary health research. BMC Medical Research Methodology, 13, 117. https://doi.org/10.1186/1471-2288-13-117
- Ritchie, J., & Spencer, L. (1994). Qualitative data analysis for applied policy research. In A. Bryman & R. G. Burgess (Eds.), Analyzing Qualitative Data (pp. 173–194). Routledge. https://doi.org/10.4324/9780203413081
- Braun, V., & Clarke, V. (2006). Using thematic analysis in psychology. Qualitative Research in Psychology, 3(2), 77–101. https://doi.org/10.1191/1478088706qp063oa
Related articles
How Reliable Is Your Coding of Open-Text Feedback? Inter-Rater Reliability and Thematic Analysis
When you turn thousands of free-text comments into themes and counts, how do you know the coding is trustworthy? O Connor and Joffe (2020) on intercoder reliability, Braun and Clarke on thematic analysis, and what rigorous qualitative QA looks like.
Why One Cruel Comment Outweighs Twenty Kind Ones: Negativity Bias in Reading Course Evaluations
Instructors and committees fixate on the harshest open-text comment and discount the praise. Baumeister''s "bad is stronger than good" and the negativity-bias literature explain why, and how to read qualitative course feedback fairly.
Can a Large Language Model Code Your Open-Text Course Feedback? What the Agreement Studies Show
Peer-reviewed evidence on how closely GPT-4-class models match human coders when categorising open-ended student comments — agreement statistics, where they fail, and how to use them responsibly in course-evaluation quality assurance.
Beyond "Positive or Negative": Aspect-Based Sentiment Analysis of Open-Text Course Feedback
Document-level sentiment scoring collapses a rich student comment into one number and hides the most useful signal. Aspect-based sentiment analysis (ABSA) separates *what* the comment is about from *how* the student feels about it. Here is the evidence on ABSA for course feedback, its accuracy, its failure modes, and how Koji applies it.