New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
analysis-reporting10 min read

How Reliable Is Your Coding of Open-Text Feedback? Inter-Rater Reliability and Thematic Analysis

When you turn thousands of free-text comments into themes and counts, how do you know the coding is trustworthy? O Connor and Joffe (2020) on intercoder reliability, Braun and Clarke on thematic analysis, and what rigorous qualitative QA looks like.

Koji Education Team

Product

In short: Open-text comments are the richest part of a course evaluation, but turning them into themes, counts, and "students said X" claims is an act of interpretation that can be done well or badly. The qualitative-methods literature gives two anchors: O Connor and Joffe (2020) on when and how to use intercoder reliability (ICR), and Braun and Clarke (2006) on doing thematic analysis systematically. The defensible position is that themes reported to an accreditation panel should be produced by a transparent, documented coding process — ideally with a reliability check on a sample — not by one reader skimming comments and reporting impressions.

What the research says

Quantitative scores get all the psychometric scrutiny, but the qualitative half of an evaluation — the free-text box — is where students explain why. The problem is that "we read the comments and the main themes were workload and feedback" is a claim with no stated method behind it. Two bodies of work tell you how to make such claims trustworthy.

Thematic analysis as a method. Virginia Braun and Victoria Clarke's "Using thematic analysis in psychology" (Qualitative Research in Psychology, 2006, 3(2), 77–101; DOI: 10.1191/1478088706qp063oa) is the most-cited articulation of how to do thematic analysis rigorously. Their six-phase process — familiarisation, generating initial codes, searching for themes, reviewing themes, defining and naming themes, and producing the report — turns "I read the comments" into an auditable procedure. They stress that themes do not simply "emerge"; the analyst actively constructs them, so the choices must be made explicit. They also distinguish semantic coding (what is on the surface) from latent coding (underlying assumptions), and inductive from deductive (theory-driven) approaches — distinctions that matter when a QA office decides whether it is counting explicit mentions or interpreting sentiment.

Intercoder reliability — and the debate about it. Cliodhna O Connor and Helene Joffe's "Intercoder Reliability in Qualitative Research: Debates and Practical Guidelines" (International Journal of Qualitative Methods, 2020, 19, 1–13; DOI: 10.1177/1609406919899220) reviews when ICR is appropriate and how to do it. ICR asks: if two trained coders independently apply the same coding scheme to the same comments, how often do they agree? Reported as a coefficient — Cohen kappa for two coders, or Krippendorff alpha for more — it tests whether a theme is a property of the data or of one reader. O Connor and Joffe are even-handed: they note that ICR is contested in parts of the qualitative community (some argue meaning is interpretive and not meant to be "replicated"), but they conclude that, used appropriately, ICR improves the systematicity, communicability, and transparency of coding and helps convince sceptical audiences — exactly the audiences a QA office faces.

For coefficient interpretation, the standard references are Jacob Cohen's "A coefficient of agreement for nominal scales" (Educational and Psychological Measurement, 1960, 20(1), 37–46; DOI: 10.1177/001316446002000104) and Klaus Krippendorff's content-analysis methodology, which set out why raw percentage agreement overstates reliability (it ignores agreement expected by chance) and why a chance-corrected coefficient is preferred.

Why it matters for course evaluation in practice

When open-text findings feed real decisions, the coding method is not an academic nicety.

Accreditation panels increasingly want the qualitative story, not just the mean. ENQA/ESG-aligned reviews ask how the institution listens to students and what it changed (see our notes on ESG accreditation evidence and closing the feedback loop). If your "top three themes" cannot be defended as the product of a documented coding process, a panel can reasonably ask how you know they are not just the comments that caught one administrator's eye.

Selective reading is a real failure mode. A reviewer who scans 400 comments tends to over-weight vivid, emotional, or recent ones — a salience bias that systematically distorts which themes get reported. A coding frame applied to all comments, with a reliability check on a sample, is the discipline that prevents this.

Counts attached to themes need a defined unit. "65% of comments mentioned assessment" requires a rule for what counts as a mention and a coder who applies it consistently. Without inter-rater agreement, the percentage is an artefact of one person interpretation. This connects to the cautions in our note on text analytics and NLP: automation does not remove the need to validate that the categories mean what you say they mean.

Limitations and honest caveats

A methodologically literate reader will push back, and several objections are legitimate.

ICR is not universally appropriate. O Connor and Joffe themselves stress this. For deeply interpretive, latent-level analysis, forcing a kappa coefficient can be inappropriate or even misleading — high agreement on a shallow code is not the same as a valid deep interpretation. Reliability is a tool for certain kinds of coding (especially semantic, category-counting work), not a universal stamp of quality. Validity and reliability are different things.

Kappa has known quirks. The "kappa paradox" means that with highly skewed category prevalence (for example, when 95% of comments are positive), kappa can be low even when raw agreement is very high, making the coefficient hard to interpret. Coefficient choice and reporting of base rates matter.

Reliability is not validity. Two coders can agree perfectly on a coding frame that misses the point entirely. A reliable-but-wrong scheme is still wrong. Reliability checks must sit alongside, not replace, careful theme construction in the Braun and Clarke sense.

Resource cost. Double-coding even a sample of comments is labour-intensive, which is why many institutions skip it and report impressions instead. The honest trade-off is between rigour and effort — and worth naming, not hiding.

How Koji incorporates this

Koji for Education treats open-text analysis as a methodological task, not a vibe — and is designed to bring the discipline above into the routine evaluation cycle. These are mitigations and aids to human judgement, not claims of perfect objectivity.

  • Automatic thematic analysis applies one consistent frame to every response. Rather than an administrator skimming a subset, Koji codes the full corpus, which directly addresses the salience/selective-reading bias. Because the same procedure is applied to all comments, theme counts rest on a defined and repeatable rule rather than one reader attention.
  • AI-moderated conversational interviews generate cleaner text to code. When a student writes "the course was confusing", the moderator asks what specifically was confusing? in the moment. The resulting transcripts contain the concrete detail that makes coding reliable, instead of terse comments that two coders would interpret three different ways.
  • Transparent, reviewable theme outputs support human verification. Koji surfaces themes with their supporting quotes and counts, so a QA officer can audit whether a theme is genuinely grounded in the data — the practical equivalent of a reliability and validity check — and can adjust the framing where the latent meaning needs a human read.
  • Triangulation across cohorts and question types lets a theme be cross-checked against structured scale, ranking, and single_choice items rather than resting on free text alone, which guards against a coding artefact driving a decision.
  • Bias-aware reporting keeps prevalence visible (how many comments, how skewed), which is exactly the information needed to interpret any agreement statistic responsibly.

Koji does not claim its automatic coding is a substitute for human judgement on contested, latent themes — and the literature is clear that it should not be. The aim is to make the qualitative half of an evaluation as auditable and defensible as the quantitative half. Koji core research platform at koji.so applies the same thematic-analysis engine to customer and product research, where defensible coding of open-ended feedback is just as critical.

A practical checklist for defensible open-text coding

When open-text findings will be reported to a panel or used to justify a change, a small amount of methodological discipline makes the claims far harder to dismiss.

  • Write down the coding frame before you start. Define each theme, give inclusion and exclusion rules, and decide the unit of analysis (a comment, a sentence, a mention) so counts mean something specific.
  • Code the whole corpus, not a memorable sample. Selective reading over-weights vivid and recent comments; applying one frame to every response is the discipline that prevents salience bias.
  • Run a reliability check on a subset. Have a second coder independently code 10 to 20 percent of comments and compute a chance-corrected coefficient (Cohen kappa or Krippendorff alpha). Report base rates alongside it so a skewed prevalence does not mislead.
  • Keep reliability and validity separate. A high coefficient confirms consistency, not correctness; revisit whether the frame actually captures what matters using the Braun and Clarke review-and-define phases.
  • Show your evidence. Report themes with representative quotes and counts so a reader can audit whether a theme is grounded in the data or in the analyst expectations.

Done this way, "the main themes were assessment and feedback" stops being an impression and becomes a defensible, auditable finding — which is exactly what an ENQA-aligned reviewer is entitled to expect.

Related Resources

References

  • O Connor, C., & Joffe, H. (2020). Intercoder reliability in qualitative research: Debates and practical guidelines. International Journal of Qualitative Methods, 19, 1–13. https://doi.org/10.1177/1609406919899220
  • Braun, V., & Clarke, V. (2006). Using thematic analysis in psychology. Qualitative Research in Psychology, 3(2), 77–101. https://doi.org/10.1191/1478088706qp063oa
  • Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20(1), 37–46. https://doi.org/10.1177/001316446002000104
  • Krippendorff, K. (2018). Content Analysis: An Introduction to Its Methodology (4th ed.). SAGE Publications.

Related articles

accreditation

Turning Student Feedback into ESG / ENQA Accreditation Evidence

A buyer's guide mapping the ESG 2015 internal quality assurance standards to concrete, accreditation-ready evidence you can generate from student feedback — and how AI-moderated evaluation closes the loop.

analysis-reporting

Text Analytics for Open-Ended Student Comments: What NLP Can and Cannot Tell You

Natural-language processing can turn thousands of free-text course-evaluation comments into themes and sentiment at scale — but the research (Cunningham-Nelson 2019; Sunar & Khalid 2023) shows where automated analysis is reliable and where human judgement is still required.

analysis-reporting

Generalizability Theory and the Reliability of Student Ratings: Why One Class Is Not Enough

Reliability is not one number. Generalizability theory (Gillmore, Kane & Naccarato 1978; Marsh 1984) decomposes the variance in student ratings into student, teacher, course and occasion components — and shows that a single class can be reliable for the course yet a poor estimate of the teacher. What that means for fair evaluation.

analysis-reporting

Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees

The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.