New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Methodology9 min read

Would Two Analysts Agree? Inter-Coder Reliability and the Standard Your Open-Text Feedback Never Meets

The debate about whether AI can be trusted to code student comments skips a prior question: could two humans doing the same job even agree? Inter-coder reliability is the standard almost no open-text analysis reports — for people or machines.

Koji Education Team

Product ·

Answer up front: When you turn hundreds of free-text student comments into themes — "assessment was unfair", "loved the labs", "pace too fast" — you are performing a coding task, and coding has a measurable quality standard: inter-coder reliability, the degree to which independent analysts assign the same comment to the same category. The convention, from Klaus Krippendorff's work on content analysis, is that his agreement coefficient alpha (α) should reach 0.80 for firm conclusions, with 0.667 to 0.80 tolerated only for tentative ones. The uncomfortable truth is that most course-evaluation open-text "analysis" — whether done by a tired administrator over a weekend or by an AI in seconds — never reports any reliability figure at all. Before you ask whether AI can be trusted to code comments, ask the prior question the sector has skipped for decades: could two humans doing this job even agree?

The question everyone skips

Course evaluations produce two kinds of data. The numbers get a great deal of statistical scrutiny — we have argued at length about why averaging Likert scores misleads. The free text gets almost none. A programme lead reads the comments, forms an impression, and reports "students mainly praised the seminars but raised concerns about feedback turnaround." That impression is a coding decision — an assignment of raw text to categories — but it is a coding decision by a single coder, unaudited, with no check on whether anyone else reading the same comments would have drawn the same map.

This matters because qualitative coding is genuinely hard and genuinely subjective. Does "the tutor was always rushing" count as a pace complaint or a teacher availability complaint? Reasonable analysts disagree. Reliability coefficients exist precisely to quantify how much they disagree, so that a theme is not just one person's reading dressed up as a finding.

What inter-coder reliability actually measures

The idea is simple: have two or more coders independently categorise the same set of comments, then measure how often they agree — corrected for the agreement you would expect by chance alone. Raw percentage agreement is misleading, because with only a few categories coders agree often by luck. Chance-corrected coefficients fix this.

  • Cohen's kappa (κ), introduced by Jacob Cohen in 1960, is the classic two-coder statistic (Cohen, 1960, Educational and Psychological Measurement). It is widely used but has known quirks — the "kappa paradox", where high agreement can produce a low kappa when one category dominates.
  • Krippendorff's alpha (α) is the more general tool: it handles any number of coders, missing data, and any level of measurement (nominal, ordinal, interval). Krippendorff's guidance is that it is "customary to require α ≥ .800", and to treat the band between .667 and .800 as fit only for drawing tentative conclusions (Krippendorff, Content Analysis: An Introduction to Its Methodology, summarised here).

The threshold is not sacred — for genuinely complex, interpretive coding some methodologists accept .667 — but the practice of reporting it is the point. A theme derived from coding with an unknown reliability is a theme of unknown trustworthiness.

The double standard hiding in the AI debate

Here is where the argument turns. When institutions consider using AI to summarise open-text feedback, the reflex objection is: can we trust a machine to code comments correctly? It is a fair question — we have written about whether AI can reliably analyse thousands of student comments and why a sentiment score is not insight. But notice the double standard. The human process the AI would replace was almost never held to a reliability standard either. Nobody asked whether the administrator's weekend read-through would replicate if a colleague did it. We demand of the machine a rigour we never demanded of ourselves.

The correct response is to apply the same standard to both. Emerging benchmark studies now do exactly this — comparing large-language-model coding of qualitative data against human expert adjudication and reporting agreement coefficients for each (recent benchmark work, 2026). This is the right frame: reliability is a property of a coding procedure, human or machine, and the honest question is whether a given procedure — whoever runs it — hits a defensible α, and whether it does so consistently across the messy, ambiguous comments that matter most.

How to actually check it

You do not need a research grant to audit your own open-text analysis:

  1. Draw a sample of, say, 100–200 comments from a real evaluation cycle.
  2. Fix a codebook — the list of themes and clear inclusion rules — before coding, so coders are aiming at the same target.
  3. Have two independent coders (or your AI plus a human) code the sample without conferring.
  4. Compute Krippendorff's alpha on the results. Free calculators exist for exactly this.
  5. Read the disagreements. Where coders diverge tells you which themes are ill-defined — often more useful than the coefficient itself.

If α clears .80, your theming is dependable enough to act on. If it sits at .5, the "themes" in your annual report are substantially the coder's construction, not the students' collective voice.

But isn't reliability the wrong goal for qualitative data?

This is the strongest counterargument, and it deserves a fair hearing. A serious strand of qualitative methodology holds that forcing rich, interpretive text into a fixed codebook and chasing an agreement coefficient betrays the whole point of qualitative inquiry — that meaning is co-constructed, that a skilled single analyst's deep reading can be more valid than two shallow readings that happen to agree, and that reliability is a positivist import ill-suited to interpretive work. There is truth here. Inter-coder reliability is the right standard when your goal is to count and compare themes — "how many students raised feedback turnaround, and is it worse than last year?" — which is exactly what institutional course evaluation does at scale. It is the wrong standard for a small, deeply interpretive study whose value is depth, not enumeration. The failure mode in course evaluation is not excessive positivism; it is the opposite — enumerative claims ("students mainly complained about X") presented with none of the reliability evidence that enumeration requires. Use the standard where you are counting; drop it where you are genuinely interpreting. Just be honest about which one you are doing.

Where Koji fits

Koji for Education treats open-text analysis as a coding procedure to be trusted only if it earns it. Its automatic thematic analysis applies a consistent, standardised categorisation across every comment — no coder fatigue, no drift between the first fifty comments and the last five hundred — which addresses one of the main threats to human reliability. Crucially, every theme is anchored to the underlying student quotes rather than asserted, so a sceptical reader can audit the coding directly; we consider this non-negotiable, and it is the argument of why you should not trust a theme the AI cannot show you the quote for. Because the same analysis runs identically each cycle, results are comparable over time in a way a rotating cast of human coders can never guarantee. Koji does not claim its coding is perfect; it claims something more defensible — that its procedure is consistent, auditable, and testable against the very reliability standard the sector has spent decades ignoring. (The same thematic-analysis engine powers the wider research platform at koji.so for teams coding customer and user interviews.)

The lesson is not "trust the AI" or "distrust the AI". It is: hold every coding procedure — yours, your colleague's, the machine's — to the standard that two competent analysts should broadly agree. Most course-evaluation open text has never once been held to it.

Closing the loop

If your open-text feedback deserves an analysis you can defend to a sceptical committee, see how Koji for Education themes student comments with quote-level provenance. Reliability is not a machine question. It is a procedure question — and it applies to all of us.