New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
analysis-reporting9 min read

High Agreement, Low Kappa: Choosing an Agreement Coefficient for Coding Course Feedback

Cohen's kappa can collapse to near zero even when two coders agree on 95 percent of comments. Here is why the kappa paradox happens and when to report Gwet's AC1 or Krippendorff's alpha instead.

Koji Education Team

Product

In brief

When two people code the same open-text course comments, you report a chance-corrected agreement coefficient so readers know the coding was not just one analyst's opinion. The default choice, Cohen's kappa, has a notorious failure mode: when one category is rare — say only a handful of comments mention "assessment fairness" — kappa can crash toward zero or go negative even though the coders agreed on 95 percent of comments. This is the kappa paradox, and it is a property of the coefficient, not of your coding. The fix is to know when it happens and to report a paradox-resistant coefficient — Gwet's AC1 or, for many coders, missing data or ordinal codes, Krippendorff's alpha — alongside the raw percentage agreement.

What the research says

Chance-corrected agreement all follows one shape: (observed agreement − expected-by-chance agreement) / (1 − expected-by-chance agreement). The coefficients differ only in how they estimate the chance term.

Cohen (1960), "A coefficient of agreement for nominal scales" (Educational and Psychological Measurement, 20(1), 37-46, doi:10.1177/001316446002000104), estimated chance agreement from the product of each coder's marginal rates. That is the source of the trouble. When a category is very common (or very rare), the marginals are skewed, the expected-chance term inflates, and subtracting a large chance term from a high observed agreement leaves a small — or negative — kappa.

Feinstein and Cicchetti (1990), "High agreement but low kappa: I. The problems of two paradoxes" (Journal of Clinical Epidemiology, 43(6), 543-549, doi:10.1016/0895-4356(90)90158-L), named and formalised this. Their first paradox: with high observed agreement, unbalanced marginals push kappa low. Their second: symmetric versus asymmetric imbalance in the disagreements can make kappa swing in unintuitive directions. Two coding exercises with identical 90 percent agreement can post kappas of 0.8 and 0.0 purely because of how prevalent the categories are. For course feedback — where "the lecturer was clear" recurs constantly and "the room was too cold" appears twice — skewed prevalence is the norm, not the exception.

Gwet (2008), "Computing inter-rater reliability and its variance in the presence of high agreement" (British Journal of Mathematical and Statistical Psychology, 61(1), 29-48, doi:10.1348/000711006X126600), proposed the AC1 coefficient, which estimates chance agreement differently — as the probability that raters agree by chance given the propensity to make random rather than deterministic judgements. AC1 is provably far less sensitive to prevalence, so it does not collapse when a category is rare. In the teacher-evaluation coding study by Wongpakaran and colleagues and later replications, AC1 stayed stable and interpretable exactly where kappa produced paradoxical near-zero values on data with 90 percent-plus agreement.

Krippendorff's alpha (Krippendorff, 2004, "Reliability in content analysis: Some common misconceptions and recommendations," Human Communication Research, 30(3), 411-433, doi:10.1111/j.1468-2958.2004.tb00738.x) is the most general option: it handles any number of coders, missing data, and nominal, ordinal, interval or ratio codes through a configurable distance function. It derives its chance baseline from the pooled observed distribution, so it shares kappa's sensitivity to skewed marginals to some degree, but its flexibility (multiple coders, partial data, weighted disagreements) makes it the reference standard in content analysis. Alpha of at least 0.80 is the customary threshold for firm conclusions, with 0.667 as the floor for tentative ones.

The honest synthesis in the recent methods literature — including Zec, Soriani and colleagues, "Gwet's AC1 is not a substitute for Cohen's kappa" (2023) — is that no single coefficient is universally correct. Each encodes a different assumption about what "chance agreement" means. The professional move is to report percentage agreement plus the coefficient whose assumptions fit your data, and to say which one and why.

Why it matters for course evaluation in practice

Open-text comments are where the actionable detail lives, and coding them reliably is what separates defensible thematic analysis from cherry-picking. If a quality office reports "inter-rater kappa was 0.31," a committee may wrongly conclude the coding was unreliable and discard the qualitative evidence — when in truth the coders agreed almost perfectly and kappa was merely paradoxical because most comments fell in one or two themes.

Practical guidance:

  • Always report raw percentage agreement next to any chance-corrected coefficient, so a paradox is visible.
  • Report prevalence. If the modal category exceeds roughly 85-90 percent of cases, expect kappa to be depressed and prefer AC1.
  • Match the coefficient to the design. Two coders, nominal themes, skewed prevalence → AC1. Many coders, missing data, or ordinal severity codes → Krippendorff's alpha. Two coders, roughly balanced categories → Cohen's kappa is fine.
  • Do not "shop" for the highest number. Choose the coefficient before seeing which flatters you, and disclose the choice.

This is the coding-reliability analogue of the method-agreement problem that Bland-Altman limits address for continuous measures: a single summary statistic can hide, or invent, disagreement.

Limitations and honest caveats

  • AC1 is not a free lunch. Because it assumes a portion of agreement is deterministic, AC1 can overstate reliability when true chance agreement really is high. Zec and colleagues argue it should not be a reflexive replacement for kappa; it answers a slightly different question.
  • Prevalence is real information, not only nuisance. A skewed category distribution sometimes reflects a genuinely low-frequency but important theme. Correcting it away statistically should not lead you to ignore the substantive rarity.
  • Alpha and kappa share a marginal-dependence weakness. Krippendorff's alpha is more flexible but not fully paradox-immune under extreme skew.
  • Coefficients do not validate a codebook. High agreement on a vague or biased coding scheme is high agreement on the wrong thing. Reliability is necessary, not sufficient — pair it with thematic saturation and a defensible framework.
  • Small comment sets give unstable estimates for every coefficient; report confidence intervals, which Gwet's and Krippendorff's frameworks both provide.

How Koji incorporates this

Koji automates the first pass of open-text coding and treats the human-versus-model comparison as an explicit agreement problem rather than a black box.

  • Transparent agreement reporting: when Koji's automated thematic analysis is checked against human coders, the comparison is framed as inter-rater agreement, with raw agreement and a prevalence-aware coefficient — so a rare-but-important theme does not produce a misleadingly low headline number. This mirrors the human-versus-LLM coding evidence.
  • Prevalence surfaced, not hidden: Koji dashboards show theme frequencies, so an analyst can see immediately when the modal category dominates and interpret any agreement statistic accordingly.
  • Consistent, auditable codebook: applying the same structured scheme across every comment reduces the coder-drift that erodes agreement in the first place, and keeps the coding auditable for accreditation review.
  • Confidence intervals, not point estimates: because class-level comment sets are small, Koji favours interval reporting over a single seductive number. Koji's core research platform at koji.so uses the same agreement machinery when reconciling AI and human coding of customer-research transcripts.

Koji is designed to mitigate the paradox-driven misreading of qualitative reliability, not to claim any one coefficient is always right — the choice still has to be justified for your data.

Frequently asked questions

Why does Cohen's kappa go down when my coders clearly agree?

Because kappa subtracts an expected-chance-agreement term estimated from the coders' marginal rates. When one category is very common or very rare, that term inflates, so a high observed agreement minus a high chance term leaves a small or negative kappa. This is the kappa paradox described by Feinstein and Cicchetti (1990).

When should I use Gwet's AC1 instead of kappa?

Use AC1 when agreement is high and category prevalence is skewed — the usual situation for course-feedback themes. AC1 estimates chance agreement in a way that is far less sensitive to prevalence, so it does not collapse the way kappa does.

What makes Krippendorff's alpha different?

Alpha handles any number of coders, missing data, and nominal, ordinal or interval codes through a distance function, which makes it the content-analysis standard. It still draws its chance baseline from the observed distribution, so it is not fully immune to skew, but it is the most flexible option.

What counts as acceptable agreement?

For Krippendorff's alpha, at least 0.80 is the customary threshold for firm conclusions and 0.667 the floor for tentative ones. Whatever coefficient you use, report raw percentage agreement and a confidence interval alongside it.

Can I just report whichever coefficient looks best?

No. Choose the coefficient based on your design and prevalence before you see which flatters the result, and disclose the choice. Shopping for the highest number is a reporting error.

Does high agreement mean the coding is valid?

No. Agreement is reliability, not validity. Two coders can agree perfectly on a vague or biased codebook. Pair the coefficient with codebook development, saturation checks, and a transparent framework.

Related resources

References

  • Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20(1), 37-46. doi:10.1177/001316446002000104
  • Feinstein, A. R., & Cicchetti, D. V. (1990). High agreement but low kappa: I. The problems of two paradoxes. Journal of Clinical Epidemiology, 43(6), 543-549. doi:10.1016/0895-4356(90)90158-L
  • Gwet, K. L. (2008). Computing inter-rater reliability and its variance in the presence of high agreement. British Journal of Mathematical and Statistical Psychology, 61(1), 29-48. doi:10.1348/000711006X126600
  • Krippendorff, K. (2004). Reliability in content analysis: Some common misconceptions and recommendations. Human Communication Research, 30(3), 411-433. doi:10.1111/j.1468-2958.2004.tb00738.x
  • Zec, S., Soriani, N., Comoretto, R., & Baldi, I. (2023). Gwet's AC1 is not a substitute for Cohen's kappa - A comparison of basic properties. MethodsX, 10, 102212. doi:10.1016/j.mex.2023.102212
  • Wongpakaran, N., Wongpakaran, T., Wedding, D., & Gwet, K. L. (2013). A comparison of Cohen's kappa and Gwet's AC1 when calculating inter-rater reliability coefficients: A study conducted with personality disorder samples. BMC Medical Research Methodology, 13, 61. doi:10.1186/1471-2288-13-61