New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods8 min read

How Many Open-Text Comments Are Enough? Thematic Saturation in Course Evaluations

The qualitative-research evidence on data saturation — Guest et al. (2006) and Hennink et al. (2017) — applied to course-evaluation free text: how many comments you need before new themes stop appearing, and why richness per comment matters more than raw count.

Koji Education Team

Product

In brief

The qualitative-methods literature gives a usable rule of thumb that almost no course-evaluation report applies: the number of open-text comments needed to capture the themes in student feedback is far smaller than most people assume — landmark experiments find the common, high-prevalence themes emerge within roughly the first 6–12 responses, with broad code saturation by around 9. But a fuller, nuanced understanding of those themes (meaning saturation) takes more — on the order of 16–24 substantive comments — and rare concerns may never saturate at all. For course evaluation this reframes two questions at once: how many comments you must read before the thematic picture stabilises, and why a few rich, probed responses can be worth more than hundreds of thin ones.

What the research says

The foundational study is Guest, Bunce, and Johnson (2006), published in Field Methods. Working with 60 in-depth interviews from a reproductive-health study across two West African countries, they coded incrementally and tracked when new themes stopped appearing. Their headline result: saturation — the point at which additional data ceased to yield new codes — occurred within the first twelve interviews, and the basic elements of the major "metathemes" were already present after about six. High-prevalence themes, the ones that recur across many participants, were identifiable very early; it was the low-prevalence, idiosyncratic content that kept trickling in.

Hennink, Kaiser, and Marconi (2017), in Qualitative Health Research, sharpened the concept by separating two kinds of saturation. Code saturation — hearing the range of issues, "heard it all" — was reached at about nine interviews in their data. But meaning saturation — developing a richly textured understanding of each issue, "understand it all" — required 16 to 24 interviews. The distinction is decisive for evaluation work: counting whether a theme appears is cheap and saturates fast; understanding what students actually mean by, say, "the assessment was unfair" saturates much more slowly.

Replication and extension work qualifies both. Hagaman and Wutich (2017), revisiting Guest's data across multiple sites in Field Methods, showed that when you need themes that hold across subgroups (for example, across faculties or campuses), 20–40 interviews may be required before cross-site metathemes stabilise. The unifying lesson is that saturation is not a fixed number; it is a function of theme prevalence, the heterogeneity of your population, and how richly each respondent speaks.

Why it matters for course evaluation in practice

Two practical consequences follow, and they pull in opposite directions from the usual instincts.

First, you need fewer comments than you think to know what the issues are. Programme directors often dismiss qualitative feedback because "only 30 students wrote anything." But 30 substantive comments are, by the saturation evidence, comfortably past code saturation for the common themes. The major recurring concerns in a course — assessment load, pacing, feedback turnaround, lab access — will almost always be visible in the first dozen or two thoughtful responses. A low comment count is not a reason to ignore the text; it is usually enough to identify the dominant themes reliably. This directly complements the quantitative question of how many responses you need for a reliable course evaluation, where the logic for Likert means is different and typically demands more.

Second, you need richer comments than you usually get to know what students mean. Meaning saturation is where most course-evaluation processes fail. A one-line free-text box produces thin codes — you learn that "assessment" is a theme but not whether the problem is volume, clarity, weighting, or timing. Reaching meaning saturation on a thin-comment corpus may be impossible at any sample size, because the depth simply is not in the data. This is why response quality per comment, not just response rate, governs how much you can actually learn.

Third, rare but serious themes never saturate — and must not be averaged away. A single credible report of discrimination, a safeguarding concern, or a hazardous lab condition is signal, not noise, even though it appears once and "fails" to saturate. Saturation logic tells you when the typical picture is stable; it says nothing about whether a rare comment matters. Treating low-frequency comments as below a reporting threshold is a methodological error, and it intersects with the duty-of-care issues we cover in abusive open-text comments and faculty wellbeing.

A worked example

Consider a second-year statistics module with 95 enrolled students, of whom 28 leave a written comment. By the saturation evidence, those 28 comments are well past code saturation: the recurring themes — pace too fast, worked examples too few, the group project weighting disputed — will all appear, probably within the first dozen. A programme director can therefore trust that the list of dominant issues is stable and act on it. What 28 thin comments will not deliver is meaning saturation on a contested theme: "the marking felt unfair" might mean the rubric was opaque, the weighting was wrong, or the feedback came too late, and a one-line comment cannot disambiguate. The correct response is not to collect more thin comments — that adds little once code saturation is reached — but to probe the existing respondents for depth, or to run a short follow-up that asks the students who flagged fairness to say more. This is the practical difference between needing more people and needing more depth per person.

Limitations and honest caveats

Several objections deserve airing for a critical reader.

Saturation is partly an artefact of the coding frame. You only saturate relative to the codebook you are building. A different analyst, or a more granular scheme, will saturate later. Saturation is therefore not a purely objective stopping rule; it is contingent on the interpretive lens, a point critics of the concept (including later reflections by Braun and Clarke) have pressed hard.

The anchor studies are interviews, not survey free-text. Guest and Hennink studied in-depth interviews — long, probed, interactive. A typical course-evaluation comment is a sentence or two written in haste. You cannot assume one written comment equals one interview; it carries far less information, so the effective number of "comments" needed for code saturation is plausibly higher than the interview figures suggest. The numbers transfer as logic, not as constants.

Heterogeneity breaks the small numbers. The 6–12 figures hold for relatively homogeneous populations. A faculty-wide evaluation spanning disciplines, languages, and cohorts is heterogeneous; Hagaman and Wutich's 20–40 range is the more honest reference for cross-group conclusions.

Saturation says nothing about representativeness. Reaching thematic saturation among the students who chose to comment does not tell you about the students who stayed silent. Non-response and self-selection biases still apply to the qualitative stream, just as they do to the quantitative one.

How Koji incorporates this

Koji is built to attack the meaning-saturation problem at its source, which the evidence identifies as the real bottleneck.

The platform replaces the thin, one-shot free-text box with an AI-moderated conversational interview: when a student raises an issue, the system asks a relevant follow-up — "you mentioned the assessment felt unfair; was that about the workload, the marking criteria, or the timing?" Each respondent therefore contributes interview-like depth rather than a single thin line, which is precisely the input the saturation studies were built on. In saturation terms, Koji is designed to push individual responses from code-level toward meaning-level richness, so meaning saturation can plausibly be reached with fewer respondents.

Koji's automatic thematic analysis then tracks the emergence of themes across responses and presents the distribution with verbatim evidence attached, so a quality officer can see when the common themes have stabilised and which comments are rare-but-serious outliers that warrant individual attention rather than aggregation. Because the platform supports mid-cycle and formative collection, an institution can run a short conversational pulse, reach code saturation on the dominant issues quickly, and act before the semester ends rather than waiting for a large end-of-term corpus.

None of this eliminates the limitations above — Koji frames saturation as a guide to interpretation, not a licence to stop reading, and it never suppresses low-frequency comments. The same conversational engine powers Koji's core research platform at koji.so, where product and customer-research teams rely on the same depth-per-respondent advantage to reach saturation efficiently.

Related resources

References

  • Guest, G., Bunce, A., & Johnson, L. (2006). How Many Interviews Are Enough? An Experiment with Data Saturation and Variability. Field Methods, 18(1), 59–82. https://doi.org/10.1177/1525822X05279903
  • Hennink, M. M., Kaiser, B. N., & Marconi, V. C. (2017). Code Saturation Versus Meaning Saturation: How Many Interviews Are Enough? Qualitative Health Research, 27(4), 591–608. https://doi.org/10.1177/1049732316665344
  • Hagaman, A. K., & Wutich, A. (2017). How Many Interviews Are Enough to Identify Metathemes in Multisited and Cross-Cultural Research? Another Perspective on Guest, Bunce, and Johnson's (2006) Landmark Study. Field Methods, 29(1), 23–41. https://doi.org/10.1177/1525822X16640447