New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
analysis-reporting10 min read

How Many Responses Do You Need for a Reliable Course Evaluation?

Reliability of a class-average rating rises steeply with the number of students who respond. Generalizability theory — from Marsh''s syntheses to Dzakadzie (2026) — shows raters matter more than items, and a class of 12 cannot be evaluated like a class of 120.

Koji Education Team

Product

Answer

The reliability of a course evaluation is governed far more by how many students respond than by how many questions you ask. Classic generalizability work synthesized by Marsh (1987) and recent generalizability-theory evidence from Dzakadzie (2026) converge on the same practical rule: a class-average rating becomes dependable somewhere around 15–25 respondents for low-stakes (formative) use, and you need still more — roughly the 25–60 range, depending on the instrument — before the average is solid enough for high-stakes (summative) decisions about an instructor or programme. Below about 10 respondents, a single class average is too noisy to stand alone and must be triangulated with other evidence. Adding more questions helps far less than hearing from more students.

What the research says

Generalizability (G) theory is the right lens here because it decomposes the variance in ratings into separate sources — the instructor, the individual student rater, the item, and their interactions — and then asks how dependable the class mean is for a given number of raters and items.

Marsh (1987), in his landmark monograph synthesizing data from roughly 5,000 classes, established that class-average student ratings are multidimensional, reliable, and stable, that they are primarily a function of the instructor rather than the course, and that their reliability climbs steeply with the number of students rating the class. The widely-reproduced generalizability estimates (originating with Gilmore, Kane and Naccarato, 1978, and summarized in Marsh''s reviews) show class-mean reliability rising from modest values for a handful of raters to roughly .90 and above once 25–50 students respond, and Marsh specifically cautions that when ratings are based on fewer than about 10 students, results from multiple classes are needed before drawing conclusions about an instructor.

Dzakadzie (2026), writing in Frontiers in Education, applied generalizability and decision (D) study analysis to modern evaluation instruments and reported two findings that update the classic picture. First, the number of raters matters far more than the number of items: increasing student raters improves dependability much more efficiently than lengthening the questionnaire. Second, the required sample depends on what the instrument measures and the class context — for a teaching-evaluation scale, roughly 12 items rated by about 20 students already produced high dependability (G ≈ .87), whereas a broader "course appraisal" needed more (in the order of 13 items and ~60 raters) and small classes (under ~25 students) remained difficult to evaluate dependably even with longer instruments. Dzakadzie adopts the standard benchmark that G ≈ 0.70 is acceptable for formative/low-stakes use, while 0.80+ is recommended for summative/high-stakes decisions — thresholds any QA office can adopt directly.

The literature is not unanimously rosy, which is the point of citing it honestly. Morley (2012) argued that most reliability claims rest on an ecological fallacy: they estimate reliability across groups of courses rather than within a single class, which inflates the apparent reliability of any one class result. Re-analyzing 1,073 course sections with intraclass and inter-rater agreement coefficients appropriate to the within-class question, Morley found that students were often unable to reliably evaluate a given instructor at the level of the individual section. The reconciliation is subtle but important: aggregate, multi-class evidence about a teacher can be reliable even when one isolated class mean — especially a small one — is not.

Why it matters for course evaluation in practice

1. Set a respondent-count gate, not a percentage gate. Because reliability is a function of the number of raters, a quality-assurance dashboard should flag results by how many students responded, not by response rate alone. A pragmatic policy: report formative results from ~10+ respondents with appropriate caveats, treat ~15–25 as reasonably dependable for formative use, and require materially more (and ideally corroborating evidence) before any summative judgment about a person.

2. Stop lengthening the questionnaire to "improve" reliability. Dzakadzie''s clearest managerial lesson is that adding items is an inefficient way to buy dependability. A shorter, well-constructed instrument that more students actually complete will be more reliable than a long one that depresses participation. This dovetails with response-rate research: the questionnaire and the turnout problem are linked.

3. Protect small classes from over-interpretation. Seminars, capstones, and specialist electives are exactly where a single class mean is least reliable and where high-stakes inferences are most tempting (small programmes, individual instructors). Morley''s critique is most biting here. The answer is triangulation: pool across the instructor''s sections and terms, and weight qualitative evidence more heavily where the quantitative average is thin.

4. Distinguish formative from summative thresholds. The same data can be perfectly adequate to give a lecturer mid-cycle feedback and wholly inadequate to inform a promotion case. Publishing one number without its reliability context invites misuse.

Limitations and honest caveats

  • G-coefficients are instrument- and population-specific. The numbers above (G ≈ .87 at 20 raters; .90+ at 25–50) come from particular scales and samples. They are sound order-of-magnitude guidance, not universal constants; your own instrument should ideally be subjected to a D-study.
  • The ecological-fallacy critique is real and unresolved. Marsh''s reassuring reliability and Morley''s skeptical reliability are answering subtly different questions (across-class vs within-class). Anyone quoting "student ratings are reliable" should be explicit about which reliability they mean.
  • Reliability is not validity. A perfectly reliable class mean can still be a biased or invalid measure of teaching quality (see the bias literature on gender, grading leniency, and the Dr. Fox/halo effects). Hearing from enough students removes noise; it does not remove systematic bias.
  • "Enough respondents" assumes they are representative. Sheer numbers do not cure non-response bias if the students who skip the survey differ systematically from those who complete it.

How Koji incorporates this

Koji for Education is designed so that reliability is something a reviewer can see and act on, not assume.

  • Respondent-count thresholds, surfaced. Koji reports each result with the number of respondents and the share of the cohort reached, and is built to apply formative-vs-summative confidence framing rather than printing a bare mean. This operationalizes the G ≈ 0.70 (formative) / 0.80+ (summative) logic that Dzakadzie and the broader literature recommend.
  • Depth per respondent, not just count. Because Marsh-style reliability rewards more raters and Dzakadzie shows items add little, Koji invests in getting more from each student rather than asking everyone more questions. Its AI-moderated conversational interview probes beyond a single Likert number with structured follow-ups (open_ended, scale, single_choice, multiple_choice, ranking, yes_no), so even a small cohort yields richer, codable evidence — partially offsetting the small-class reliability problem the research flags.
  • Triangulation by design. Where one section is under-powered, Koji supports pooling across an instructor''s sections and across mid-cycle and end-of-cycle collection — exactly the multi-class aggregation Marsh recommends for classes under ~10 and the antidote to Morley''s single-section unreliability.
  • Bias-aware, not just noise-aware. Koji''s thematic analysis and bias-aware reporting are designed to keep reviewers from confusing a reliable number with a valid one.

These mechanisms are designed to mitigate the small-sample and over-interpretation risks the research identifies; they do not manufacture reliability where too few students have responded. Koji''s core research platform at koji.so applies the same reliability-aware, AI-moderated engine to product and customer research, where small-sample over-reading is an equally common failure mode.

A worked example: reading three class results

Consider three results that all show an identical 4.2/5 mean — and why a reliability-aware office would treat them completely differently.

  • Class A: 220 enrolled, 96 respondents. With nearly a hundred raters, the class mean sits comfortably in the high-dependability range (well above G ≈ 0.80). A 4.2 here is solid enough to inform a summative judgement, especially when corroborated across the instructor''s other sections.
  • Class B: 40 enrolled, 17 respondents. Seventeen raters is roughly at the formative threshold (around G ≈ 0.70 on a teaching scale, per Dzakadzie''s D-study findings). The 4.2 is reasonable direction for a mid-cycle conversation but too fragile to anchor a personnel decision on its own.
  • Class C: 11 enrolled, 6 respondents. Six voices is below the level at which a single class mean is dependable. Marsh''s guidance is explicit: pool across multiple classes before concluding anything about the instructor, and weight the qualitative comments — read individually — more heavily than the average.

The lesson is that the same number carries three different evidential weights, governed almost entirely by how many students responded. A dashboard that prints "4.2" three times without the respondent count actively misleads its readers. Building the denominator and a formative/summative confidence band into every report is the cheapest, highest-impact reliability intervention a quality-assurance office can make — and it requires no change to the instrument itself.

Related Resources

References

  • Marsh, H. W. (1987). Students'' evaluations of university teaching: Research findings, methodological issues, and directions for future research. International Journal of Educational Research, 11(3), 253–388. https://doi.org/10.1016/0883-0355(87)90001-2
  • Dzakadzie, Y. (2026). Identifying optimal number of raters and items for dependable student evaluation of teaching: evidence from generalizability theory. Frontiers in Education, 11, 1814494. https://doi.org/10.3389/feduc.2026.1814494
  • Morley, D. D. (2012). Claims about the reliability of student evaluations of instruction: The ecological fallacy rides again. Studies in Educational Evaluation, 38(1), 15–20. https://doi.org/10.1016/j.stueduc.2012.01.001
  • Gilmore, G. M., Kane, M. T., & Naccarato, R. W. (1978). The generalizability of student ratings of instruction: Estimation of teacher and course components. Journal of Educational Measurement, 15(1), 1–13. https://doi.org/10.1111/j.1745-3984.1978.tb00051.x