New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
analysis-reporting10 min read

Can You Report a Class Mean? ICC(1), ICC(2), and r_wg for Aggregating Student Ratings

Before you average student ratings into a class score, three organisational-psychology indices decide whether you can: r_wg (within-class agreement), ICC(1) (how much variance is between classes), and ICC(2) (the reliability of the class mean).

Koji Education Team

Product

Course-evaluation reporting almost always ends with a single number per class: the mean of the students who responded. Attach a decimal, colour it red or green, put it on a dashboard, and a career decision follows. But averaging turns many individual opinions into one group-level statistic — and that move is only defensible if two conditions hold: students within a class agree enough with each other for a mean to represent them, and the class means differ enough from one another for a comparison to carry signal. The framework that answers both questions comes not from education research but from organisational psychology, where the problem of "when can I aggregate individual ratings to a group score" has been worked out in detail. Its three core indices are ICC(1), ICC(2), and r_wg.

The short answer (BLUF)

Before you report or compare class means, check three things. r_wg tells you whether the students in a given class agree enough to be summarised by one number. ICC(1) tells you how much of the total variation in ratings is between classes rather than within them — that is, how much a single student's response is "about the class" versus "about the student". ICC(2) tells you how reliably the class means can be distinguished from one another given your response counts. Low ICC(2) — common when response rates are low — means the ranking of class means is mostly noise, and you should not act on small differences no matter how precise the decimal looks. These indices are standard, computable from data you already hold, and directly limit what your evaluation dashboard is allowed to claim.

What the research says

The modern treatment is Bliese's chapter on within-group agreement, non-independence, and reliability, and LeBreton and Senter's question-and-answer synthesis, which together formalise the aggregation decision. LeBreton and Senter (2008) distinguish interrater agreement (do raters give the same absolute value?) from interrater reliability (are raters rank-ordering targets consistently?), and show that these are different questions requiring different indices — a distinction routinely muddled in evaluation practice.

r_wg (James, Demaree, and Wolf, 1984) measures within-group agreement for a single target by comparing the observed variance of ratings within a class against the variance you would expect if students answered at random. When observed variance is far smaller than random-response variance, r_wg approaches 1 (strong agreement); when ratings are as spread out as random, r_wg approaches 0. Crucially, r_wg requires you to specify a null distribution — usually a uniform distribution across scale points — and the result is sensitive to that choice, a point James and colleagues were explicit about.

ICC(1) and ICC(2) come from a one-way random-effects analysis of variance in which students are nested in classes. ICC(1) is the proportion of total rating variance attributable to class membership; Bliese (2000) notes it can also be read as the expected correlation between two randomly chosen students in the same class, and that values around 0.05–0.20 are typical for group-level perceptions. ICC(2), by contrast, is the reliability of the class mean, and — through the Spearman-Brown logic — depends on both ICC(1) and the number of respondents per class: even a modest ICC(1) yields a trustworthy mean if enough students respond, while a high ICC(1) can still produce an unreliable mean if only three students answered. Koo and Li (2016), writing for a different applied field, give the widely used interpretive bands: below 0.50 poor, 0.50–0.75 moderate, 0.75–0.90 good, above 0.90 excellent reliability — bands that map naturally onto ICC(2) for class means.

The relevance to teaching evaluation is direct: Marsh's programme of research on students' evaluations of teaching repeatedly framed reliability in terms of the number of students per class, showing that class-average ratings become stable only once responses accumulate — the same relationship ICC(2) formalises.

Why it matters for course evaluation in practice

Three practical consequences follow.

First, a class mean is not automatically meaningful. If r_wg is low, students in that class genuinely disagree — some loved the course, some hated it — and the mean hides a bimodal reality that a location-scale or distributional view would expose. Reporting only the average here is not neutral; it erases the disagreement that is the actual finding.

Second, small classes produce unreliable means even with high agreement. Because ICC(2) scales with respondent count, a department that compares a 5-respondent seminar against a 120-respondent lecture on the same decimal scale is comparing a noisy estimate against a precise one as if they were equivalent. The fix is not to hide small classes but to report the reliability of each mean and widen the interval around sparse ones.

Third, if ICC(1) is near zero, between-class comparison is close to meaningless — almost all the variance is within classes, so ranking instructors on their means is ranking on sampling noise. This is the empirical test of whether your evaluation instrument discriminates between teaching contexts at all.

These indices also tell you how many responses you need. Setting a target ICC(2) (say, 0.70) and estimating ICC(1) from historical data lets you back out a minimum response count per class — a principled alternative to arbitrary "50% response rate" rules.

Limitations and honest caveats

The indices are not a validity guarantee. High agreement (r_wg near 1) can reflect a genuine shared experience or a shared bias — every student anchoring on the same halo, the same grading expectation, or the same salient incident. Agreement is necessary for a defensible mean but says nothing about whether the mean measures teaching quality rather than, say, leniency.

r_wg is sensitive to the assumed null distribution; a uniform null is conventional but debatable, and a skewed null (students rarely use the bottom of the scale) can push r_wg down. It can also fall outside [0,1] and require truncation, which some methodologists dislike. ICC(1) and ICC(2) assume the one-way random-effects model is appropriate; when students are cross-classified across several teachers, or when non-response is informative, the simple ICC understates the uncertainty and a multilevel or cross-classified model is more honest. Finally, all three indices describe this administration; they do not license extrapolation to a different cohort, term, or delivery mode.

How Koji incorporates this

Koji is built to make the aggregation decision explicit rather than hiding it inside an average. Because Koji collects structured responses (scale, single_choice, multiple_choice, yes_no, ranking) alongside open text, it has the item-level data needed to compute within-class agreement and between-class reliability, so a class mean can be reported with an agreement flag and a reliability-aware interval rather than as a bare decimal. Where within-class agreement is low, Koji's bias-aware reporting is designed to surface the split — showing the distribution and the competing narratives from its automatic thematic analysis of open text — instead of collapsing a divided class into one misleading number. For sparse classes, Koji's mid-cycle and formative collection is designed to accumulate enough responses to lift the reliability of the mean before it is used for any consequential decision, and its triangulation across cohorts lets a programme pool evidence rather than over-read a single small section. The AI-moderated conversational interview adds a further layer the ICC framework cannot: when a student rates the course a 3, Koji probes why, so a low-agreement class can be understood rather than merely detected. Koji's core research platform at koji.so applies the same aggregation-aware reporting to product and customer research, where the "can I trust this segment average" question is identical.

None of this eliminates the underlying limits — Koji is designed to mitigate over-interpretation of unreliable means, not to manufacture reliability that the data does not contain.

Frequently asked questions

What is the difference between ICC(1) and ICC(2) in course evaluation?

ICC(1) is the proportion of total variance in individual student ratings that lies between classes rather than within them — roughly, how much one student's answer reflects the class versus the student. ICC(2) is the reliability of the class mean and depends on both ICC(1) and how many students responded. You can have a low ICC(1) but a reliable mean (ICC(2)) if enough students answer, and a respectable ICC(1) with an unreliable mean if only a handful respond.

What is a good ICC(2) value for reporting a class average?

Applying the Koo and Li (2016) bands, ICC(2) below 0.50 indicates poor reliability, 0.50–0.75 moderate, 0.75–0.90 good, and above 0.90 excellent. Many institutions treat roughly 0.70 as a working floor for using a class mean in consequential decisions, and report sparser classes with wider intervals or not at all.

What does r_wg tell me that the mean does not?

r_wg measures whether students within a single class actually agree with one another. Two classes can share the same 3.8 average while one is a genuine consensus and the other is a 50/50 split between fives and ones. r_wg distinguishes them; the mean cannot. Low r_wg is a signal to report the distribution, not the average.

Can a high response rate fix a low ICC(1)?

No. A high response rate raises ICC(2) — the reliability of the mean — but ICC(1) is a property of how much classes differ. If ICC(1) is near zero, the classes barely differ, so even perfectly reliable means carry little comparative signal; collecting more responses just measures a near-nonexistent difference more precisely.

Why not just report every class mean to one decimal place?

Because a decimal implies a precision the data may not support. A mean from five respondents in a class where students disagree is a noisy estimate dressed up as an exact figure. Reporting agreement (r_wg) and mean reliability (ICC(2)) alongside the number keeps the report honest about what it can and cannot claim.

Which index should trigger a manual review rather than an automatic flag?

Low ICC(2) (an unreliable mean) argues against any automatic flag — the ranking is too noisy to act on. Low r_wg with an adequate response count argues for a manual, qualitative review, because it signals a genuinely divided class whose story the average is hiding.

Related resources

References

  • James, L. R., Demaree, R. G., & Wolf, G. (1984). Estimating within-group interrater reliability with and without response bias. Journal of Applied Psychology, 69(1), 85–98. https://doi.org/10.1037/0021-9010.69.1.85
  • Bliese, P. D. (2000). Within-group agreement, non-independence, and reliability: Implications for data aggregation and analysis. In K. J. Klein & S. W. J. Kozlowski (Eds.), Multilevel Theory, Research, and Methods in Organizations (pp. 349–381). Jossey-Bass.
  • LeBreton, J. M., & Senter, J. L. (2008). Answers to 20 questions about interrater reliability and interrater agreement. Organizational Research Methods, 11(4), 815–852. https://doi.org/10.1177/1094428106296642
  • Koo, T. K., & Li, M. Y. (2016). A guideline of selecting and reporting intraclass correlation coefficients for reliability research. Journal of Chiropractic Medicine, 15(2), 155–163. https://doi.org/10.1016/j.jcm.2016.02.012
  • Marsh, H. W. (1987). Students' evaluations of university teaching: Research findings, methodological issues, and directions for future research. International Journal of Educational Research, 11(3), 253–388. https://doi.org/10.1016/0883-0355(87)90001-2