Generalizability Theory and the Reliability of Student Ratings: Why One Class Is Not Enough
Reliability is not one number. Generalizability theory (Gillmore, Kane & Naccarato 1978; Marsh 1984) decomposes the variance in student ratings into student, teacher, course and occasion components — and shows that a single class can be reliable for the course yet a poor estimate of the teacher. What that means for fair evaluation.
Koji Education Team
Product
Answer box
Student ratings can be highly reliable in one sense and unreliable in another at the same time. Generalizability theory (G-theory) shows that the average rating from a single class is a stable estimate of that class's experience, but a much weaker estimate of an instructor's general teaching effectiveness, because a large share of the variance lives in the specific teacher-by-course-by-occasion combination rather than in the teacher alone. Gillmore, Kane & Naccarato (1978) found that obtaining a dependable estimate of an individual instructor typically requires averaging across several courses or occasions, not one. The practical lesson: judge instructors on a body of evidence collected over time, never on a single semester's mean.
What the research says
Classical test theory treats reliability as a single coefficient — it asks only how much "true score" sits above "error." Generalizability theory, developed by Cronbach, Gleser, Nanda & Rajaratnam (1972), reframes the question: error is not one undifferentiated thing but the sum of identifiable sources (students, items, occasions, courses, raters), each contributing its own variance component. A G-study estimates how big each source is; a follow-up decision (D-) study then tells you how many students, items or occasions you need to reach a target dependability for a specific decision.
Gillmore, Kane & Naccarato (1978), in the Journal of Educational Measurement, applied this directly to student ratings of instruction, estimating the teacher and course variance components. Their central result is widely cited: when the object of measurement is the individual teacher, a meaningful portion of variance attaches to the particular teacher-and-course combination and to the occasion, so a single course rating generalises only weakly to the teacher in general. To reach the dependability levels institutions assume when they use ratings for personnel decisions, you must aggregate across multiple sections or terms. The number of students within one class is rarely the binding constraint; the number of independent classes usually is.
This dovetails with Marsh's (1984) landmark synthesis in the Journal of Educational Psychology. Marsh reported that student ratings are among the most thoroughly studied of all personnel evaluation measures and are, by usual standards, reliable when averaged over enough students — internal-consistency and inter-rater agreement within a class are strong once roughly 10–15 students respond, and class-average ratings are notably stable over time (the same instructor tends to receive similar ratings across years and even when re-rated by former students). Crucially, Marsh's evidence on dimensionality matters here too: ratings are multidimensional (a structure such as the SEEQ's nine factors), so a single overall number collapses distinct, separately reliable signals.
The two findings are complementary, not contradictory. Within a single class, ratings are reliable. Across the inference most institutions actually want — "how good is this teacher, generally?" — reliability depends on sampling enough occasions, exactly as G-theory predicts. Subsequent G-theory work on student ratings (e.g., studies of item specificity and section effects, and of generalizability across courses and teachers) has repeatedly reinforced this teacher-versus-course-versus-occasion decomposition.
Why it matters for course evaluation in practice
Most evaluation policies quietly conflate two questions:
- How was this particular offering of this particular course experienced? (a course/occasion question)
- How effective is this instructor as a teacher? (a person question)
G-theory says these need different amounts of evidence. A single well-sampled class answers (1) reliably. Answering (2) reliably requires several classes, because teacher-by-course and teacher-by-occasion interactions are real: the same lecturer can score 4.6 in a small senior seminar and 3.9 in a large required first-year module, and neither number is "noise" — both are valid for their context, but neither alone is a dependable estimate of the teacher.
This has concrete consequences for European QA and personnel committees:
- Single-semester judgements are statistically indefensible for high-stakes decisions about an individual. A reappointment case built on one term's mean ignores the dominant variance source.
- Sample size within a class is the wrong lever to obsess over. Beyond ~10–15 respondents, adding students yields diminishing returns; adding independent course offerings is what actually raises dependability for teacher-level inference.
- Collapsing dimensions throws away reliable signal. Because ratings are multidimensional, reporting only a single global score discards separately-reliable information about organisation, clarity, workload and feedback that a D-study would tell you to keep.
Limitations and honest caveats
- Variance components are not universal constants. The exact size of teacher, course and occasion components depends on the instrument, the discipline, the institution and the rating culture. Gillmore et al.'s specific numbers should not be transplanted wholesale; the structure of the argument generalises, the parameters must be re-estimated locally.
- Reliability is necessary, not sufficient. A rating can be perfectly reliable and still biased or invalid (see the evidence on gender, race and grading leniency). G-theory tells you how consistent a measure is, not whether it measures teaching effectiveness rather than charisma or leniency.
- D-study assumptions. Decision studies assume the facets you sample (courses, occasions) are exchangeable. If an instructor only ever teaches one atypical course, "averaging across courses" is not available and the dependability ceiling is lower — a limitation no amount of student volume can fix.
- Stability can mask drift. Marsh's stability findings are reassuring for aggregate use but can lull committees into ignoring genuine improvement or decline; longitudinal tracking is needed to separate stable signal from real change.
How Koji incorporates this
Koji is designed so that the unit of evidence matches the inference being made:
- Longitudinal, multi-cohort evidence by default. Rather than treating each end-of-term survey as a standalone verdict, Koji accumulates an instructor's evidence across courses, cohorts and cycles — the multiple "occasions" G-theory says you need before drawing teacher-level conclusions. Reports can show whether a pattern holds across contexts or is specific to one course.
- Distribution- and dimension-aware reporting. Koji preserves the multidimensional structure of feedback (using scale, single_choice, ranking and open_ended question types) instead of collapsing everything to one mean, retaining the separately-reliable signals Marsh documents.
- Probing reduces occasion noise. Because the AI moderator follows up on each answer, Koji captures why a cohort responded as it did, helping evaluators tell a stable teaching characteristic from a one-off contextual reaction (a difficult timetable slot, a room change), which is exactly the teacher-by-occasion interaction G-theory isolates.
- Mid-cycle and formative collection. Adding formative touchpoints increases the number of independent observations available for an instructor without simply lengthening one end-of-term form, improving dependability for the person-level question while keeping each instrument short.
- Guardrails against single-class ranking. Koji's reporting is built to discourage comparing instructors on one course's mean, nudging evaluators toward the body-of-evidence interpretation the measurement theory requires.
Framed honestly, Koji does not raise the reliability coefficient of any single survey — that is fixed by the data. What it changes is the evidence architecture around the survey, so that decisions are made on the sample of occasions the inference actually demands. The same engine powers general user and customer research at koji.so, where distinguishing a stable preference from a one-session reaction is the identical statistical problem.
A worked illustration
Imagine a faculty whose instructors each teach several courses, and you collect ratings from many students per course. A generalizability study partitions the total variance in ratings into components. A common (illustrative) pattern looks like this: a sizeable share of variance sits between students within a class (individual students simply rate differently); a share sits between courses taught by the same teacher (the teacher-by-course interaction); and a comparatively modest share sits in the stable teacher main effect — the thing personnel committees actually want to measure.
The between-students component is tamed quickly: average over 10–15 respondents and that source shrinks toward zero in the class mean. But averaging students does nothing for the teacher-by-course and teacher-by-occasion components — only sampling more courses and occasions reduces those. That is the whole point of a decision study: if the teacher-by-course interaction is large, the dependability of a one-course estimate of the teacher will be low no matter how many students answered, and you will need three, four or more course offerings to reach a defensible coefficient (often cited targets are around 0.80 for high-stakes use).
The practical reading is counter-intuitive but important: chasing an extra five responses in a single class buys you almost nothing for a teacher-level decision, whereas waiting for a second and third course offering can move the dependability from "inadequate for personnel use" to "adequate." Institutions that set a fixed minimum response rate per course but never specify a minimum number of independent courses are optimising the wrong facet entirely.
Related Resources
- How Many Responses Do You Need for a Reliable Course Evaluation?
- Can You Fairly Rank Instructors by Their Scores?
- What Do Student Evaluations Actually Measure? (Marsh & the SEEQ)
- Do Student Evaluations Measure Learning? (Uttl meta-analysis)
- How Many Scale Points Should a Question Have?
References
- Gillmore, G. M., Kane, M. T., & Naccarato, R. W. (1978). The generalizability of student ratings of instruction: Estimation of the teacher and course components. Journal of Educational Measurement, 15(1), 1–13. https://doi.org/10.1111/j.1745-3984.1978.tb00051.x
- Marsh, H. W. (1984). Students' evaluations of university teaching: Dimensionality, reliability, validity, potential biases, and utility. Journal of Educational Psychology, 76(5), 707–754. https://doi.org/10.1037/0022-0663.76.5.707
- Cronbach, L. J., Gleser, G. C., Nanda, H., & Rajaratnam, N. (1972). The Dependability of Behavioral Measurements: Theory of Generalizability for Scores and Profiles. New York: Wiley.
Related articles
How Many Responses Do You Need for a Reliable Course Evaluation?
Reliability of a class-average rating rises steeply with the number of students who respond. Generalizability theory — from Marsh''s syntheses to Dzakadzie (2026) — shows raters matter more than items, and a class of 12 cannot be evaluated like a class of 120.
Can You Fairly Rank Instructors by Their Course-Evaluation Scores? What Esarey & Valdes (2020) Proved
Even if course evaluations were unbiased, reliable, and valid, ranking instructors by their scores would still misclassify many good teachers. A walk through the Esarey & Valdes (2020) simulation and what it means for how you report and use evaluation data.
Do Student Evaluations Measure Learning? The Uttl Meta-Analysis Revisited
A research-grounded reading of the Uttl, White & Gonzalez (2017) meta-analysis on the SET–learning relationship, with implications for European course evaluation and quality-assurance policy.
How Many Scale Points Should a Course-Evaluation Question Have?
What the measurement literature — Preston & Colman (2000), Weng (2004), Dawes (2008) — says about the optimal number of response categories on rating scales, and why the answer for course evaluation is not just a number but a question about what a Likert item can and cannot capture.