New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Methodology9 min read

How Many Students, How Many Questions? Generalizability Theory and the Dependability of Course Evaluation

A course-evaluation score is only as trustworthy as the number of raters and items behind it. Generalizability theory tells you exactly how much of your data is signal and how much is noise — and how many responses you actually need.

Koji Education Team

Product · July 4, 2026

Before you rank an instructor on a 4.1, ask a prior question: if you re-ran the same evaluation with a different group of students and a slightly different set of items, would you get 4.1 again — or 3.6? If the answer is "we have no idea", then the number is not a measurement, it is an anecdote with a decimal point. Generalizability theory (G-theory) is the framework that answers this question directly, and for course evaluation the answer is often uncomfortable: a large share of the variance in your scores is measurement error, not teaching quality.

This is the single most important idea a quality-assurance office can borrow from psychometrics, and it is routinely ignored in favour of averaging.

What classical reliability gets wrong

Most evaluation reports, if they mention reliability at all, quote Cronbach's alpha — a single number describing internal consistency. The problem is that alpha answers only one question ("do the items hang together?") and silently assumes there is one undifferentiated source of error. Real course-evaluation data has several error sources at once: which students happened to respond, which items you happened to ask, which section or occasion you sampled. Classical test theory cannot separate them. It lumps everything not-true-score into one bucket and hopes.

This matters because the distinction between reliability and validity is already doing heavy lifting in evaluation debates — as we discuss in reliability vs validity and why the difference matters. G-theory sharpens the reliability half into something actionable.

Generalizability theory in one paragraph

G-theory was developed by Cronbach, Gleser, Nanda and Rajaratnam (1972) and is set out in the standard reference by Shavelson, Webb and Rowley. Its core move is to treat any observed score as a sample from a universe of admissible observations and to use analysis of variance to decompose the observed-score variance into components: variance due to the object of measurement (the instructor or course — the part you want), and variance due to each source of error (items, students/raters, occasions, and their interactions — the part you want to quantify and minimise).

The framework splits neatly into two studies:

  • A G-study (generalizability study) estimates the magnitude of as many sources of measurement error as possible. It tells you, for your instrument, how much of the total variance is instructor, how much is item, how much is student, how much is unexplained interaction.
  • A D-study (decision study) uses those variance components to answer the practical question: given this error structure, how many students and how many items do I need for a dependable score? It lets you simulate configurations before you spend the effort.

The output is a generalizability coefficient (for relative, norm-referenced decisions) or a dependability coefficient (for absolute, criterion-referenced decisions like "is this course above our quality threshold?"). Both range 0 to 1 and answer the question classical alpha cannot: how well does this score, from this many raters and items, generalise to the universe of scores you could have observed?

What G-studies of student ratings actually find

The empirical literature on applying G-theory to student evaluations of teaching is now substantial, and two findings recur.

First, the instructor/course variance component — the thing you are trying to measure — is often modest relative to the error components. A large fraction of the variance in raw ratings is attributable to which students responded and to student-by-item interactions, not to stable differences between instructors. This is the psychometric restatement of a point we make elsewhere: instructor averages are mostly noise, and multilevel models exist to prove it. G-theory reaches the same conclusion by a different route and quantifies it precisely.

Second, dependability is highly sensitive to the number of raters. Because student sampling is usually the dominant error source, dependability climbs steeply as response numbers rise and collapses when they fall. A D-study will typically show that a handful of respondents produces a dependability coefficient far below any defensible threshold, which is why small classes are a genuine measurement problem, not just a small-sample inconvenience — a theme we develop in what is a fair score for a class of 12. Recent work using G-theory to identify the optimal number of raters and items for dependable evaluation (Frontiers in Education, 2026) continues to show that the required configuration depends entirely on the underlying variance structure — there is no universal "enough".

Item specificity matters too: earlier research (Marsh and colleagues; Gillmore, Kane and Naccarato) found that generalizability is substantially influenced by how specific the items are, while section effects tend to be small. The design of the instrument, not just the size of the sample, sets the ceiling on how dependable your scores can be.

The strongest counterargument — "This is psychometric over-engineering"

The honest objection is administrative: We are a teaching-and-learning office, not a measurement lab. Running variance-component analyses on every course is unrealistic, and deans want a number, not a dependability coefficient.

Three responses.

First, you do not need to run a G-study on every course — you need to run one on your instrument, once. A single well-designed G-study on a representative slice of your data tells you the error structure of your evaluation system: how many responses buy you a dependable score, and which items pull their weight. That result then governs policy for thousands of courses. It is a one-off investment, not a per-course tax.

Second, the alternative is not "no psychometrics" — it is bad psychometrics. Every time an office reports a raw mean to three significant figures and lets a committee compare a 4.1 to a 3.9, it is making an implicit reliability claim: that the number is stable enough to compare. G-theory simply makes that claim explicit and usually refutes it. Choosing not to quantify error does not make the error go away; it just hides it behind a decimal point, which is precisely the trap behind reading a 0.3-point difference as real.

Third, dependability is exactly what accountability decisions require. If student ratings feed into probation, promotion or programme review, the burden of proof is higher, not lower. A dependability coefficient is the minimum evidence that a score is fit for a high-stakes purpose. Refusing to compute it while still using the scores for consequential decisions is the indefensible position.

What to do with this

  • Commission one G-study on your evaluation instrument using existing data. Decompose variance into instructor, student, item and interaction components.
  • Run a D-study to set a response threshold. Instead of an arbitrary "50% response rate" rule, derive the number of respondents that achieves an acceptable dependability coefficient given your variance structure. This puts response-rate policy on an evidentiary footing.
  • Report dependability alongside the mean, and refuse to compare or rank scores whose dependability is below threshold — especially for small classes.
  • Prune low-information items. If a G-study shows certain items add error without adding instructor variance, they are costing response quality for nothing.

Where Koji fits

G-theory tells you that dependability rises with the number and quality of observations and with well-designed items. Koji for Education is built to improve both inputs. Its AI-moderated conversational interviews raise the information content of each response — a probed, elaborated answer carries more signal per respondent than a single tick on a Likert scale, which lifts the effective dependability of a given number of participants. Koji's six structured question types and quality scoring let you design instruments whose items actually discriminate, rather than padding the form with low-information questions that a G-study would flag as pure error. Its programme- and institution-level reporting is the natural home for reporting dependability and uncertainty alongside point estimates, rather than a bare mean. Koji does not repeal the mathematics of measurement — a class of eight is still a class of eight — but it mitigates the two levers G-theory identifies: richer observations and better-targeted items.

The same interview engine underpins the main Koji platform, where the dependability of customer-research findings hinges on exactly the same question — how many respondents, how good the instrument.

The bottom line

A course-evaluation score without a dependability estimate is a claim without a confidence. Generalizability theory is the honest accounting of where your variance comes from — and its verdict, again and again, is that raw student-rating means carry more error than institutions like to admit, especially in small classes. You do not have to become a psychometrician to use it. You have to run one G-study, set your thresholds from evidence rather than folklore, and stop treating a number as a measurement until you know how much of it is real.

Want evaluation data whose dependability you can actually defend? Explore Koji for Education.