Your Instructor Averages Are Mostly Noise: The Case for Multilevel Models in Course Evaluation
A departmental league table of mean evaluation scores treats sampling noise as if it were teaching quality. Variance-decomposition research shows only a modest share of the variation sits between instructors. Multilevel models and generalizability theory are the honest way to separate signal from noise.
Koji Education Team
Product ·
Bottom line up front: When a quality-assurance office ranks lecturers by their mean course-evaluation score, it is treating a number as if it were a clean measurement of teaching quality. It is not. Decades of variance-decomposition research show that only a modest share of the variation in student ratings sits between instructors; large portions belong to individual students, to the specific student-teacher pairing, and to the course context. A raw departmental league table mistakes much of this noise for signal. Multilevel (hierarchical) models — and their close cousin, generalizability theory — are the statistically honest way to ask a deceptively simple question: how much of this score is actually about the teacher? The answer, uncomfortably often, is "less than the spreadsheet implies."
The problem with a column of averages
A course-evaluation score is never a single, free-standing measurement. It sits at the top of a hierarchy. Individual students rate a course; those students are clustered within a section; sections are taught by an instructor; instructors sit within a programme and a discipline. Each of those levels contributes its own variation. When you collapse all of it into one mean — 4.1 out of 5 — and place that mean next to another lecturer's 3.9, you are implicitly claiming that the 0.2 gap is a property of the teachers. Statistically, that claim is usually unsupported.
The cleanest way to see this is to decompose the total variance in ratings into its sources. A widely cited variance-components analysis published in Assessment & Evaluation in Higher Education (2016) did exactly that and found something awkward for league-table thinking: a substantial share of variance was attributable to students rather than teachers, and the single strongest source of variation was the student-by-teacher interaction — the idiosyncratic fit between a particular student and a particular instructor. In other words, much of what an average captures is not "this lecturer is good" but "this particular mix of students happened to click, or not, with this particular lecturer."
This is not an isolated finding. A multivariate generalizability-theory study in Frontiers in Psychology (2018) reached compatible conclusions, partitioning rating error into student, item, teacher, and interaction components and showing that the teacher signal competes with several large error sources. And the most pointed evidence comes from Uttl, White and Gonzalez's 2017 meta-analysis in Studies in Educational Evaluation: once small-sample bias is corrected, student-evaluation ratings explain at most around 1% of the variability in how much students actually learn, with a corrected ratings-learning correlation of roughly r = 0.08 — and near zero (r = -0.02) after accounting for prior ability. If the construct itself is only weakly tied to learning, the third decimal place of a mean cannot bear the weight institutions place on it.
What a multilevel model actually does
Multilevel modelling takes the hierarchy seriously instead of flattening it. Three of its features matter directly for evaluation:
- Variance partitioning and the intraclass correlation (ICC). The model estimates how much of the total variance lives at each level. The ICC for the instructor level tells you the share of variation genuinely attributable to which teacher you got. When that share is small, instructor-to-instructor differences in raw means are mostly noise, and ranking on them is indefensible.
- Partial pooling (shrinkage). Rather than trusting each lecturer's raw mean equally, the model "shrinks" estimates toward the overall average in proportion to how little data supports them. A lecturer with 12 responses and a 4.6 is pulled toward the centre far more than one with 120 responses and a 4.6, because the small sample is less trustworthy. This is the same statistical humility that underlies regression to the mean in year-over-year scores and the uncertainty in small-class evaluations.
- Adjusting for context. Random and fixed effects let you hold class size, discipline, course level, and modality constant before comparing instructors — controlling for the structural factors that systematically depress or inflate scores independent of teaching.
A worked intuition
Picture two lecturers in the same department. Anna scores 4.5 on a section of 12 students who returned evaluations; Ben scores 4.2 on a section of 120. A naive table ranks Anna above Ben. A multilevel model does something different: it notices that Anna's estimate rests on a handful of responses with wide uncertainty, shrinks it toward the departmental mean, and reports a credible interval that comfortably overlaps Ben's. The "gap" evaporates. The model has not hidden the truth; it has stopped manufacturing a difference the data never contained. This is precisely the discipline missing when an institution publishes a ranked list and asks promotion committees to read meaning into a 0.3 difference that may not be real.
"But doesn't this just let weak teachers off the hook?"
This is the strongest objection, and it deserves a direct answer. Critics worry that talking about variance components and shrinkage is a sophisticated way of saying "you can never hold anyone accountable." It is not.
Multilevel modelling makes accountability fairer and more defensible, not impossible. It does not claim teaching is unmeasurable or that all lecturers are equal; it claims that a single ranked mean is the wrong instrument. Where genuine, reliable instructor-level signal exists, a properly specified model will find it — and a difference that survives partial pooling and adjustment for context is far more credible evidence for a tenure or promotion file than a raw average ever was. The honest version of accountability is better evidence, not more of it.
Two limitations cut the other way and should be stated plainly. First, multilevel models are data-hungry; a small institution with few sections per instructor may not be able to estimate stable instructor effects at all — which is itself useful information, because it means raw averages there are even less trustworthy. Second, reliability is necessary but not sufficient. A reliably-estimated instructor effect can still encode demonstrated biases in student ratings — gender, accent, grading leniency — rather than teaching quality. Partitioning variance tells you how much signal is there; it does not tell you the signal is valid. That is why variance modelling has to be paired with triangulation across multiple evidence sources, not treated as a final verdict.
What this means in practice
For institutional-research and QA teams, the implications are concrete:
- Stop ranking on raw means. Report distributions and uncertainty, not point estimates to two decimals. A score without a confidence interval is a rumour.
- Set minimum-reliability thresholds. Don't report or compare instructor effects below a response count where the estimate is dominated by noise.
- Adjust before comparing. Hold discipline, level, class size, and modality constant — these are well-documented structural confounds, not teaching.
- Treat the number as a prompt, not a conclusion. The most useful question a low score raises is why, and a Likert mean cannot answer it.
This last point is where modern evaluation infrastructure earns its place. Legacy SET platforms — EvaSys, Qualtrics survey exports, paper forms — are built to produce the very averages this article warns against: a section mean, a department mean, a trend line, all stripped of the structure beneath them. Koji for Education is designed around the opposite premise. Its AI-moderated conversational interviews probe beyond the number — when a student rates "assessment and feedback" low, the moderator asks what specifically, surfacing the mechanism a Likert item conceals. Automatic thematic analysis turns thousands of open-text responses into structured, auditable themes; quality scoring flags low-information responses; and programme- and institution-level reporting can present findings as distributions and weighted evidence rather than naked, falsely-precise means. Because the AI moderation is standardized, it removes the human-moderator inconsistency that adds yet another variance component to traditional interview-based evaluation. Collection can be formative and mid-cycle, and the whole pipeline is GDPR/AVG-compliant and built for European data-handling expectations.
The same conversational interview engine underpins the main Koji platform for general user and customer research — the statistical humility is the same wherever you are trying to turn messy human feedback into a defensible decision.
A column of averages feels like measurement. Multilevel thinking is what turns it into one. If your evaluation system can only give you the mean, it is hiding most of what you need to know — and quietly inviting you to act on noise.
Ready to move beyond the average? See how Koji for Education surfaces the structure beneath your scores.