New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Methodology8 min read

Why You Can't Compare a 4.1 in Engineering to a 4.4 in History: The Benchmarking Trap

Ranking instructors and departments against each other on a shared evaluation scale feels objective. It is one of the most statistically indefensible things quality offices routinely do — and the discipline-baseline evidence explains why.

Koji for Education

Research & Editorial Team ·

Bottom line: Average student-evaluation scores differ systematically by discipline, level, class size, and assessment type — for reasons that have nothing to do with teaching quality. A 4.1 in a large quantitative engineering module and a 4.4 in a small humanities seminar are not on a comparable scale, so ranking them against a single institutional benchmark is measuring the subject, not the teacher. If you must benchmark, benchmark within like-for-like reference groups and report uncertainty — or, better, stop relying on a number whose baseline you cannot hold constant.

The benchmark feels objective. That is the problem.

Almost every quality-assurance dashboard does the same thing: it places every module's mean rating on one axis, draws an institutional or faculty average as a reference line, and flags whoever falls below it. It looks rigorous. It produces a clean ranking. And it quietly assumes the one thing that is demonstrably false — that a point on the scale means the same thing in every subject, at every level, in every class size.

It does not. The baseline against which you are measuring teaching moves around for reasons the teacher does not control. When you ignore that, you are not benchmarking teaching effectiveness. You are benchmarking the discipline, the cohort, and the format, and attributing the result to the individual.

What systematically shifts the baseline

Discipline. This is the best-established confound. Evaluation scores tend to run higher in the arts, humanities and social sciences and lower in the natural sciences, engineering and mathematics. The University of California, Berkeley's Center for Teaching and Learning, reviewing the evidence on its own evaluations, notes plainly that scores differ by discipline and that comparing an instructor's average to a department or campus average is, by itself, uninformative. Students in quantitative fields also write less, and less enthusiastically, in open comments — so even the qualitative data is not comparable across the divide.

Class size. Smaller classes tend to receive higher ratings than large ones. A seminar of 12 and a lecture of 300 are different measurement instruments with different baselines, even before you ask anything about teaching.

Grades and workload. There is a persistent, if contested, association between expected grades, perceived workload and ratings. Modules that are harder or graded more strictly carry a baseline headwind unrelated to instructional quality.

The deeper validity problem. Underneath all of this sits a finding that should give any benchmarker pause. The most rigorous meta-analysis of the multisection literature, Uttl, White and Gonzalez (2017), Studies in Educational Evaluation, re-analysed decades of studies and found that once you account for small-sample bias, the correlation between average evaluation ratings and actual student learning is essentially zero (best estimate around r = 0.08, and near zero after controlling for prior ability). If the thing you are ranking barely correlates with learning within a comparable setting, ranking it across incomparable settings compounds an already weak signal with a structured artefact.

A worked illustration

Imagine an engineering lecturer running a required 250-student second-year module on thermodynamics, assessed by a hard final exam, and a literature tutor running an optional 14-student final-year seminar assessed by a reflective essay. Suppose both are, by every independent measure — peer observation, learning gain, alumni testimony — excellent.

On a shared 5-point scale, the engineer might land at 4.0 and the literature tutor at 4.5. Each number is "true" in the sense that students really did circle those responses. But the half-point gap is almost entirely manufactured by discipline, class size, electivity and assessment type. Put them in a single ranked list under a 4.2 institutional benchmark, and you have just told a panel that the engineer is "below average" and the tutor is "above" — a conclusion the underlying teaching does not support. Repeat this across a faculty and you systematically disadvantage exactly the people teaching large, compulsory, quantitative service courses: often the hardest and most thankless teaching in the institution.

"But we have to compare somehow — and we adjust for it"

This is the fair counterargument, and it has two parts.

First: we cannot run a university without some comparison. Correct. The argument here is not against all comparison; it is against comparison on a scale whose baseline you have not held constant. The defensible version is like-for-like benchmarking: compare a module only against other modules of similar discipline, level, size and assessment, and never against a single global mean. Even then, report the comparison with its uncertainty — a difference of 0.1 or 0.2 between two modules of 20 students each is, statistically, almost always within noise. Treating it as a ranking is false precision.

Second: we already statistically adjust for discipline and class size. Some institutions do, and that is genuinely better than a naive average. But adjustment models inherit the validity ceiling above them. If the raw measure correlates near zero with learning, a regression-adjusted version of it is a more sophisticated estimate of a weak construct. Adjustment narrows the artefact; it does not manufacture validity. And most adjustment models still cannot touch the things that move feedback most — what students were actually responding to, and whether their concern was real.

There is a third, blunter objection worth naming: isn't this just protecting underperformers from accountability? No — the opposite. Honest benchmarking protects the credibility of the cases where action is genuinely warranted. When a panel knows the institution rank-orders raw means across incomparable contexts, every flagged academic can reasonably claim the process is rigged, and they are sometimes right. Defensible comparison is what lets you act with confidence when action is due.

What to do instead

The practical path is not to abandon student feedback — it is to stop asking a single averaged number to do work it cannot do.

  1. Benchmark within reference groups, never against one global mean. Group by discipline, level, class size and assessment mode before you compare anything.
  2. Report uncertainty, not just a point. A mean from 18 respondents is an estimate with a wide interval; present it that way and most "rankings" dissolve into ties.
  3. Shift weight from the score to the substance. What a student says, and whether it is corroborated, travels across disciplines far better than where a 4.1 sits relative to a 4.4.

That last point is where a conversational approach earns its place. Koji for Education does not begin from a single comparable digit; it runs a standardised, AI-moderated interview that probes each student's reasoning across six structured question types and then applies automatic thematic analysis across the whole cohort. Because the moderation is consistent and bias-aware for every respondent and every department, the comparison that emerges is one of themes and their prevalence — "across these five modules, the recurring issue is feedback turnaround" — rather than a spurious league table of means on a scale that was never comparable to begin with. It also surfaces a quality score and lets you track whether issues are being closed at programme and institution level, with GDPR/AVG-compliant data handling appropriate to European institutions. None of this eliminates the comparability problem — students in different subjects are genuinely different populations — but it stops you from laundering that difference into a ranking and lets you compare the things that actually transfer.

The same AI interview engine powers the main Koji platform for general customer and user research, where the comparability-across-segments problem looks remarkably similar.

The reframe

A benchmark is only meaningful if the scale underneath it is stable. In course evaluation it is not: the baseline moves with discipline, size, level and assessment, and the underlying measure barely tracks learning. Comparing a 4.1 to a 4.4 across that gap is not objectivity — it is precision without accuracy. Benchmark like with like, carry the uncertainty, and move your decisions onto evidence that survives the crossing between subjects.

Want comparisons your academic board will actually trust? See how Koji for Education works.