New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Sector trends10 min read

What Mean Satisfaction Cannot See: Inclusive Teaching, the Awarding Gap, and the Case for Disaggregated Evaluation

A course can score 4.2 out of 5 overall and still be failing a fifth of its students. Persistent awarding gaps prove that averaged satisfaction hides exactly the inequities quality assurance claims to care about. Evaluating inclusive teaching requires disaggregation and qualitative depth — the two things a single mean is built to destroy.

Koji Education Team

Product · July 14, 2026

The short answer: A single averaged satisfaction score is structurally incapable of detecting whether a course serves all its students equally — because averaging is, by definition, the act of erasing subgroup differences. Yet the equity gaps that averaging hides are large, persistent and well documented: in UK higher education the degree-awarding gap between white and Black, Asian and minority ethnic students has sat around ten percentage points for two decades. If a course's teaching is experienced very differently by different groups of students, an overall mean of 4.2 will report 'good' while quietly averaging over a much worse experience for some. Evaluating inclusive or culturally responsive teaching therefore requires two things the standard instrument actively suppresses: disaggregation of results by student group, and qualitative depth that can surface how the course lands for students unlike the majority. This is both an equity argument and a validity argument.

The number that hides the thing you care about

Start with the arithmetic, because it is the whole problem. A mean is a compression: it takes a distribution and throws away everything except its centre of mass. If sixty students rate a course 4.6 and fifteen rate it 2.8, the reported mean lands somewhere reassuring — and the fifteen disappear. When those fifteen are not random, but cluster by ethnicity, first-generation status, disability, or being an international or commuting student, the mean is not merely lossy; it is systematically blind to inequity. It reports the majority experience and calls it the course experience.

This is not a hypothetical. Attainment differences by student background are among the most durable findings in the sector. Advance HE's analysis of UK higher education found a white / Black, Asian and minority ethnic degree-awarding gap of around 9.9 percentage points, with far larger gaps for some groups — on the order of 28 percentage points for Black students in the year analysed — and, tellingly, a gap that has persisted since the sector's first statistical reports two decades ago (Advance HE, degree attainment gaps). Whatever is producing those outcome gaps — and the causes are contested and multiple — the student experience of teaching is plausibly part of the picture. And it is precisely the part that an averaged satisfaction score is designed not to show.

Culturally responsive teaching, and what it asks of evaluation

The pedagogical response to these gaps travels under several names — culturally responsive teaching, inclusive teaching, equity-minded practice — but the common thread is that good teaching is not identical for every student, and that a course can be experienced as welcoming and legible by some students and alienating or opaque by others. As work on culturally responsive assessment argues, traditional methods can 'inadvertently reinforce inequalities by overlooking cultural differences and diverse ways of learning' (Every Learner Everywhere, on equity-minded assessment and culturally responsive teaching). The evaluation implication is direct: an instrument that measures teaching quality as a single number for 'the students' assumes the very homogeneity that inclusive teaching denies.

A key methodological move in the equity-assessment literature is therefore to disaggregate — to examine engagement and experience broken down by student identity characteristics rather than in aggregate (see this systematic review of culturally responsive practice measures). Disaggregation is the exact inverse of what a mean does. Where the average erases the subgroup, disaggregation restores it. The methodological argument and the equity argument converge on the same instruction: stop reporting only the centre of mass.

This is a validity problem, not only a fairness problem

It would be easy to file this under 'equity, do-the-right-thing'. But it is also a straightforward validity failure. If a course's teaching quality genuinely differs across student groups, then a single mean is not just ethically incomplete — it is an invalid summary of the construct, because it reports a parameter (the average) that does not describe the reality (a bimodal or group-structured distribution). This is the ecological fallacy wearing an equity coat: the aggregate does not represent the subgroups nested within it, and inferring 'the course works' from the mean is a cross-level error. It connects, too, to a broader problem we have written about — differential non-response, where the students who are least well served are also least likely to respond, doubly hiding their experience.

But won't disaggregation produce tiny, unreliable, identifiable subgroups?

This is the serious counterargument, and it deserves a serious answer. Break a class of eighty into ethnicity-by-gender-by-disability cells and you get cells of two or three students: statistically unstable and a genuine risk to anonymity. Three responses to it.

First, disaggregation does not require reporting every cell as a mean. The point is to detect group-structured differences, and qualitative, thematic evidence can surface a pattern ('several international students described the assessment briefs as culturally opaque') without publishing a fragile subgroup average or exposing individuals. Second, where quantitative disaggregation is used, it must be paired with the statistical disclosure controls and small-sample uncertainty that any responsible small-n reporting demands — suppression thresholds, interval estimates, and no naming of near-identifiable individuals. Third, the alternative — refusing to look because looking is hard — is not neutral. It is a decision to keep the inequity invisible. The methodological difficulty of disaggregation is real; it is an argument for doing it carefully, not for continuing to hide behind the mean.

A second objection is that asking students about identity in an evaluation is intrusive, particularly under GDPR, where much of this is special-category data. This is a real constraint, not a fatal one: it means disaggregation should often lean on data the institution already holds (linking evaluation to existing, lawfully processed demographics under appropriate governance) rather than interrogating students afresh, and it means qualitative surfacing of themes is frequently the more proportionate route.

How Koji surfaces what the mean hides

Koji for Education is built to resist the compression that makes inclusive teaching invisible. Its automatic thematic analysis of open-text and conversational feedback is designed to surface patterns — including patterns concentrated in a subset of students — rather than dissolving them into an average, so a concern raised by a minority of the cohort registers as a theme instead of being outvoted by the majority mean. Its AI-moderated conversational interviews probe experience rather than collecting a rating, which is exactly the depth needed to understand why a course lands differently for different students — the 'the examples never reflected students like me' insight that no five-point scale will ever yield. Reporting is programme- and institution-level as well as course-level, letting quality teams see whether an experience gap recurs across a programme rather than treating each course in isolation. And because Koji is GDPR/AVG-compliant and built for EU data handling, disaggregation and free-text collection can be governed proportionately, with special-category data treated with the care it requires.

Koji does not claim to close awarding gaps — no evaluation tool does, and the causes run far wider than teaching. What it does is refuse to let a reassuring average conceal an unequal experience, giving quality teams the disaggregated, qualitative signal that inclusive teaching requires. The same engine underpins Koji's main research platform, where the identical discipline applies: a headline satisfaction score that hides how a product fails a segment of users is worse than useless, because it looks like reassurance.

The takeaway

An averaged satisfaction score is not a neutral summary. It is an active erasure of subgroup difference, and subgroup difference is exactly where the sector's most persistent inequities live. Evaluating inclusive teaching honestly means disaggregating results and collecting qualitative depth — carefully, proportionately, and with due regard for small samples and data protection. The alternative is to keep certifying courses as 'good' on the strength of a number built to look away from the students they serve least well.

For European quality assurance under the ENQA Standards and Guidelines (ESG), where the social dimension and equitable provision are explicit expectations rather than aspirations, this is not an optional methodological refinement. An evaluation system that structurally cannot see differential experience is, by construction, unable to evidence inclusive provision — and unable to tell a programme team where its inclusive-teaching effort is actually working. Disaggregation is how 'we treat all students the same' is tested against 'all students experience the course the same', and only one of those is a defensible claim.