New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Sector trends11 min read

Whose Voice Is in Your Course Evaluation? Equity, the Social Dimension, and the Students You Don't Hear

Course evaluation is skewed toward the students who were already doing well. With awarding gaps of nearly 19 points for some groups and selection bias worth a quarter of a standard deviation, the quiet non-responders are often the students your programme most needs to hear — and Europe's social-dimension commitments make that a governance problem, not just a methodological one.

Koji Education Team

Product ·

Short answer: Course evaluation systematically over-represents students who are engaged and already succeeding, and under-represents the disadvantaged, disengaged, and lower-attaining students who are most likely to be struggling — or to have already left. This is not a minor sampling nuisance: at a large European university, selection into who responds shifts average evaluation scores by about 28% of a standard deviation (Goos & Salomons, Research in Higher Education, 2017). Given Europe's explicit social-dimension commitments and documented awarding gaps, "whose voice is in the data?" is a governance question. But — and this is the honest complication — the usual fix, breaking results down by demographic, collides with the anonymity that makes feedback trustworthy in the first place.

The evidence that non-response is not random

The reassuring assumption behind any evaluation dashboard is that responders are a fair sample of the class. They are not. Goos and Salomons analysed more than 3,000 courses and found that correcting for who chooses to respond changes both average scores and course rankings — the total selection bias is worth roughly 28% of a standard deviation, comparable to the effect of a one-standard-deviation higher average grade. Evaluations are, in their finding, upward biased: the students who respond are disproportionately those who are doing well.

Supporting work in the survey-methodology literature points the same way, with non-responders in student evaluations tending to have lower grades and lighter course loads (Studies in Educational Evaluation, 2014). The pattern is consistent with what we have written about response-rate and non-response bias and its close cousin, survivorship bias: the students who withdrew, disengaged, or are quietly failing are exactly the ones missing from the feedback, so the instrument paints the course in the colours of its most successful survivors.

Why this is a social-dimension problem, not just a statistics problem

If under-response were random, it would cost you precision but not fairness. It is not random — it is patterned by exactly the characteristics European higher education has committed to caring about.

The Bologna Process has, since the 2007 London Communiqué, held that the student body should reflect the diversity of the population. The 2020 Rome Ministerial Communiqué reaffirmed this with an annex of "Principles and Guidelines to Strengthen the Social Dimension of Higher Education in the EHEA." And the students the social dimension is about are precisely those an evaluation under-hears. According to EUROSTUDENT, students whose parents do not hold a higher-education qualification are under-represented across most European systems — the minority of the student body in around 60% of countries — and are materially poorer, being on average 15 percentage points more likely to report their parents as not well-off.

That these groups experience higher education differently is not speculation. Advance HE's data on ethnicity awarding gaps in UK higher education for 2019/20 found an overall white/BME gap of 9.9 percentage points in the proportion awarded a first or 2:1 — and a gap for Black students specifically of 18.7 percentage points. Commuter students, meanwhile, show lower rates of engagement and belonging; the Higher Education Policy Institute called them a "much misunderstood and underappreciated group". When different groups are having measurably different experiences and one of those groups responds at lower rates, the average on your dashboard is not neutral — it is quietly weighted toward the students for whom the course already works.

The accountability stakes are rising

This is no longer only an ethical concern; it is increasingly a regulatory one. In England, every provider seeking access to public funding must hold an Office for Students-approved Access and Participation Plan setting out targets and interventions to close gaps in access, success, and progression for disadvantaged groups. An evaluation system that structurally cannot hear those groups is a poor foundation for the evidence such plans demand — and the same logic follows from the ESG's requirement that institutions gather information to serve "the needs of students and society."

But doesn't the obvious fix break anonymity? The strongest counterargument

Here the honest analysis has to slow down, because the intuitive remedy — disaggregate results by ethnicity, first-generation status, or commuter status — runs straight into a genuine dilemma, and pretending otherwise would be vendor hype.

Anonymity versus breakdown. Course evaluations are usually anonymous by design, precisely to elicit candour. That anonymity is what prevents disaggregation by demographic. So "does it hear at-risk students?" is often structurally unanswerable within the instrument itself, and bolting on demographic questions risks deanonymising respondents in small classes, chilling the very honesty that makes feedback worth collecting. We have written about this tension in the context of GDPR and anonymity in course evaluation.

Small subgroup numbers. Even where a breakdown is possible, widening-participation subgroups within a single module are often small enough that results are statistically unstable, suppressed for disclosure control, or dominated by one or two vocal respondents — the opposite of hearing them reliably.

The fix assumes you can reach them. Boosting response rates or weighting for non-response assumes you can identify and reach the under-responding groups, which for anonymous evaluation you usually cannot. And mandatory or incentivised response can simply trade non-response bias for coercion bias.

Tokenism. Foregrounding "at-risk voices" can slide into treating a handful of minority responses as representative, or into performative equity metrics that satisfy a reporting requirement without changing any teaching.

The defensible thesis the evidence actually supports is therefore narrower and more useful than a slogan: awarding gaps and EUROSTUDENT data prove different groups have systematically different experiences; selection-bias evidence proves standard evaluation skews toward the more engaged and higher-attaining; but the anonymity and small-N problems mean the answer is not simply "break the data down by demographic." It is a design problem before it is a reporting problem. The goal is an instrument that lowers the barrier to responding for everyone — so the sample self-corrects — rather than one that labels people after the fact.

The cost of a blind spot

Consider what the gap conceals. A module that quietly loses its commuter and first-generation students by the fifth week can still post a healthy average, because the students who remain — and who respond — are disproportionately the ones for whom it worked. The programme team reads a reassuring number and changes nothing, while the retention and awarding gaps the institution is publicly accountable for widen out of view. Differential non-response is not merely a measurement flaw; it is a mechanism by which an evaluation system can launder inequity into a clean-looking dashboard, term after term, and call it evidence.

Designing for the students you don't currently hear

If the lever is design, the questions become concrete. Does the instrument take five seconds or fifteen minutes? Is it available in a language the student is comfortable in? Does it work on a phone during a commute? Does it feel like a form to endure or a conversation worth having? These are the differences that move response among exactly the disengaged, time-pressured, and less-confident students who currently opt out. We have argued elsewhere that accessible survey design is itself a non-response intervention.

This is where an AI-native, conversational approach has a fair claim. Koji for Education replaces the long static grid with a short, adaptive, AI-moderated interview that meets students on mobile, in the moment, and adapts its follow-ups — lowering the effort and self-presentation cost that suppresses response among less-engaged students. Because thematic analysis works on open conversation, a minority experience can surface as a theme even when the numbers are too small for a demographic cross-tab — you can hear the signal without deanonymising the student. Formative, mid-cycle collection reaches students before the disengaged ones drift away and become survivorship-bias casualties, and its EU/GDPR-appropriate data handling keeps that reach compatible with the anonymity that candour depends on. To be clear about the limit: Koji does not "solve" non-response, and no honest tool claims to — it mitigates the design barriers that drive differential drop-off, which is the part actually within an instrument's control. The same conversational engine underpins inclusive user and customer research at koji.so.

The average on your dashboard answers a question you did not ask: how did the course go for the students who were already doing well enough to fill in the form? The students whose experience your quality commitments most depend on are the ones the old instrument was built to miss. Design for them first.