New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Methodology9 min read

What Course Evaluations Miss: Measuring Belonging, Not Just Satisfaction

A satisfaction average tells you whether students liked a course. It does not tell you whether they felt they belonged in it — and belonging is the better predictor of whether they stay, achieve, and graduate. Here is the evidence, and how to measure belonging without tokenism.

Koji Education Team

Product ·

Most course evaluations measure the wrong construct. They ask whether students were satisfied, whether the lecturer was clear, whether the assessment was fair — and then average the answers. Yet the variable that most consistently predicts whether a student persists, achieves, and graduates is not satisfaction. It is their sense of belonging: the felt sense of being a legitimate, valued member of an academic community. Belonging is measurable, it is distributed unequally across student groups, and it is far more actionable than a 4.1-out-of-5 satisfaction mean. If your evaluation instrument cannot see it, you are flying blind on the thing that matters most.

Why belonging, and why now

The construct has a long pedigree. Terrell Strayhorn defines sense of belonging in higher education as students' "perceived social support on campus, a feeling or sensation of connectedness, and the experience of mattering or feeling cared about, accepted, respected, valued by, and important to the campus community." The empirical case is strong. A 2022 systematic-style study in the Journal of Further and Higher Education (Pedler, Willis & Nieuwoudt) found that students with a greater sense of belonging report higher motivation, greater academic self-confidence, more engagement, and — critically — a lower likelihood of considering leaving before they complete (Pedler et al., 2022). A 2025 systematic literature review synthesising 66 empirical studies on belonging and retention reached the same conclusion: belonging is one of the more robust predictors of persistence in the literature (Tandfonline, 2025).

Belonging matters precisely because it is not the same as satisfaction. A student can rate a course 4 out of 5 — the content was fine, the slides were clear — while quietly concluding that "people like me don't really fit here." That student is at risk in a way no satisfaction item will detect.

Belonging is an equity variable, not a wellness nicety

The reason belonging belongs in your quality-assurance toolkit, and not just your student-wellbeing strategy, is that it is unequally distributed — and that inequality tracks attainment gaps. In England, the degree-awarding gap is stubborn: roughly 68% of Black, Asian and minority-ethnic students are awarded a first or upper-second, compared with around 81% of white students, a gap of about 13 percentage points that has barely moved since 2017/18 (Office for Students degree-awarding gaps). A widely-accepted strand of the explanation is that minoritised students more often report feeling "out of place" — a belonging deficit that HEFCE research as far back as 2015 linked to retention and success. If your course evaluation reports a single satisfaction mean per module, it cannot tell you which students feel they belong and which do not. It averages the gap away.

This is the methodological heart of the argument. Averaging a Likert scale destroys the distributional information that matters for equity (a point we make in detail in why averaging Likert scores misleads). Belonging, measured well, is reported as a gap between groups, not a single number — and a gap is something a programme team can act on.

How to measure belonging without tokenism

Three principles separate credible belonging measurement from box-ticking.

1. Use validated multi-item scales, not a single "Do you feel you belong?" question. Belonging is multi-dimensional — it includes acceptance, mattering, and academic fit. Single-item measures are noisy and easy to game. Established instruments adapt items such as "I feel like I am part of this learning community" and "I feel my contribution is valued in this course," answered on agreement scales, then analysed for internal consistency.

2. Disaggregate, then protect anonymity. The entire value of belonging data is in comparing groups — first-generation vs continuing-generation, domestic vs international, by ethnicity. But disaggregation in small cohorts risks identifying individuals. This is a genuine tension that demands careful minimum-cell-size rules and confidentiality safeguards (see anonymity vs confidentiality under GDPR).

3. Pair the number with the "why." Knowing that international students score belonging 0.6 points lower is a prompt, not an answer. You need the lived texture — what, specifically, makes them feel peripheral. A group-work norm? An assessment that assumes cultural reference points? Numbers locate the problem; open, probing dialogue explains it.

But doesn't this just turn evaluation into a feelings survey?

This is the strongest objection, and it deserves a direct answer. Critics argue that universities should measure teaching, not emotion — that belonging is a vague, soft construct, that students will conflate it with whether they liked the lecturer, and that chasing belonging scores invites grade inflation and pandering.

Three responses. First, belonging is not vague when measured with validated instruments; it has decades of construct-validation behind it and predicts hard outcomes — retention and attainment — better than satisfaction does. Second, the concern about conflation is real but is an argument for better measurement, not none: well-designed items separate "I felt I belonged" from "I enjoyed the class." Third, and most importantly, belonging is not a popularity contest. A demanding course with high standards can produce strong belonging precisely because students feel the challenge is fair and that they are trusted to meet it. Belonging is about legitimacy and support, not leniency. The active-learning literature shows students sometimes enjoy effortful learning less even as they learn more (see the active-learning penalty) — belonging, unlike satisfaction, does not punish rigour.

A fair limitation: belonging is one signal among several. It should triangulate with attainment data, engagement analytics, and direct teaching observation — never stand alone. We are arguing for adding a predictive, equity-sensitive lens, not replacing every other measure with it.

Where Koji fits

Measuring belonging well is exactly the kind of task that exposes the limits of a static Likert form and plays to the strengths of an AI-native approach. Koji for Education runs AI-moderated conversational interviews that can ask a validated belonging item and then probe — "You said you don't always feel part of the group; can you tell me about a moment that captured that?" — at the scale of an entire cohort, with standardized, bias-aware moderation rather than the inconsistency of dozens of human interviewers. Its six structured question types (including scale and open-ended) let you combine a quantified belonging measure with the qualitative "why" in a single instrument, and automatic thematic analysis surfaces the recurring drivers of belonging gaps across hundreds of transcripts without a research assistant reading every line. Reporting rolls up to programme and institution level, so a dean can see the belonging gap by student group, and all of it is handled in a GDPR/AVG-compliant, EU-appropriate way with the confidentiality safeguards disaggregated data demands. The same AI interview engine underpins the main koji.so platform used for general customer and user research — belonging, after all, is a special case of understanding how people experience a system.

The goal is not to chase a higher belonging score. It is to see the students your satisfaction average has been quietly hiding — and to act before they leave.

A note on cadence: belonging is not a one-shot measure

One practical caution before the summary. Belonging is not a fixed trait a student either has or lacks; it shifts over a course and a degree, often dipping at predictable pressure points — the first weeks, the first piece of summative assessment, the transition into a final-year project. A single end-of-module belonging score, like a single satisfaction mean, collapses that trajectory into one blurred number. The more useful instrument samples belonging formatively — early enough that a programme team can intervene for the current cohort, not only redesign for the next one. An early-semester signal that commuter students or international students feel peripheral is a problem you can still act on in week five; the same signal in an end-of-term report is an epitaph. Measuring belonging well therefore means measuring it more than once, lightly, and acting between measurements — which is exactly the formative posture that static end-of-term surveys make impractical.

The bottom line

Satisfaction tells you whether students liked the course. Belonging tells you whether they will still be here next year, and whether your attainment gaps have a cause you can address. Measure it with validated items, disaggregate it carefully, pair it with the "why," and triangulate it with hard outcomes. Then do something with what you find — because an unmeasured belonging gap is still a belonging gap.