Should Course Evaluations Measure Whether Students Feel They Belong? The Sense-of-Belonging Evidence
Sense of belonging predicts persistence, engagement, and mental health in higher education — often more reliably than satisfaction. Here is what the evidence says about measuring belonging in course evaluation, and how to do it without turning a survey into a diagnostic instrument.
Koji Education Team
Product
In brief: A student's sense of belonging — the feeling of being accepted, valued, and a legitimate member of an academic community — is one of the best-evidenced predictors of persistence, engagement, and wellbeing in higher education. Most course evaluations do not measure it at all; they measure satisfaction with the lecturer. The research (Gopalan & Brady, 2020; Pedler, Willis & Nieuwoudt, 2022; Strayhorn, 2018) suggests belonging is a distinct construct worth capturing at the module and programme level, especially for first-generation and minoritised students. But belonging is a climate variable, not an instructor-performance variable — measure it to improve courses and support students, never to rank staff.
The question this answers
End-of-term course evaluations were built around a simple, largely tacit theory of what makes a good course: a clear, organised, responsive lecturer produces satisfied students. That theory is not wrong, but it is narrow. A generation of research on student retention and engagement points to a variable the standard questionnaire almost never touches: whether the student felt they belonged in the room. This article asks whether course evaluation should measure sense of belonging, what the evidence actually supports, and where the idea breaks down.
What the research says
Belonging predicts the outcomes universities care about. The most authoritative recent evidence comes from Gopalan and Brady (2020), who analysed the Beginning Postsecondary Students Longitudinal Study — a nationally representative U.S. sample of roughly 23,750 first-time students drawn from more than 4,000 institutions. On average, first-year students only "somewhat agree" that they belong. At four-year institutions, sense of belonging predicted better persistence, engagement, and mental health even after extensive covariate adjustment, and racial-ethnic minority and first-generation students reported lower belonging than their peers. Critically, the pattern differed by sector: at two-year institutions the demographic gaps did not run the same way, a reminder that belonging is context-dependent rather than a fixed trait.
Pedler, Willis and Nieuwoudt (2022), writing in the Journal of Further and Higher Education (46(3), 397–408), surveyed 578 students and found sense of belonging positively correlated with motivation and enjoyment, and lower among students who had considered leaving. First-generation students again reported weaker belonging. Their framing is useful for evaluation designers: belonging is not a mood but a relational appraisal that tracks the outcomes — effort, persistence, self-esteem — a quality process is ultimately trying to protect.
Underpinning both is Terrell Strayhorn's synthesis, College Students' Sense of Belonging (2nd ed., Routledge, 2018), which argues belonging is a basic human need that becomes especially salient in novel or marginalising contexts and, when satisfied, drives positive outcomes such as achievement and retention. This connects the education-specific work to a much older psychological literature (Baumeister & Leary's 1995 "need to belong").
The through-line across these sources is consistent and replicated: belonging is distinct from satisfaction, it is unevenly distributed across student groups, and it predicts behaviour — dropping out, disengaging, seeking help — that a satisfaction score does not capture.
Why it matters for course evaluation in practice
Three practical implications follow.
First, belonging is a leading indicator; satisfaction is a lagging one. A student can rate a lecturer 4/5 for clarity and still be quietly disengaging because they feel like an impostor in the discipline. Because belonging predicts considered-leaving before withdrawal happens, a belonging item at mid-semester functions as an early-warning signal in a way an end-of-term satisfaction mean cannot. This connects directly to retention-focused evaluation and Tinto's model of departure (see our piece on course feedback as an early-warning system).
Second, belonging is an equity lens. Because first-generation and minoritised students systematically report lower belonging, an aggregate course mean can hide a belonging gap between groups. A programme that looks healthy on average may be quietly failing the students most at risk. This is where belonging data earns its place in quality assurance: it makes visible a disparity that a single average erases.
Third, belonging is actionable at the course level. Belonging is shaped by concrete, teachable practices — how a lecturer handles questions, whether names are learned, whether group work is structured to prevent exclusion, whether assessment feedback is framed as "you can meet this standard" rather than as a verdict. These are course-design choices, which means a belonging signal can be closed like any other feedback loop (see closing the feedback loop).
Belonging also sits naturally alongside constructs the corpus already covers. It overlaps conceptually with the "relatedness" component of Self-Determination Theory (see autonomy, competence, relatedness) and with the affective dimension captured by control-value theory (see achievement emotions). Designers should treat these as a family of related, partially overlapping climate measures rather than independent scales — and avoid asking six near-duplicate questions.
Limitations and honest caveats
A PhD reader will raise several objections, and they are right to.
Correlation, self-report, and confounding. Almost all of the belonging-outcomes evidence is observational. Gopalan and Brady adjust for extensive covariates, but adjustment is not randomisation; students who feel they belong may differ in unmeasured ways (prior attainment, family support, financial security) that independently drive persistence. Belonging is measured by self-report on short scales, inheriting all the reference-bias and response-style problems the rest of this knowledge base documents (see reference bias). Reported associations are real and replicated, but the causal arrow is not settled.
The construct is contested and multidimensional. "Belonging" is measured differently across studies — some use a single item, others multi-item scales (e.g. the PSSM tradition, or Freeman-style class-belonging items). Sector differences in Gopalan and Brady's data warn against assuming a universal effect. A borrowed U.S. scale may not be measurement-invariant across European systems and languages, so cross-context comparison requires the same caution we describe for cross-language equivalence.
The misuse risk is serious. The single most important caveat is governance. Belonging is a climate variable driven by peers, discipline culture, institutional signals, and the student's own history — not primarily by one instructor's competence. Attaching a belonging score to an individual staff member's appraisal would be a category error and would incentivise exactly the wrong behaviours. Belonging data belongs at course, cohort, and programme level, read for patterns and gaps, never used to rank teachers. It is also sensitive: a student disclosing that they feel they do not belong may be flagging distress, which raises a duty-of-care question about anonymity and follow-up.
How Koji incorporates this
Koji for Education is built to capture climate constructs like belonging without collapsing them into a lecturer-performance number.
- Structured belonging items. Belonging can be measured with
scaleitems (e.g. "I feel like I belong in this course / discipline") andsingle_choiceitems, positioned as course-climate questions distinct from instructor-performance questions so they are read and reported separately. - Conversational probing beyond the number. A low belonging rating is diagnostically thin on its own. Koji's AI-moderated conversational interview can follow a low score with a neutral, non-leading probe ("What would have made you feel more part of this course?"), surfacing the mechanism — an unwelcoming seminar, unstructured group work, a first-assessment shock — that a Likert item cannot. This is designed to mitigate the reference-bias problem: the student explains what belonging meant to them in context.
- Bias-aware, disaggregated reporting. Because the equity value of belonging comes from disaggregation, Koji is designed to report belonging by cohort or subgroup where sample sizes permit safe, non-re-identifying analysis, rather than reporting only a course mean — while respecting the small-class anonymity limits described in our k-anonymity guidance.
- Mid-cycle collection. Because belonging is a leading indicator, Koji supports formative, mid-semester collection so a belonging gap can be acted on within the same term, not diagnosed after the students who felt excluded have already left.
- Thematic analysis of open text. Automatic theming of open responses can quantify how often belonging-related concerns (isolation, not fitting in, feeling behind) appear, turning scattered comments into a trackable signal.
We are deliberate about scope: Koji is designed to measure and surface belonging as course-improvement and student-support evidence, not to eliminate belonging gaps or to score individual staff on them. The same AI-moderated interview engine powers Koji's core research platform at koji.so for product and customer research, where "why did you feel that way?" is the same question in a different domain.
Phrasing belonging items so they measure belonging
Belonging items fail quietly when they drift into satisfaction or into instructor performance. "I enjoyed this course" is satisfaction; "The lecturer was approachable" is an instructor-behaviour item; neither measures belonging. Items that do the work name the relational appraisal: "I felt like I belonged in this course", "I felt accepted as a member of this discipline", "I felt comfortable contributing in this class", "I felt like the kind of person who succeeds in this subject". Keep them at the level of the course community and the discipline, not the individual teacher, so the resulting signal cannot be mistaken for a staff-performance score. Pair the scale items with a single, neutral open-text probe — "What, if anything, made you feel more or less part of this course?" — because the reasons for a low belonging rating (an unwelcoming seminar dynamic, a bruising first assessment, being the only student from one's background) are what a programme can actually act on. Finally, resist stacking six near-duplicate items drawn from belonging, relatedness, and achievement-emotion scales at once: they correlate heavily, inflate survey length, and invite straightlining without adding information.
Related Resources
- Autonomy, Competence, Relatedness: Self-Determination Theory as an Evaluation Lens
- Course Feedback as an Early-Warning System (Tinto)
- Achievement Emotions and Control-Value Theory
- Students as Partners: Co-Designing Course Evaluation
- Universal Design for Learning as an Inclusive-Evaluation Lens
- Reference Bias in Self-Reported Learning
References
- Gopalan, M., & Brady, S. T. (2020). College students' sense of belonging: A national perspective. Educational Researcher, 49(2), 134–137. https://doi.org/10.3102/0013189X19897622
- Pedler, M., Willis, R., & Nieuwoudt, J. E. (2022). A sense of belonging at university: student retention, motivation and enjoyment. Journal of Further and Higher Education, 46(3), 397–408. https://doi.org/10.1080/0309877X.2021.1955844
- Strayhorn, T. L. (2018). College Students' Sense of Belonging: A Key to Educational Success for All Students (2nd ed.). Routledge. https://doi.org/10.4324/9781315297293
- Baumeister, R. F., & Leary, M. R. (1995). The need to belong: Desire for interpersonal attachments as a fundamental human motivation. Psychological Bulletin, 117(3), 497–529. https://doi.org/10.1037/0033-2909.117.3.497
Related articles
Course Feedback as an Early-Warning System: What Retention Research Says About Listening Before Students Leave
Tinto's integration model and the persistence literature (Pascarella & Terenzini 1980; Bean 1980) show dropout is driven by academic and social integration — experiences course evaluation could detect, if it did not arrive after the semester ends. How to redesign feedback timing for retention.
Students as Partners, Not Just Respondents: Co-Designing Course Evaluation
Healey, Flint and Harrington's students-as-partners framework reframes evaluation from something done *to* students into something done *with* them. What partnership adds to course feedback, its limits, and how to build it in.
Is Your Course Accessible to Every Learner? Universal Design for Learning as an Inclusive-Evaluation Lens
Universal Design for Learning asks whether a course is usable by the full range of learners by default, not by retrofitted accommodation. Here is how to evaluate inclusivity with UDL — and why the evidence demands you do it honestly.
Autonomy, Competence, Relatedness: Evaluating a Course Through Self-Determination Theory
Most course evaluations ask whether the teaching was good. Self-determination theory suggests a more predictive question: did the course support students' needs for autonomy, competence and relatedness? Those three needs, meta-analytic evidence shows, drive the motivation that in turn drives learning — and almost no institutional survey measures them.