Why "Assessment and Feedback" Consistently Scores Lowest in Student Surveys
A research-grounded explanation of why assessment and feedback is the perennial weak spot in the UK NSS and comparable instruments, and what a low score actually tells course-evaluation and QA teams.
Koji Education Team
Product
In brief. Assessment and feedback has been the lowest- or near-lowest-scoring domain in the UK National Student Survey (NSS) almost every year since 2005, and comparable instruments show the same tilt. The evidence attributes this less to the quality of the comments academics write and more to timeliness, students not recognising or acting on feedback (the "feedback gap"), and mismatched expectations. For course-evaluation and QA teams, a low score is best read as a prompt to diagnose why — not as a settled verdict on teaching quality.
The pattern is real, robust, and old
Few findings in higher-education survey research are as stable as this one. Since the NSS launched in 2005, the assessment and feedback section has, in the words of one review, seen "students least satisfied with their assessment and feedback experience than with any other aspect of their course" — year after year, across almost all discipline areas and almost all institutions. This is not a one-off dip or an artefact of a single cohort; it is a structural feature of the data.
The picture requires one important qualification, because the NSS questionnaire was substantially redesigned for 2023, breaking a time series that had run since 2017. Under the older questionnaire, "assessment and feedback" was consistently the weakest block. Under the new questionnaire, a closely related theme — student voice, which asks whether students can see their feedback being acted upon — now often scores lowest overall, while assessment and feedback persists as a "relative weakness." In NSS 2024, the single lowest-scoring item out of the 28 in the survey, by some distance, was "how clear is it that students' feedback on the course is acted on?" at just 63.3 per cent positive. The assessment and feedback items themselves sat only modestly higher: 72.7 per cent of students agreed that feedback had helped to improve their work, and 76 per cent felt marking criteria were clear.
So the honest headline is this: whether you look at the old questionnaire or the new one, the cluster of questions about feedback and whether it is used is where UK students are least satisfied. The wording has moved; the weak spot has not. That stability is exactly what makes the phenomenon worth explaining rather than explaining away.
What the research says
Three strands of scholarship converge on the same diagnosis, and none of them locate the primary problem in lazy or absent marking.
Feedback as a process, not a transmission (Nicol & Macfarlane-Dick, 2006). In one of the most-cited papers in the field, Nicol and Macfarlane-Dick reconceptualise formative assessment around self-regulated learning and set out seven principles of good feedback practice. Their central move is to stop treating feedback as something delivered to students and start treating it as something that must feed forward into the learner's own regulation of their work. On this account, comments that are technically accurate but arrive too late, are too generic to act on, or are not connected to a future task simply cannot function as feedback in the strong sense — the loop never closes. This reframing predicts precisely the survey pattern we observe: students discount feedback they cannot use, however diligently it was written.
The feedback gap and feedback literacy (Carless & Boud, 2018). Carless and Boud sharpen the argument by introducing student feedback literacy — "the understandings, capacities and dispositions needed to make sense of information and use it to enhance work or learning strategies." They propose four inter-related features: appreciating feedback, making judgements, managing affect, and taking action. The implication is uncomfortable for the "just write better comments" reflex: even high-quality feedback fails if students lack the literacy to appreciate it, judge their own work against it, manage the emotional sting of criticism, and convert it into action. Much of the dissatisfaction captured by a survey item, on this view, is a recognition and uptake problem, not a provision problem.
Designing the whole process (Winstone & Carless, 2019). In Designing Effective Feedback Processes in Higher Education: A Learning-Focused Approach, Winstone and Carless extend feedback literacy into a design agenda. They deliberately shift emphasis away from what teachers do when they comment and towards how students generate, make sense of, and use feedback for ongoing improvement — through proactive feedback-seeking, peer feedback, and technology-enabled processes. The corollary for evaluation is that a feedback score is a property of a system (curriculum sequencing, turnaround times, whether students are taught to use feedback), not merely of individual markers.
Running underneath all three is the dialogic critique: feedback as monologue (a document handed back) reliably underperforms feedback as dialogue (an exchange in which students respond, question, and demonstrate uptake). And qualitative work confirms the texture of the complaint. Harkin and colleagues (2022) analysed students' written NSS responses on assessment and feedback — the free-text most institutions ignore in favour of the Likert numbers — and found the dissatisfaction clustering around concrete, recognisable issues: feedback that was late, too vague to act on, inconsistent between markers, or disconnected from the grade. These are not diffuse grumbles; they are specific, fixable process failures.
Why it matters in practice
The practical question for a dean, programme director, or QA officer is deceptively simple: when the assessment-and-feedback score comes in low again, what does it actually mean? The research supports a three-part reading.
First, part of the signal is genuine and teaching-related: real problems with turnaround time, clarity of criteria, and the actionability of comments. These are the drivers Harkin et al. surface in the free text, and they are addressable through course design — earlier and more frequent low-stakes feedback, exemplars and rubrics that make standards visible, and explicit "feed-forward" links from one assignment to the next.
Second, part of the signal is a feedback-gap effect — students receiving perfectly good feedback but not appreciating, recognising, or using it. This is where feedback literacy interventions matter, and where simply working markers harder produces no movement in the score.
Third, part is measurement and expectation (developed below). The danger for QA teams is collapsing all three into "our marking is bad" and either over-correcting or, worse, gaming the item. The UCL case study widely cited in the sector — a 26 per cent rise in NSS feedback and assessment scores over three years — is instructive precisely because the gains came from redesigning feedback processes and student communication, not from a single tactical fix.
The actionable stance is therefore diagnostic. A low score should trigger a question — why is it low? — decomposed into timeliness, clarity, actionability, and closure. A number alone cannot tell you which lever to pull; it tells you a lever needs pulling.
Limitations and honest caveats
Rigour requires naming what the score cannot do.
- It measures perception, not quality. The item captures how students experienced and recognised feedback, not the intrinsic pedagogical quality of the comments. Two modules with identical feedback practices can score differently if one cohort was better prepared to use it.
- Construct and wording sensitivity. The 2023 NSS redesign changed item wording and split feedback-adjacent content across "assessment and feedback" and "student voice," which is why the nominal lowest theme shifted without any real change in student experience. Small wording changes move scores; this is a property of the instrument, not of teaching.
- Reference and expectation bias. Students rate against their own expectations. Higher-tariff institutions have historically seen lower satisfaction with feedback partly because expectations are higher — a documented "elite penalty" in the NSS. A low score can reflect ambition and standards, not failure.
- The single-number trap. A domain average hides the distribution. The same mean can come from uniformly mediocre feedback or from excellent feedback on some modules and late feedback on a few, which require opposite responses.
- Generalisability. The specific percentages are UK-NSS-specific. The pattern recurs in the Course Experience Questionnaire and Australian student-experience surveys, and the feedback-literacy literature is international, so the explanations travel further than the figures do. Treat cross-instrument comparisons of raw numbers with caution.
None of these caveats dissolve the finding. They discipline its interpretation: the low score is real and worth acting on, but it is a compound of teaching signal, measurement artefact, and expectation effect that must be separated before it drives decisions.
How Koji incorporates this
Koji for Education is designed to attack the diagnostic problem the NSS exposes — turning a low aggregate number into an actionable understanding of why — while being honest that no instrument can eliminate perception effects or measurement bias.
- AI-moderated conversational interviews that probe the "why." When a student rates feedback low, Koji's interview engine can follow up in natural language to disentangle the drivers the literature identifies: was the problem timeliness, clarity of criteria, actionability, or whether they acted on it and saw the loop close? This is designed to mitigate — not remove — the single-number trap by capturing the mechanism behind the rating rather than only the rating.
- Structured question types for low-inference items. Alongside open text, Koji supports
open_ended,scale,single_choice,multiple_choice,ranking, andyes_noquestions, which lets evaluation designers write behaviourally specific feedback items ("Did you receive feedback before your next assignment was due?") rather than the high-inference, easily-misread wording that makes survey scores volatile. - Automatic thematic analysis of open text. Following the logic of Harkin et al. (2022), Koji analyses free-text responses at scale, so the qualitative signal most institutions leave unread becomes usable evidence about what "assessment and feedback" actually means to a given cohort.
- Mid-cycle and formative collection. Because feedback problems are most fixable while a module is running, Koji supports mid-cycle collection so a timeliness or clarity issue surfaces in time to correct it — rather than arriving as a post-hoc NSS number a year later.
- Closing-the-loop action tracking. Given that the lowest single NSS item concerns whether students see their feedback acted on, Koji supports recording and communicating the actions taken in response — directly targeting the student-voice weak spot rather than only measuring it.
The same AI-moderated interview engine underpins Koji's core research platform at koji.so, where it is applied to product and customer research; the education product points that engine at course evaluation and the feedback loop specifically.
Used well, these mechanisms are designed to convert an inert, perennially low score into a diagnosis a programme team can act on within a cycle — with the caveat, consistent with the caveats above, that they surface and mitigate the drivers of dissatisfaction rather than promising to eliminate them.
References
- Carless, D., & Boud, D. (2018). The development of student feedback literacy: enabling uptake of feedback. Assessment & Evaluation in Higher Education, 43(8), 1315–1325. https://doi.org/10.1080/02602938.2018.1463354
- Harkin, B., Paltoglou, A. E., Tariq, K., Watkin, M., Ashfaq, S., Yates, A., & Jacobs, C. (2022). Student experiences of assessment and feedback in the National Student Survey: an analysis of student written responses with pedagogical implications. International Journal of Management and Applied Research, 9(2), 115–139. https://doi.org/10.18646/2056.92.22-006
- Nicol, D. J., & Macfarlane-Dick, D. (2006). Formative assessment and self-regulated learning: a model and seven principles of good feedback practice. Studies in Higher Education, 31(2), 199–218. https://doi.org/10.1080/03075070600572090
- Office for Students. (2024). National Student Survey results. Retrieved from https://www.officeforstudents.org.uk/for-students/understanding-students/national-student-survey/latest-nss-results/
- Winstone, N., & Carless, D. (2019). Designing effective feedback processes in higher education: a learning-focused approach. Routledge. https://doi.org/10.4324/9781351115940
Related Resources
- Closing the feedback loop: course-evaluation evidence
- Do student evaluations improve teaching? The feedback-intervention evidence
- The Course Experience Questionnaire (CEQ): perceived teaching quality and learning
- Student feedback as ESG accreditation evidence
- Course-evaluation evidence for UK TEF and QAA
- Open-ended course-evaluation question design: specific prompts
Related articles
Turning Student Feedback into ESG / ENQA Accreditation Evidence
A buyer's guide mapping the ESG 2015 internal quality assurance standards to concrete, accreditation-ready evidence you can generate from student feedback — and how AI-moderated evaluation closes the loop.
Course Evaluation Evidence for the UK TEF and QAA Quality Review
A practical buyer''s guide for UK universities: how to turn course and module evaluation into defensible evidence for the OfS Teaching Excellence Framework (TEF), the OfS B-conditions, and the QAA UK Quality Code - mapping each requirement to concrete, accreditation-ready outputs.
Does Closing the Feedback Loop Actually Matter? The Evidence on Acting on Student Evaluations
Universities are good at collecting student feedback and bad at acting on it visibly. The research — Watson (2003), Leckey & Neill (2001), Shah et al. (2017) — shows that failing to close the loop drives the scepticism and declining response rates that quietly destroy your evaluation data.
Do Student Evaluations Actually Improve Teaching? The Feedback-Intervention Evidence
Student ratings handed back as a bare number barely change teaching (d ≈ 0.2). Paired with structured consultation, the same data produces moderate, durable improvement (d ≈ 0.6–0.7). What the intervention meta-analyses tell quality teams about closing the loop.