New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Methodology8 min read

Why "Assessment and Feedback" Always Scores Lowest—and What Your Survey Can't See

Year after year, assessment and feedback is the weakest theme in national student surveys. A low score is reliable; what it means is not. Here is why a Likert average can't diagnose the problem, and what evaluation has to do instead.

Koji Education Team

Product ·

Bottom line up front: Across the UK's National Student Survey, assessment and feedback is one of the most persistently low-scoring areas—in the 2024 results, just 72.7% of students agreed that feedback had helped improve their work, a figure broadly unchanged from the year before (Advance HE, NSS 2024). The score is reliable and the pattern is real. But the number itself is diagnostically empty: a 72.7% does not tell you whether feedback was late, vague, too harsh, never read, or perfectly good but invisible to students who didn't recognise it as feedback. Each of those has a different fix. If your evaluation produces the score but not the reason, you have measured a problem you cannot act on.

A remarkably stable weak spot

Anyone who works in European quality assurance knows the shape of this finding. In national instruments such as the UK's NSS, the themes covering teaching, academic support, and learning resources tend to score well, while the questions about assessment, feedback, and student voice trail behind year after year (Office for Students, latest NSS results). In 2024, the headline assessment-and-feedback items sat around the low-to-mid 70s—72.7% agreeing feedback helped them improve, roughly 76% finding the marking criteria clear—and these moved only marginally from 2023.

The stability is itself informative. When a theme scores low everywhere, every year, across institutions with wildly different practices, the explanation is unlikely to be that every assessment regime in the country is simultaneously failing. Something more structural is going on—and a single satisfaction percentage is exactly the wrong tool to uncover it.

The five problems a low score could mean

"Assessment and feedback: 72%" is a symptom with at least five distinct underlying causes, and they call for opposite responses:

  1. Feedback is late. Returned after the next assignment is already submitted, so it cannot inform anything. Fix: turnaround deadlines.
  2. Feedback is vague. "Good effort, could be clearer" tells a student nothing. Fix: feedback literacy for markers, exemplars, rubrics.
  3. Feedback is not perceived as feedback. A substantial body of work shows students often do not recognise dialogue, in-class comments, model answers, or peer review as feedback, and so under-report receiving it. Fix: framing and signposting, not more feedback.
  4. Assessment feels misaligned. Students sense the assessment did not match what was taught or what they were told mattered. Fix: constructive alignment.
  5. Feedback is good but emotionally hard to receive. Critical feedback on work students cared about lands badly, and the rating reflects the sting, not the quality. Fix: how feedback is delivered, not whether.

A Likert average cannot distinguish any of these from the others. Two programmes can both score 72% for completely opposite reasons—one because feedback arrives six weeks late, the other because excellent feedback is never recognised as such. Give them the same intervention and you will help one and annoy the other.

Why this theme in particular defeats numeric scales

Assessment and feedback is uniquely badly served by Likert scoring for three reasons.

First, it is multi-dimensional masquerading as one thing. Timeliness, clarity, usefulness, fairness, and tone are separate constructs that students fuse into a single grumpy number. Averaging that number, as we have argued more generally, destroys the only information you needed.

Second, it is emotionally loaded. Feedback is the moment a student's effort meets judgement. Expectation-disconfirmation effects are strong here: a student who expected a 70 and received a 58 will rate the feedback poorly regardless of how constructive it was. The rating measures the gap between hope and grade as much as it measures feedback quality.

Third, it suffers a recognition problem that no rating scale can detect. If students do not count something as feedback, they will mark the feedback question low while sitting on a pile of feedback they did not register. You cannot fix a perception gap you cannot see, and a number cannot show you a perception gap—only language can.

"But isn't the NSS score still useful as a benchmark?"

Yes—and this is the honest counterpoint. A persistently low, stable score is genuinely useful for one thing: telling you where to look. As a smoke detector, the assessment-and-feedback theme works. The mistake is treating the smoke detector as a diagnosis. Knowing that a programme scores 68% versus a sector average of 73% tells you to investigate; it tells you nothing about what you'll find. Benchmarking identifies the where; it cannot supply the why, and chasing the score directly—pressuring markers to be "nicer" or gaming turnaround stats—risks Goodhart's Law: the metric improves while the underlying experience does not.

There is also a real limit to qualitative depth: open comment boxes attached to national surveys are notoriously thin, because students are tired by the time they reach them and have no prompt to elaborate. So the sector ends up with a reliable number and no usable explanation—the worst of both worlds.

What evaluation has to do instead

If the score tells you where and never why, the evaluation's job is to recover the why—at the module level, where action happens, not the institutional level, where the NSS lives. That requires three things a static survey rarely delivers: questions that separate the dimensions (was it late, unclear, or unfair?), follow-up that probes a low rating instead of recording it, and analysis that turns hundreds of comments into named, countable issues.

This is the work Koji for Education is designed for. When a student rates feedback poorly, Koji's AI moderator asks the obvious next question—what would have made it more useful?—and does so consistently for every student, so a flat 72% resolves into "feedback was thorough but arrived after the resit deadline" for one cohort and "students didn't realise the in-seminar comments counted as feedback" for another. Its automatic thematic analysis quantifies which of the five causes actually dominates, and its structured question types let you measure timeliness, clarity, and fairness as distinct items rather than one fused score. Because collection can run mid-cycle, a programme can fix a feedback-timing problem in the current term rather than reading about it in next year's NSS. And because the moderation is standardised and GDPR/AVG-compliant, you get scale without sacrificing consistency or data protection.

Koji will not magic away the structural difficulty of feedback—students will always find judgement hard, and some will always under-recognise feedback. What it does is mitigate the diagnostic blindness of the single score, replacing "assessment and feedback: 72%" with a ranked list of specific, fixable causes tied to the students' own words. The same conversational interview engine runs the main Koji platform for customer and product research, where teams face the identical trap: a satisfaction metric that flags a problem and stays silent on its cause.

The feedback-literacy angle the score hides

There is a further reason the assessment-and-feedback theme resists improvement: feedback is a two-sided act. The influential concept of student feedback literacy—developed by David Carless and David Boud (2018)—describes the capacities students need to make sense of feedback and use it to improve. Where that literacy is low, even excellent feedback lands as noise: students cannot decode it, do not know how to act on it, and rate it poorly. This reframes the low score entirely. Some of what looks like a teaching problem is a feedback-literacy problem, and the intervention is to help students learn how to use feedback, not to ask staff to produce more of it. A numeric rating cannot tell these two cases apart; only asking students what they did with the feedback can. It is one more reason the score names a symptom while staying silent on the cure.

The takeaway

Assessment and feedback scores low everywhere, every year, for reasons that contradict each other. The NSS number is a reliable alarm and a useless diagnosis. Treat the score as a prompt to investigate, not as a finding to act on—and build evaluation that asks the follow-up question the rating scale never can. The programmes that improve fastest are not the ones that chase the percentage; they are the ones that find out, in students' own words, which of the five problems they actually have.