National Student Surveys Are Not Course Evaluations: What NSS and Studiebarometeret Can and Cannot Tell You
National student surveys sit at the wrong level of analysis to diagnose a course. Cheng and Marsh showed that most apparent difference between UK universities is not reliable variance at all — here is how to use national data alongside your own instrument instead of in place of it.
Koji Education Team
Product
In short: National student surveys (the UK NSS, Norway's Studiebarometeret, Eurostudent) are institution- and programme-level accountability instruments, not diagnostic course evaluation. Cheng and Marsh (2010) modelled NSS data properly as three levels — student, course, university — and found that once you account for that nesting, differences between universities are small and largely unreliable, while meaningful variance sits at the course level. The practical implication is that a national survey can tell you where to look and almost never what to change.
The mistake institutions make with national survey data
A national survey lands. It is published, ranked, written about in the press, and arrives on a senior team's agenda with an implied instruction: improve the number. Somebody is then asked to work out which programmes are dragging the score down, and an improvement plan is written against items a national instrument was never designed to diagnose.
The error is one of level of analysis. A national survey is engineered to compare institutions and broad subject areas for public information and regulatory purposes. A course evaluation is engineered to tell a programme director what to change in a particular module next term. These are different measurement problems with different sampling, different item granularity, and different acceptable error. Using one for the other's job produces confident action on noise.
What the research says
Most of the between-university difference is not real difference
Cheng and Marsh (2010) is the foundational analysis here. They argued that the standard single-level treatment of NSS data is inappropriate because students are nested within courses, which are nested within universities. Fitting the correct three-level multilevel model (student → course → university), they found that the variance attributable to the university level is small, and that differences between universities are largely unreliable — far less meaningful than the league-table use of the data implies. Substantial and interpretable variance sits at the course level.
This is the same statistical logic covered in our piece on multilevel models for nested evaluation data, applied to national data: ignoring nesting understates standard errors and manufactures apparent precision at exactly the level where precision is weakest.
What national items actually predict
Langan, Dunleavy and Fielding (2013) modelled NSS ratings for science and engineering subjects, relating the 21 core items to overall satisfaction. Two findings matter for practice. First, subjects showed consistent differences from one another — discipline is a structural feature of the data, not a quality signal, echoing the discipline-bias problem in institutional evaluation. Second, the strongest single predictor of overall satisfaction was that "the course was well designed and running smoothly" — an organisational property — while feedback items, despite being the lowest-scoring domain, were among the weakest predictors of overall satisfaction.
That gap is instructive. The item scoring worst is not necessarily the item to fix first if the aim is the headline number — and the fact that those two can diverge is itself a reason not to manage to the headline number. See Campbell's Law and SET for where that road ends.
The Nordic case: coverage, participation, and what gets published
Norway's Studiebarometeret, administered by NOKUT on behalf of the Ministry of Education and Research, surveys students in second-year bachelor and second-year master programmes across effectively all Norwegian higher education institutions. NOKUT reports around 26,000 respondents in the 2024 round, covering roughly 1,900 study programmes, and publishes results openly at programme level via the studiebarometeret.no portal.
The design is deliberately programme-level, and NOKUT is explicit that background data on the student population is used to check representativeness. Norwegian scholarship has nonetheless raised validity concerns about its use as a quality indicator, noting a participation rate of 49% in 2019 with 17% of questionnaires incomplete. A programme-level instrument with a half-responding sample is a reasonable public-information tool and a poor basis for module-level decisions.
Why it matters for course evaluation in practice
National data is a screening instrument, not a diagnostic one. Treat a low national score the way a clinician treats an abnormal screening result: as a reason to run a specific test, not as a diagnosis. The specific test is your own evaluation instrument, run at course level, with items written for the programme in question.
Do not cascade national items into local surveys. A common and costly response is to adopt the national questionnaire's wording for internal use so the numbers are "comparable". They will not be comparable — different population, different timing, different response rate — and you will have replaced diagnostic items with accountability items. Item-specific, low-inference questions do the local job better; see low-inference teaching-behaviour items.
Respect the level when you report. If the reliable variance is at course level, then reporting that presents institution-level movement as evidence of an intervention's success is misleading. Year-over-year national movement in a single programme with 40 respondents is usually regression to the mean.
Use national data for the thing it is uniquely good at: external reference. No internal instrument tells you whether your assessment-and-feedback scores are unusual for the sector. National data does, and that is genuinely valuable context for a self-evaluation report or an accreditation panel.
A worked example of the level error
A faculty sees its national score for "assessment and feedback" fall four points. An action plan is written: all programmes will introduce a two-week marking turnaround and a standard rubric template.
Three things are wrong with this response. First, the four-point movement may be within the uncertainty of a programme-level estimate built on a partial sample — the national instrument was not designed to resolve movements of that size for a single unit. Second, the domain scoring lowest is not necessarily the domain with the most leverage, as Langan et al. found for feedback items. Third, and most importantly, "assessment and feedback" is a national reporting category, not a diagnosis. It bundles turnaround time, feedback specificity, assessment scheduling, criteria clarity and marker consistency into one number. The action plan picks two of those five, essentially at random, because the national instrument never disaggregated them.
The correct response is cheaper and slower: run a course-level instrument in the affected programmes that asks about each of those five things separately, find which one students are actually describing, and act on that. The national score told you where to point the instrument. It could not tell you what the instrument would find.
Limitations and honest caveats
- Cheng and Marsh analysed early NSS rounds. The NSS has since been revised — the 2023 redesign changed items and removed the overall-satisfaction question in England. The structural argument about levels of analysis survives instrument revision; the specific variance estimates should not be quoted as current.
- "Unreliable" is not "meaningless". The finding is that between-university differences are small relative to their uncertainty, not that no institution ever differs from another. Extreme cases with large samples may be real.
- National surveys differ substantially. NSS (final-year, UK), Studiebarometeret (second-year bachelor and master, Norway) and Eurostudent (social and economic conditions, cross-national) sample different populations for different purposes. Findings about one do not automatically transfer.
- Response rates cut both ways. A 49% national response rate is high by course-evaluation standards. The nonresponse concern is about composition, not volume — see wave analysis for nonresponse bias.
- Publication changes behaviour. Because national results are public and consequential, they are more exposed to gaming and to institutional campaigning than internal instruments are. That is a validity threat internal evaluation does not share to the same degree.
How Koji incorporates this
Koji for Education is built for the level of analysis where the usable variance actually lives — the course and the programme — and is designed to complement national instruments rather than imitate them.
- Programme- and cohort-level structuring means evaluations are collected and reported at the unit where the research says differences are interpretable, with nesting preserved rather than flattened into a single institutional average.
- AI-moderated conversational interviews do the diagnostic work a national item cannot. Where a national survey records that "assessment and feedback" scored poorly, the interview asks which assessment, at what point, and what specifically was missing from the feedback — turning a screening flag into an actionable finding.
- Structured question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) support locally written, item-specific questions rather than borrowed accountability wording.
- Automatic thematic analysis aggregates open text across a programme so that a pattern flagged nationally can be characterised internally at scale, with quality scoring to down-weight low-effort responses.
- Closing-the-loop action tracking records what was changed in response and when — the evidence chain an accreditation panel asks for and that a national score, by itself, cannot supply. Our guides to ESG/ENQA evidence and the self-evaluation report cover how this is presented.
Koji is designed to mitigate the diagnostic gap in national survey data. It does not replicate national benchmarking, and institutions should continue to use official national data for external reference. Teams running customer or product research alongside this work use the same AI-moderated interview engine at koji.so.
Related resources
- Students Are Nested in Courses: Why Multilevel Models Beat Raw Averages
- Why "Assessment and Feedback" Consistently Scores Lowest in Student Surveys
- Engagement Surveys vs Course Evaluations: What NSSE Measures
- Course Evaluation Evidence for Nordic Accreditation (NOKUT, UKÄ & FINEEC)
- When a Course Evaluation Becomes a Target, It Stops Being a Measure
- Course Evaluation Evidence for the UK TEF and QAA Quality Review
References
- Cheng, J. H. S., & Marsh, H. W. (2010). National Student Survey: Are differences between universities and courses reliable and meaningful? Oxford Review of Education, 36(6), 693–712. https://doi.org/10.1080/03054985.2010.491179
- Langan, A. M., Dunleavy, P., & Fielding, A. (2013). Applying models to national surveys of undergraduate science students: What affects ratings of satisfaction? Education Sciences, 3(2), 193–207. https://doi.org/10.3390/educsci3020193
- NOKUT. Studiebarometeret – higher education (national student survey portal and reports). https://www.nokut.no/en/publications/studiebarometeret--higher-education/ and https://studiebarometeret.no/en/
- Office for Students. The National Student Survey: Consistency, controversy and change. https://www.officeforstudents.org.uk/publications/the-national-student-survey-consistency-controversy-and-change/
Related articles
Course Evaluation Evidence for the UK TEF and QAA Quality Review
A practical buyer''s guide for UK universities: how to turn course and module evaluation into defensible evidence for the OfS Teaching Excellence Framework (TEF), the OfS B-conditions, and the QAA UK Quality Code - mapping each requirement to concrete, accreditation-ready outputs.
Course Evaluation Evidence for Nordic Accreditation (NOKUT, UKÄ & FINEEC)
How to turn student course evaluations into accreditation-ready evidence for the Norwegian (NOKUT), Swedish (UKÄ) and Finnish (FINEEC/Karvi) quality-assurance systems — with each requirement mapped to concrete outputs.
Students Are Nested in Courses: Why Multilevel Models Beat Raw Averages for Course Evaluation
Course-evaluation data has a nested structure — students within sections within instructors — and a flat average ignores it. This guide explains how multilevel (hierarchical linear) models partition variance, why ignoring clustering understates uncertainty, and what it means for fair reporting.
Why "Assessment and Feedback" Consistently Scores Lowest in Student Surveys
A research-grounded explanation of why assessment and feedback is the perennial weak spot in the UK NSS and comparable instruments, and what a low score actually tells course-evaluation and QA teams.