Course Evaluation Evidence for the UK TEF and QAA Quality Review
A practical buyer''s guide for UK universities: how to turn course and module evaluation into defensible evidence for the OfS Teaching Excellence Framework (TEF), the OfS B-conditions, and the QAA UK Quality Code - mapping each requirement to concrete, accreditation-ready outputs.
Koji Education Team
Product
In one sentence: UK quality review - the OfS Teaching Excellence Framework (TEF), the OfS quality and standards B-conditions, and the QAA UK Quality Code - rewards providers who can show not just that they collect student feedback, but that they act on it and can prove the loop closed; this guide maps each requirement to the concrete evaluation evidence assessors look for, and shows how to produce it without manual re-keying.
This is written for UK QA directors, heads of teaching and learning, academic registrars and institutional-research leads preparing a TEF submission or an internal periodic review.
The UK quality landscape in 2026 (get the framing right)
The English system changed in 2023, and getting the framing right matters for a credible submission:
- The Quality Assurance Agency (QAA) ceased to be the Designated Quality Body for England with effect from 1 April 2023. Since then the Office for Students (OfS) has undertaken quality and standards assessment in England directly.
- The QAA UK Quality Code remains the sector-owned reference framework and is still the basis for quality assurance in Scotland, Wales and Northern Ireland, where QAA continues to operate. Most English providers also continue to use the Quality Code as good practice.
- In England, quality is regulated through the OfS conditions of registration B1-B5 - notably B1 (academic experience), B2 (resources, support and student engagement) and B4 (assessment and awards) - and through the Teaching Excellence Framework (TEF).
So your evaluation evidence needs to serve two masters: the TEF (a periodic, rated exercise) and your ongoing B-condition / Quality Code obligations. Good news: the same well-designed evaluation data feeds both.
What the TEF 2023 framework actually asks for
TEF awards an overall rating plus two underpinning ratings - one for Student Experience and one for Student Outcomes (ratings: Gold, Silver, Bronze, or Requires improvement). The evidence is a deliberate mix of quantitative indicators and qualitative narrative:
- Student Experience indicators draw on five areas of the National Student Survey (NSS): teaching on my course; assessment and feedback; academic support; learning resources; and student voice.
- Student Outcomes indicators draw on continuation, completion and progression (the Graduate Outcomes survey).
- Crucially, numerical indicators contribute no more than half of the evidence. The rest is the provider submission (your written narrative and evidence) plus an optional independent student submission.
That "no more than half" rule is the strategic opening. The provider submission is where you demonstrate educational gain, student engagement and how you use feedback to improve - and a thin, Likert-only evidence base is exactly where weaker submissions fall down. Assessors reward specific, longitudinal, closed-loop evidence.
What the QAA UK Quality Code asks for
The Quality Code''s Monitoring and Evaluation and Student Engagement themes expect providers to:
- apply monitoring and evaluation systematically and consistently, not as a one-off;
- engage students individually and collectively in assuring and enhancing their experience;
- involve students in the analysis and communication of findings, not just data collection; and
- act on what is found and feed it back - the closing-the-loop expectation that also underpins the European Standards and Guidelines (ESG / ENQA evidence is covered here).
Requirement-to-evidence mapping
This is the heart of a credible submission. Map each requirement to a concrete output - and to what a modern evaluation platform such as Koji can produce.
| UK requirement | What assessors want to see | Concrete evaluation output | How Koji produces it |
|---|---|---|---|
| TEF Student Experience - "student voice" | Evidence that students shape their experience and that the provider responds | Documented feedback cycles with student input on themes and follow-up actions | Standardized conversational evaluations that capture specific student voice, auto-themed for the submission narrative |
| TEF provider submission (the >50% qualitative half) | Specific, evidenced claims about educational gain and enhancement | Longitudinal, cohort-level evidence with quotes and themes tied to actions | Automatic thematic analysis across cohorts + exportable quotes and trend lines |
| TEF - "assessment and feedback" / "academic support" | Granular insight into where the experience works and fails | Module-level qualitative findings, not just NSS averages | Conversational probing surfaces why a module scored low and what would help |
| OfS B1 (academic experience) | A high-quality academic experience, evidenced and monitored | Routine, standardized evaluation across all modules | Consistent, bias-aware moderation every term, every module |
| OfS B2 (support and student engagement) | Effective engagement of students with their studies | Demonstrable engagement and response, with reach across cohorts | Higher-quality engagement than a static form; reporting on participation |
| QAA Monitoring & Evaluation | Systematic, consistent, repeatable process | An auditable evaluation cycle with comparable data over time | Standardized instrument + longitudinal cohort reporting |
| QAA / TEF closing the loop | Proof that feedback led to change | Action log: theme - decision - action - outcome, fed back to students | Built-in closing-the-loop action tracking linked to the originating themes |
Why static Likert evidence underperforms in UK review
A TEF panel or QAA reviewer reading "our average satisfaction was 4.1/5" learns almost nothing actionable. The weaknesses of averaging are well documented - see why averaging Likert scores misleads and the problem of non-response bias. Worse, an NSS-style number is already counted in the indicator half of TEF; repeating it in your submission wastes the qualitative half where you could be winning marks.
What strengthens a submission is specific, evidenced, closed-loop narrative: "In 2024-25, conversational evaluation across 18 modules surfaced that 27% of comments on assessment cited unclear marking criteria; the School revised rubrics and ran a calibration workshop; in 2025-26 the same theme fell to 9% and students confirmed the change." That sentence is exactly the form of evidence TEF and the Quality Code reward - and it is far easier to produce when the platform themes comments automatically and tracks the resulting actions.
Building an audit-ready evidence base: a checklist
- Standardize the instrument across modules so data is comparable over time (QAA "systematic and consistent").
- Capture qualitative depth, not just ratings, so you can explain why indicators move.
- Theme automatically so cohort-level patterns are defensible and reproducible, not one analyst''s reading.
- Track actions against themes - the single most-cited gap in weak submissions is the missing "and then we did X."
- Report longitudinally by cohort and programme, so you can show change over the review period.
- Keep data in the UK/EU with a clear processing basis - relevant to both GDPR and your data-protection narrative.
When a different tool may fit better
In the spirit of an honest guide: if your institution''s binding constraint is in-class paper capture or a deeply embedded student-information-system automation you already run at very large scale, an established incumbent such as EvaSys or Explorance Blue may serve that operational need today - see the full course evaluation software comparison. Those tools can supply the quantitative spine. Koji''s advantage is the qualitative, closed-loop evidence that the TEF provider submission and QAA monitoring expectations specifically reward - and many institutions run a conversational layer alongside their existing survey number.
How Koji fits
Koji for Education is built to produce UK-review-ready evidence: standardized, bias-aware conversational evaluations; automatic thematic analysis across cohorts; longitudinal reporting; and closing-the-loop action tracking that links each change back to the feedback that prompted it. The same AI interview engine powers the main Koji platform (koji.so) for customer and user research, so the conversational methodology is proven well beyond higher education. The result is a submission narrative you can evidence line by line, rather than a folder of Likert averages.
Common pitfalls in UK evaluation evidence (and how to avoid them)
Reviewers and TEF panels see the same recurring weaknesses. Avoiding them is often the difference between a Silver and a Gold narrative:
- Reporting the indicator twice. Quoting your NSS satisfaction figure in the provider submission wastes the qualitative half - the panel has already seen it in the indicator data. Use the submission to explain and evidence movement, not to restate it.
- "We collect feedback" with no "so we changed X." Collection is assumed; the closing-the-loop evidence is what scores. Every theme you cite should carry a decision, an action and, ideally, a measured outcome in a later cycle.
- One analyst''s reading of free text. A hand-coded summary of comments is hard to defend as systematic. Automatic, consistent thematic analysis across all transcripts is reproducible and survives scrutiny under the QAA "systematic and consistent" expectation.
- No longitudinal comparison. A single snapshot cannot show enhancement. Comparable, standardized data across the review period lets you show a theme falling - or a new one emerging and being addressed.
- Thin module-level granularity. Programme averages hide where the experience actually breaks. Module-level qualitative findings let you target B1 and B2 evidence precisely.
- Data-protection ambiguity. Be explicit about where student data is processed and on what lawful basis; a clean GDPR narrative removes an easy line of challenge.
A practical sequencing tip: start the conversational evaluation layer on two or three programmes a full review cycle before submission, so by the time you write the narrative you already hold year-on-year closed-loop evidence rather than a single term''s data. That lead time is what turns a plausible claim into a demonstrated one.
Related Resources
- Turning Course Evaluations into NVAO Accreditation Evidence
- Course Evaluation Evidence for AACSB & EQUIS Accreditation
- Turning Student Feedback into ESG / ENQA Accreditation Evidence
- What Open-Text Student Comments Tell You That Likert Scores Cannot
- Best Course Evaluation Software for European Universities (2026)
Next step: See how Koji generates closing-the-loop evidence for your TEF submission and QAA monitoring on a single programme.
Related articles
Turning Course Evaluations into NVAO Accreditation Evidence (Netherlands & Flanders)
A buyer-and-practitioner guide to NVAO accreditation: how course-evaluation evidence maps to the standards, why closing the loop is the hard part, and where an AI-moderated approach helps — with an honest note on when a traditional survey tool suffices.
Turning Student Feedback into ESG / ENQA Accreditation Evidence
A buyer's guide mapping the ESG 2015 internal quality assurance standards to concrete, accreditation-ready evidence you can generate from student feedback — and how AI-moderated evaluation closes the loop.
Course Evaluation Evidence for AACSB & EQUIS Business-School Accreditation
A buyer''s guide for business schools: how to turn student course-evaluation data into AACSB Assurance of Learning and EQUIS quality-assurance evidence, with a requirement-to-output mapping and an honest view of what evaluation tools can and cannot do.
What Can Open-Text Student Comments Tell You That Likert Scores Cannot?
A research-grounded guide to open-ended student comments in course evaluation: what Alhija and Fresko found about who writes them and what they contain, how thematic analysis surfaces issues numbers miss (Stupans et al.), how rare abusive comments actually are (Tucker), and how to turn free text into reliable evidence.