New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Sector trends9 min read

What the TEF Actually Rewards — and Why Course-Level Evidence Is the Missing Piece

England's Teaching Excellence Framework rates whole providers Gold, Silver or Bronze — but it runs on institution-level proxies. The evidence that actually explains a rating lives at course level, and most providers cannot produce it.

Koji Education Team

Product · August 12, 2026

Bottom line up front: The Teaching Excellence Framework (TEF) awards a single Gold, Silver or Bronze rating to an entire English provider, built largely from National Student Survey (NSS) responses and administrative outcome data. Those are institution-level proxies. The thing a TEF panel is actually trying to infer — whether teaching is excellent, and why — is a course-level, mechanism-level question that provider-wide averages cannot answer. The providers that fare best in the qualitative "provider submission" are the ones that can show granular, credible evidence of what works in specific courses and how they act on it. That evidence is exactly what most course-evaluation systems fail to capture.

What the TEF measures, precisely

The TEF is run by the Office for Students (OfS) and, since the reforms following Dame Shirley Pearce's independent review, produces an overall rating plus two "aspect" ratings — one for student experience and one for student outcomes. Ratings run Gold, Silver, Bronze, and "Requires improvement." In the TEF 2023 exercise, 228 providers took part; 37 received Gold, 79 Silver, and 11 Bronze, with the remainder split across other categories and some initially pending (Advance HE summary of the TEF 2023 outcomes).

The student-experience aspect draws heavily on the NSS — students' aggregated agreement with statements about teaching, assessment and feedback, academic support and organisation. The student-outcomes aspect uses continuation, completion and progression data. On top of the numbers sit two narrative submissions: one from the provider and, importantly, one from the provider's students. The panel weighs the indicators against the narrative.

Read that structure carefully and the design tension is obvious. The quantitative core is satisfaction and administrative outcomes; the qualitative core is narrative and advocacy. Neither layer contains rigorous, comparable evidence about teaching itself at the level where teaching happens — the module, the seminar, the lab.

Why institution-level proxies mislead

Two decades of research warn against treating student satisfaction as a proxy for teaching quality. The largest meta-analysis of the literature, Uttl, White and Gonzalez (2017), re-analysed decades of "multisection" studies and found that once sample sizes and prior ability are properly accounted for, the correlation between student ratings and actual learning is effectively zero. If satisfaction barely tracks learning at the course level, aggregating it to the provider level does not fix the problem — it launders it. A Gold provider can contain mediocre modules; a Bronze provider can contain genuinely excellent ones. The rating compresses a distribution into a point, and the distribution is where the truth lives.

This is the same critique the sector has aimed at the NSS itself and at league tables. We have covered why averaging Likert scores misleads, why the NSS reform debate matters for anyone who relies on satisfaction data, and how rankings launder satisfaction into apparent quality. The TEF inherits every one of those weaknesses and then adds a high-stakes, publicly branded rating on top.

It also sits directly alongside the OfS's regulatory conditions. Condition B1 asks for a high-quality academic experience — not high satisfaction, and Condition B3 treats student outcomes as the destination, with course-level signals as leading indicators. The TEF is the reputational shop-window for the same underlying question those conditions probe.

Where ratings are actually won: the submission

Here is the part institutions under-invest in. Because the indicators are largely fixed by the NSS and outcome data, the lever a provider can still pull is the narrative submission — the case it makes, with evidence, that its educational gains are excellent and that it understands why. Panels reward specificity: a claim that "we redesigned first-year assessment and continuation improved" is only persuasive if the provider can show the mechanism, the student voice behind it, and the follow-through.

That is a course-level evidence problem. A provider that can point to structured, longitudinal, course-level feedback — showing what students actually experienced in a redesigned module, in their own words, and what changed as a result — has a submission with tendons. A provider whose only evidence is a 4.1/5 average on a legacy end-of-term form has assertions.

This is where the action gap in the quality cycle becomes a TEF liability, not just a governance one. "You said, we did" is a narrative the TEF explicitly rewards — and one you cannot write convincingly if your evaluation data never captured what students meant in the first place.

But doesn't the TEF already include qualitative evidence?

A fair objection: the TEF is not purely metric-driven. It has a student submission and a provider narrative, and the Pearce review deliberately strengthened the qualitative side. Doesn't that already fix the "averages hide teaching" problem?

Partly, and it is a genuine improvement. But there is a difference between qualitative advocacy and qualitative evidence. A student submission is, by design, a representative-body's account — valuable, but not a systematic sample of the student population's experience across courses. A provider narrative is a self-advocacy document. Neither is a rigorous, low-bias reading of what students across many modules actually experienced. Panels know this, which is why credible, well-sourced course-level evidence in a provider submission is disproportionately persuasive: it is the scarce ingredient. The TEF creates demand for exactly the evidence most institutions are not equipped to supply.

A second objection: won't better course-level data just be gamed, becoming another target? That risk is real — it is Goodhart's Law, and it applies to any metric attached to stakes. The mitigation is not to collect less evidence but to collect evidence that is harder to game: open-ended, probed, thematically analysed feedback resists a single number's gravity in a way a headline Likert mean never can. It also aligns with responsible-metrics thinking — quantitative indicators supported, not replaced, by expert and student judgement.

A four-year rating raises the stakes on evidence

The TEF is not an annual event. Ratings persist for several years, which means a single exercise fixes how your provider is publicly described — to applicants, to partners, to league-table compilers — for a long time. That longevity cuts both ways. It rewards institutions that have been quietly building a credible, course-level evidence base all along, because a strong submission cannot be assembled in the weeks before a deadline from data you never collected. And it penalises those who treat evaluation as an end-of-term compliance ritual, because when the submission window opens they discover their entire evidence base is a folder of Likert averages that says what students rated but nothing about why — and why is what a panel is reading for. The institutions that will do well in the next cycle are the ones treating the intervening years as the time to gather the mechanism-level evidence, not the submission fortnight.

What Koji contributes to a TEF evidence base

Koji for Education does not rate your provider — the TEF panel does. What Koji does is help you build the course-level evidence a strong submission needs, without adding survey fatigue.

  • AI-moderated conversational interviews replace the static end-of-term form. Instead of a 4.1 average, you get why a cohort rated assessment poorly, probed follow-up by follow-up, at a scale surveys cannot reach. That is the raw material of a "you said, we did" narrative.
  • Automatic thematic analysis of open-text feedback turns thousands of comments into named, quantified themes across modules and programmes — the comparable, course-level picture a provider submission can cite with confidence, rather than cherry-picked quotes.
  • Programme- and institution-level reporting lets you show the distribution beneath a provider average: which courses drive an outcome, where a redesign landed, where it did not.
  • Formative, mid-cycle collection and closing-the-loop action tracking document the improvement cycle itself — the evidence that you understood a problem and acted, which is what the outcomes aspect ultimately rewards.
  • GDPR-compliant, EU/UK-appropriate data handling keeps this defensible under the OfS's own regulatory expectations.

To be precise about the claim: better evidence does not earn a rating on its own, and no tool can promise a Gold. What it does is let your genuine teaching improvements be seen and believed by a panel that is actively looking for mechanism, not just metrics. Legacy tools like EvaSys or generic survey platforms give you the average; the point of the exercise is everything the average hides.

Many teaching-and-learning teams that run course evaluation this way also run broader user and staff research on the shared AI interview engine behind koji.so — the same conversational method, pointed at different questions.

If your next TEF submission will lean on course-level evidence you do not yet have, it is worth seeing how Koji for Education captures it.

Frequently asked questions

What is the Teaching Excellence Framework (TEF)?

The TEF is a national scheme run by England's Office for Students that awards providers an overall rating of Gold, Silver, Bronze or Requires Improvement, plus separate ratings for student experience and student outcomes. It combines National Student Survey data and administrative outcome metrics with narrative submissions from the provider and its students. In TEF 2023, 228 providers took part, with 37 rated Gold, 79 Silver and 11 Bronze.

Does a high TEF rating mean every course is excellent?

No. The rating is a provider-level summary built largely from aggregated proxies. A Gold provider can contain weak modules and a Bronze provider can contain excellent ones, because averaging compresses the distribution of course quality into a single point. Course-level evidence is what reveals where teaching is actually strong or weak.

Why is course-level evidence important for a TEF submission?

Because the quantitative indicators are largely fixed by the NSS and outcome data, the lever a provider controls is the narrative submission. Panels reward specific, credible evidence of what works in particular courses and how the provider acted on it. Structured, course-level feedback is the raw material for a persuasive "you said, we did" case; a single Likert average is not.

Isn't student satisfaction a good enough measure of teaching quality?

The evidence says no. Meta-analytic work by Uttl and colleagues (2017) found that once sample size and prior ability are accounted for, the correlation between student ratings and actual learning is effectively zero. Satisfaction is worth measuring, but it is a weak proxy for teaching quality, and aggregating it to provider level does not fix that.

How does Koji help with TEF evidence without adding survey fatigue?

Koji replaces static end-of-term forms with AI-moderated conversational interviews that probe why students rated something as they did, then applies automatic thematic analysis across modules and programmes. That produces comparable, course-level evidence and documented action-tracking — the mechanism-level material a strong provider submission needs — while collecting it in a single conversational instrument rather than repeated surveys.