New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Sector trends10 min read

The Regulator Doesn't Grade You on Satisfaction Anymore. It Grades You on Whether Students Finish.

The Office for Students judges quality on continuation, completion and progression — hard outcome metrics, not satisfaction means. But those numbers are lagging by years. Course evaluation's real job now is to be the leading indicator that lets you intervene before a cohort trips a B3 threshold.

Koji Education Team

Product · August 10, 2026

The Office for Students no longer judges the quality of your provision on how satisfied students say they are. Under condition B3, it judges you on whether students continue, complete, and progress — hard outcome metrics with published numerical thresholds you can be sanctioned for missing. If your quality dashboard is still a five-point satisfaction average, you are tracking something the regulator has explicitly moved away from as the test. And there is a deeper problem: B3 metrics are lagging indicators, reflecting cohorts that entered years ago. By the time the data shows a threshold breach, those students have already left. The real job of course evaluation in a B3 world is to be the leading indicator — the in-course signal that lets you act while the cohort is still here.

What B3 measures — and what it doesn't

The OfS groups its quality-and-standards requirements as the "B" conditions, B1 to B5. B3 is student outcomes, and its core requirement is blunt: the provider "must deliver positive outcomes for students on its higher education courses" (condition B3.3). The condition sits on section 5 of the Higher Education and Research Act 2017, the revised version came into force on 3 October 2022, and the numerical thresholds behind it were published on 6 July 2023.

The condition defines outcomes through three indicators:

  • Continuation — students continuing their course from year to year.
  • Completion — students completing their course.
  • Progression — students progressing into managerial or professional employment, or further study.

Notice what is not on that list: student satisfaction. B3's indicators are outcome metrics. A glowing National Student Survey score does not help you meet B3, and a mediocre one does not breach it — the two live in different worlds, a gap we have explored in the lessons from NSS reform. The regulator has, in effect, decided that whether students finish and get somewhere is the test, not whether they enjoyed the module.

Those thresholds have teeth. For a full-time first degree, the OfS set minimum numerical baselines of roughly 80% continuation, 75% completion, and 60% progression, varying by mode and level of study. Being below a threshold is not an automatic breach — under B3.4 the OfS also weighs a provider's context, benchmarks, and credible improvement plans. But where the numbers are low and context does not justify them, the OfS can act: under section 15 of HERA 2017 it can impose a monetary penalty, alongside other sanctions up to deregistration, with providers given at least 28 days to make representations and a right of appeal to the First-tier Tribunal.

The lag problem — why the number tells you too late

Here is the trap. Every B3 indicator is inherently retrospective. Continuation and completion are measured across a cohort's journey through the degree. Progression is measured by the Graduate Outcomes survey, run by HESA (part of Jisc), which surveys graduates approximately 15 months after they complete their studies. The data feeding your B3 assessment is drawn from the designated data body and official sources — which means it reports on students who entered years before the number lands on your desk.

So B3, for all its rigour, is a rear-view mirror. When a completion rate finally dips below threshold, the cohort that produced the dip has already graduated or dropped out. You cannot retain a student who left eighteen months ago. The regulatory metric tells you, with authority and a lag, about a problem you can no longer fix for the people it happened to.

This is exactly the gap a well-designed course evaluation is positioned to fill — if you stop asking it to measure satisfaction and start asking it to measure the drivers of the outcomes B3 cares about.

Evaluation as the leading indicator

Continuation and completion do not collapse for mysterious reasons. Students leave, or fail to progress, because of things that are visible in-course months before they show up in HESA data: a sense of not belonging, assessment bunched into an impossible fortnight, feedback too thin or too late to act on, a support service they could not navigate, a module that assumed knowledge they did not have. Every one of those is a leading indicator of a lagging B3 metric — and every one is surfaceable through evaluation, if the evaluation asks about it.

That "if" is the whole point. A generic "rate this module 1–5" tells you almost nothing about retention risk, and because averaging Likert responses hides the distribution, a comfortable mean can sit on top of a cluster of students who are quietly about to leave. To function as a leading indicator, evaluation has to move beyond satisfaction to the mechanisms — belonging, workload distribution, clarity of expectations, usefulness of feedback — and it has to identify which students are at risk and why, not just report a cohort average. Done that way, evaluation becomes an early-warning system aligned to the exact outcomes the regulator now measures, and it complements the longer-run graduate-tracer evidence that feeds the progression picture.

But doesn't satisfaction predict retention?

The strongest objection is that satisfaction and retention are correlated, so a satisfaction survey already is a leading indicator. There is something to this — deeply dissatisfied students are more likely to leave — but it is a weak foundation for three reasons. First, the correlation is inconsistent and confounded: plenty of satisfied students leave for financial or personal reasons, and plenty of dissatisfied ones stay. Second, a satisfaction mean does not tell you who is at risk or why, which is the only information that lets you intervene. Third, satisfaction ratings carry their own biases — leniency, halo, the tendency for the disengaged simply not to respond — so the students most at risk of dropping out are often precisely the ones missing from your satisfaction data. Predicting retention from a satisfaction average is predicting the outcome you care about from a noisy, biased proxy of a loosely related construct. The fix is not a better average; it is diagnostic evaluation that maps to the drivers of continuation and completion. And to be clear about the boundary: B3 is a regulatory floor, distinct from the Teaching Excellence Framework's separate Gold/Silver/Bronze recognition — evaluation is not a route to gaming either, but a route to genuinely improving the underlying student experience.

Where Koji fits

Koji for Education is built to be that leading indicator. Its AI-moderated conversational interviews probe the drivers of continuation — does the student feel they belong, is the workload survivable, is feedback useful, do they know where to get help — and follow up in the student's own words rather than stopping at a number. Formative, mid-cycle collection means the signal arrives while you can still act on it, not in a post-hoc survey after the at-risk students have gone. Automatic thematic analysis turns open text into ranked, cohort-level themes so a programme lead can see a retention risk building — assessment bunching, a support gap — and programme- and institution-level reporting lets you connect those in-course signals to the B3 indicators they precede. Because Koji tracks closing-the-loop actions, it also documents the intervention trail that the OfS's contextual assessment under B3.4 rewards. Institutions that run wider student-success and retention research use the same conversational engine on the main Koji platform.

Koji does not change your B3 numbers directly, and it would be dishonest to claim it retains students on its own. What it does is give you the early, specific, actionable signal that a lagging regulatory metric cannot — in time to act for the cohort that is still in front of you.

Frequently asked questions

What does OfS condition B3 measure? Condition B3 requires providers to deliver positive student outcomes, measured through three indicators: continuation (students staying on course year to year), completion (students finishing), and progression (into managerial or professional employment or further study). Student satisfaction is not one of the B3 indicators.

What are the B3 numerical thresholds? The OfS published numerical thresholds in July 2023, set separately by mode and level of study. For a full-time first degree the minimum baselines are approximately 80% continuation, 75% completion and 60% progression. Being below a threshold is not an automatic breach — the OfS also considers the provider's context and improvement plans under B3.4.

Can the OfS penalise a provider under B3? Yes. Where outcomes are below threshold and context does not justify it, the OfS can act under the Higher Education and Research Act 2017, including imposing a monetary penalty under section 15, with providers given at least 28 days to make representations and a right of appeal to the First-tier Tribunal.

Why is course evaluation relevant to B3 if B3 isn't about satisfaction? Because B3 metrics are lagging — they report on cohorts that entered years earlier, and progression is measured about 15 months after graduation. Course evaluation can be the leading indicator that surfaces the drivers of continuation and completion (belonging, workload, feedback, support) while students are still enrolled and problems can still be fixed.

Isn't a satisfaction survey already a leading indicator of retention? Only weakly. Satisfaction and retention are loosely and inconsistently correlated, a satisfaction mean does not tell you which students are at risk or why, and satisfaction ratings are biased in ways that often exclude the most at-risk students. Diagnostic evaluation of retention drivers is a far stronger early-warning signal than a satisfaction average.

How is B3 different from the TEF? B3 is a mandatory baseline condition of registration that every provider must meet to stay registered. The Teaching Excellence Framework is a separate, opt-in exercise that awards Gold, Silver or Bronze recognition for excellence above the baseline. They are distinct regimes.


Want your evaluation to predict a B3 problem instead of confirming it too late? See how Koji for Education surfaces the drivers of continuation and completion while your students are still here to help.