OfS Condition B4 Grades Whether Your Assessment Is Valid and Reliable. Students Are Its Early-Warning System, Not Its Judges.
OfS Condition B4 requires assessment to be valid, reliable, credible and misconduct-resistant, judged by expert design, not student satisfaction. But students are the witnesses to the conditions that make assessment sound. How evaluation becomes a leading indicator of B4 risk.
Koji Education Team
Product · August 16, 2026
OfS Condition B4 Grades Whether Your Assessment Is Valid and Reliable. Students Are Its Early-Warning System, Not Its Judges.
Bottom line up front: The Office for Students' ongoing condition B4, updated on 24 November 2022, requires that students are "assessed effectively," that each assessment is "valid" and "reliable," and that academic regulations keep awards "credible" and holding their value over time — with a specific duty to take "reasonable steps to detect and prevent" academic misconduct, including essay mills. That is an absolute standard, judged by expert design and comparability, not by whether students liked their assessments. But students are the only witnesses to the conditions that make assessment valid or invalid in practice: whether briefs were clear, whether workload was bunched into an unmanageable fortnight, whether feedback arrived in time to use, whether the assessment design left the door open to shortcuts. A satisfaction score cannot answer any B4 question. Well-designed evaluation is the leading indicator that surfaces B4 risk before it becomes an appeal, a complaint, or a breach.
What B4 actually demands
The OfS defines the terms precisely, and the definitions are the whole point. An assessment is "valid" when it "results in students demonstrating knowledge and skills in the way intended by design." It is "reliable" when it "requires students to demonstrate knowledge and skills in a manner which is consistent as between students" and "over time." "Academic misconduct" means "any action or attempted action that may result in a student obtaining an unfair academic advantage," expressly including "plagiarism, unauthorised collaboration" and "use of services offered by an essay mill." Providers must ensure assessments are "designed in a way that minimises the opportunities for academic misconduct and facilitates the detection of such misconduct where it does occur."
Notice what is — and is not — in that list. Nowhere does B4 mention student satisfaction, module ratings, or how enjoyable an assessment was. The regulator is asking whether the assessment tested the right thing, marked it consistently, and resisted cheating. Those are questions of assessment design and integrity, answered by expert judgement, external examiners and moderation — the same "absolute standard, not a popularity contest" logic that runs through the B1 academic-experience condition. A programme could post glowing satisfaction scores and still breach B4 if its marking is inconsistent or its coursework is trivially outsourceable to an essay mill or a generative-AI tool.
Why the satisfaction mean is the wrong instrument — and student voice is still essential
It would be a mistake to conclude that because B4 is not about satisfaction, students have nothing to contribute. They have a great deal — just not as judges of standards. They are witnesses to the conditions that determine whether an assessment is valid and reliable in practice.
Consider validity. An assessment is invalid if students cannot tell what it is asking for. The people who know whether the brief, the rubric and the criteria were intelligible are the students who sat it. Consider reliability. Perceived inconsistency between markers, or between this year's cohort and last year's, often shows up first in student reports long before an external examiner formalises it. Consider misconduct. The pressure that pushes students toward essay mills and AI shortcuts — assessments all falling in the same week, unrealistic timeframes, briefs that reward reproduction over understanding — is felt by students first. This is why "assessment and feedback" is perennially the lowest-scoring domain in national surveys: it is the part of the student experience most sensitive to design flaws, and therefore the richest source of early-warning signal.
The catch is that a five-point mean captures none of this. "Rate your satisfaction with assessment: 3.2" tells a course team nothing about which brief was unclear, where marking felt inconsistent, or why students felt cornered. To be useful for B4 you need the mechanism, and the mechanism lives in open-text detail, not in an average — the core of why averaging Likert scores misleads. Well-designed evaluation is also the leading indicator to B4's lagging outcomes: by the time an academic-misconduct case or an appeal lands, the flawed assessment has already run. The same leading-vs-lagging logic we set out for the B3 student-outcomes condition applies to assessment integrity.
The regulatory context has hardened around exactly this point. Since the Skills and Post-16 Education Act 2022, it has been a criminal offence in England to provide, arrange or advertise paid contract-cheating ("essay mill") services to students — the government announced the ban as essay mills becoming illegal in April 2022. Criminalising the supplier does not, however, discharge a provider's own B4 duty to design assessments that resist misconduct in the first place.
Generative AI sharpens all of this. B4's requirement to design assessments that "minimise the opportunities for academic misconduct" now has to contend with tools that can produce a passable essay in seconds — a live design challenge we explore in how to evaluate courses that use generative AI. Students are your fastest source of intelligence on which assessment formats have quietly become AI-trivial.
But doesn't student feedback risk lowering standards under B4?
The strongest counterargument is real: if you feed student voice into assessment, won't students push for easier assessments, more lenient marking and lighter loads — the very things that erode the rigour B4 protects? Grade-leniency and workload pressures are well documented, and B4 explicitly guards award credibility over time.
The answer is to use student evaluation for what it is good at and refuse to use it for what it is not. Students should not set standards, marking criteria or the difficulty of an award — that is expert and external-examiner territory, and constructive alignment of assessment to learning outcomes is a matter of design, not preference. Students should be heard on whether the assessment was intelligible, fairly administered, consistently marked and free of the design flaws that invite misconduct. Those are precisely the validity-and-reliability conditions B4 cares about, and hearing them strengthens standards rather than diluting them. The failure mode to avoid is treating a satisfaction mean as a mandate to make assessment easier; the correct posture is treating structured student evidence as a diagnostic that makes assessment sounder.
We keep the claim precise. Evaluation does not measure whether an assessment is valid or reliable in the technical sense B4 defines — that is judged by design, moderation and external examining. It surfaces the student-side conditions and warning signs that predict where validity, reliability and integrity are at risk.
Where Koji fits
Koji for Education replaces the static assessment-satisfaction item with an AI-moderated conversational interview that produces B4-relevant intelligence:
- It probes validity conditions directly. Rather than "rate the assessment," the interview can ask whether the brief and criteria were clear and follow up on exactly where they were not — turning a vague complaint into a specific, fixable design flaw.
- It surfaces reliability and fairness concerns — perceived marker inconsistency, unclear standards between cohorts — as structured, quality-scored themes through automatic thematic analysis, giving a course team early warning an external examiner might only formalise a year later.
- It flags misconduct-risk design — bunched deadlines, AI-trivial tasks, unrealistic timeframes — as themes, supporting the "reasonable steps ... by design" duty B4 imposes.
- Formative, mid-cycle collection catches these problems before the assessment runs, not after the appeal.
- Programme- and institution-level reporting aggregates assessment themes across modules so a quality team can evidence, to the OfS, that it monitors and acts on assessment risk — with standardised, bias-aware AI moderation and GDPR-compliant handling that keeps the evidence defensible.
Quality and standards teams that also run wider institutional research can use the same AI interview engine on the main Koji platform, keeping methodology consistent across the institution.
Koji does not mark work, moderate grades or certify that an award meets sector-recognised standards — those remain the province of examiners and academic regulations. What it does is give you the leading, mechanistic evidence about assessment conditions that turns B4 from a compliance risk you discover too late into one you can see coming.
Frequently asked questions
Does OfS Condition B4 mention student satisfaction? No. B4 requires that assessment is valid and reliable, that awards are credible and hold their value, and that providers take reasonable steps against academic misconduct. Satisfaction, module ratings and enjoyment appear nowhere in the condition — it is about assessment design and integrity, judged by expert and external-examiner standards.
If B4 isn't about satisfaction, why collect student feedback on assessment at all? Because students are the witnesses to the conditions that make assessment valid or invalid in practice — whether briefs were clear, marking felt consistent, workload was manageable, and designs invited shortcuts. That evidence is a leading indicator of B4 risk, distinct from judging the standards themselves.
Won't listening to students on assessment lower academic standards? Only if you misuse it. Students should not set difficulty, criteria or marking standards. They should be heard on clarity, fairness, consistency and misconduct-inviting design. Used that way, student evidence makes assessment sounder, not easier, and B4's award-credibility duty is preserved.
How does B4 relate to generative AI and essay mills? B4 requires assessments to be designed to minimise opportunities for misconduct and to facilitate detection. Generative AI and essay mills make that harder, and students are often the fastest source of intelligence on which assessment formats have become trivially outsourceable.
How does Koji help with B4 compliance specifically? Koji's AI interview probes assessment clarity, perceived fairness and workload, and its thematic analysis surfaces reliability and misconduct-risk concerns as structured evidence. That gives quality teams an auditable, leading-indicator record that they monitor and act on assessment risk — without ever using it to set standards.
Want to see B4 risk before it becomes an appeal? Explore Koji for Education — AI-moderated interviews and thematic analysis that turn assessment feedback into leading-indicator evidence.