New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Methodology10 min read

What Score Counts as a Course of Concern? The Case for Standard-Setting, Not Arbitrary Cut-Offs

Flagging a course because it fell below 3.5, or in the bottom 10%, is a statistically indefensible decision dressed up as objectivity. Educational measurement has a whole discipline — standard-setting — for deriving cut scores that can survive a challenge. Course evaluation should borrow it.

Koji Education Team

Product ·

The short version: most universities decide a course or instructor is "of concern" by comparing a mean rating to a threshold — below 3.5, below the faculty average, in the bottom decile. Almost none can say why that number. In educational measurement, choosing the score that separates acceptable from unacceptable performance is a formal discipline called standard-setting, with established methods (Angoff, modified-Angoff, Ebel, bookmark) precisely because arbitrary cut scores do not survive scrutiny. If a course-evaluation threshold triggers a real consequence — a review, a note in a file, a probation conversation — it deserves the same defensibility you would demand of a professional licensing pass mark. A round number picked for convenience does not have it.

The arbitrary-cutoff problem

Walk into most quality-assurance offices and you will find a rule of thumb: courses below some line get flagged. The line is usually a tidy number (3.5 out of 5) or a relative rule (the bottom 10% each cycle). Both feel objective because they are numeric. Neither is.

The tidy-number cut is arbitrary in the literal sense: there is no argument connecting 3.5 to any defensible notion of "unacceptable teaching." As the standard-setting literature bluntly puts it, if you are making a criterion-referenced decision "it is not legally defensible to just conveniently pick a round number like 70%; you need a formal process" (Assessment Systems, modified-Angoff guide). The relative rule is worse: flagging "the bottom 10%" guarantees that exactly 10% of courses are branded as concerns every single year, regardless of their absolute quality — a department of uniformly excellent teachers will still be forced to nominate a scapegoat. That is the difference between criterion-referenced and norm-referenced standards, and using the wrong one turns evaluation into a ranking tournament nobody can win.

A cut score is a decision, and decisions have error

The deeper issue is that any threshold converts a noisy continuous measurement into a binary judgement — acceptable or not — and every such conversion produces two kinds of error: flagging a course that is actually fine (false positive) and clearing one that is not (false negative). The quality of a cut score is not whether the number looks reasonable; it is how often those two errors occur, and whether they are distributed fairly.

Course-evaluation data makes this acute because the underlying measurement is so imprecise. A class of twelve produces a mean with a confidence interval so wide that a 3.4 and a 3.8 are statistically indistinguishable — yet a 3.5 cut cleanly separates them into "concern" and "fine." Thresholding without reporting the interval around each score manufactures certainty that the data does not contain, which is the same error as reading a 0.3-point gap as a real difference. And because scores carry documented bias — Boring's natural experiment at a French university found male students roughly 30% more likely to award top overall marks to male instructors (Boring, 2017) — a raw-score cut bakes that bias directly into who gets flagged.

What standard-setting actually is

Standard-setting is the set of methods measurement professionals use to derive a defensible cut score. Its core idea is that the threshold should be set independently of the results, by informed judges reasoning about what minimally acceptable performance looks like, through a documented and repeatable process (Angoff methods systematic review, BMC Medical Education, 2025).

In the most common approach — the modified-Angoff method, "by far the most popular" for high-stakes credentialing — a panel considers each item and estimates how a just-barely-acceptable candidate would perform on it; those judgements are aggregated, often across rounds, sometimes with a "reality check" against real data. Crucially, the method reports its own reliability. The same review found that method choice matters enormously: a modified-Angoff with a reality check achieved excellent inter-panel reliability (r = 0.917), while a cruder Angoff yes/no variant was far weaker and more variable (r = 0.536). The classic reference text, Cizek and Bunch's Standard Setting (2007), catalogues the full family of methods and the validity evidence each requires.

Translated to course evaluation, this means: instead of decreeing "below 3.5 is a concern," a panel of experienced academics and QA staff defines what a just-acceptable course looks like on each dimension, derives the implied aggregate threshold, and — this is the part that separates rigour from theatre — reports the decision consistency: how reliably the process would classify the same course the same way on a re-run. A flag then comes with an error rate, not a false air of precision.

Making the cut honest

Three practices turn a threshold from a liability into a defensible instrument:

  1. Set the standard before you see the scores, by criterion (what acceptable teaching looks like), not by rank.
  2. Attach uncertainty to every classification. A course whose confidence interval straddles the cut is unclassified, not "just below." This alone would remove a large share of unfair flags, especially in small classes.
  3. Never flag on a bare number. Because scores are contaminated and weakly related to learning — student ratings explain at most about 1% of variance in actual learning (Uttl et al., 2017) — a threshold should trigger a look, corroborated by other evidence, not a verdict. This is also why student evaluations alone are a poor basis for tenure and promotion.

The strongest counterarguments, taken seriously

"Standard-setting is subjective — panels just disagree." Partly true, and beside the point. Every cut score is a judgement; the choice is only between judgement that is hidden and undocumented (a manager picking 3.5) and judgement that is explicit, panel-based, and auditable. Standard-setting does not remove subjectivity — it governs it, and reports how much of it remains as decision reliability. That is strictly more honest than a round number.

"It is disproportionately expensive for course evaluation." A full modified-Angoff exercise is heavy, and it would be absurd to run one for every low-stakes module report. The proportionality principle is the answer: the rigour of the standard should scale with the stakes of the decision. A dashboard that merely prompts a teacher's own reflection needs no panel; a threshold that feeds a personnel process needs a defensible one. Running high-stakes consequences off a convenience cutoff is the mismatch to fix.

"Course evaluation is too weak to justify any cut score at all." This is the most serious objection, and it is substantially right for summative, high-stakes use. The correct conclusion is not "no thresholds ever" but "thresholds should trigger support, not sanction." A cut score that routes a course toward formative help, teaching consultation, and a closer look at corroborating evidence is defensible in a way that one triggering punishment is not — and it sidesteps the multiple-comparisons problem where flagging enough instructors guarantees some look bad by chance.

Where an AI-native platform helps

Better standard-setting needs cleaner inputs and honest uncertainty — both areas where the measurement tool matters.

Koji for Education supports defensible thresholds rather than arbitrary ones:

  • Uncertainty by default. Koji's programme- and institution-level reporting can present each score with its interval, so a classification near the cut is shown as uncertain rather than falsely decisive — the single biggest fix for unfair flagging.
  • A cleaner score to threshold. Standardised, bias-aware AI moderation applies the same probing to every student and every instructor, removing the human-moderator inconsistency that adds construct-irrelevant variance to raw ratings. The number a panel sets a standard against is less contaminated to begin with.
  • Evidence to accompany a flag, not a bare mean. Koji's automatic thematic analysis and quality scoring mean any threshold breach arrives with the qualitative "why" attached — turning a flag into a triangulated case rather than a number, exactly what a defensible decision requires.
  • Thresholds that trigger help. Because Koji supports formative, mid-cycle collection and closing-the-loop action tracking, a cut score can route a course toward improvement while the term is still running, not toward a retrospective sanction.

Institutions that run high-stakes assessment or accreditation research elsewhere can apply the same conversational engine on the main Koji platform; the discipline of setting defensible thresholds transfers directly.

The bottom line

If your evaluation policy flags courses at a round number or by fixed rank, you are making consequential decisions on an indefensible basis — and you would not accept that standard for a pass mark on a professional exam. Borrow the discipline that field already built: set the standard by criterion, report the error around it, and let it trigger support rather than sanction. See how Koji for Education pairs cleaner scores with the uncertainty and evidence a defensible threshold needs.

Koji provides cleaner measurement and honest uncertainty; setting the standard itself remains a documented human judgement, as it should be.