New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods11 min read

Where Should the Concern Threshold Sit? Standard-Setting Methods (Angoff, Bookmark) for Course-Evaluation Triggers

Most institutions pick a review trigger - below 3.5, act - out of thin air. Standard-setting, the discipline that decides pass marks on high-stakes exams, offers a defensible, panel-based way to set a criterion-referenced course-evaluation threshold that will survive an appeal.

Koji Education Team

Product

The short answer

Almost every quality-assurance office has a rule like "any course scoring below 3.5 out of 5 is flagged for review." Ask where the 3.5 came from and the honest answer is usually: it felt about right. That is a standard-setting problem, and educational measurement has spent decades building defensible procedures for it — chiefly the Angoff and Bookmark methods used to set pass marks on licensure and certification exams. Borrowing them lets you replace an arbitrary trigger with a criterion-referenced cut score derived from the structured judgement of a panel, documented well enough to survive an instructor's appeal or an accreditor's question.

BLUF: A defensible course-evaluation threshold is not discovered in the data; it is set by a procedure. Standard-setting convenes a panel (faculty, students, QA staff), defines the "borderline acceptable course," and derives the cut score from structured judgements about that borderline case — the Angoff method aggregates experts' expected item responses for a borderline course; the Bookmark method has experts place a bookmark in difficulty-ordered items. The cut score is an argument to be validated (Kane, 1994), not a fact, and different methods can yield different cuts — so the process, panel, and rationale must be recorded, not just the number.

What the research says

Standard-setting is the branch of educational measurement concerned with how to set defensible cut scores — the marks that separate "pass" from "fail," "proficient" from "basic." Its methods transfer directly to the question "what evaluation score should trigger action?"

The Angoff method (Angoff, 1971, in a now-famous footnote to his chapter in Educational Measurement) is the most widely used approach in modern certification and licensure. Panellists first agree on a description of the minimally acceptable candidate — for us, a borderline acceptable course: not excellent, not failing, exactly on the line. Each panellist then estimates, item by item, how that borderline case would respond, and the estimates are averaged into a cut score. Its dominance, as Cizek & Bunch (2007) document in the standard reference text, comes from concrete virtues: it is tied directly to the content of the instrument, it produces a clear audit trail, it works with panels of any size, and it has decades of guidance for running it defensibly.

The Bookmark method (Mitzel, Lewis, Patz & Green, 2001) was designed for the era of item-response theory. Panellists receive items — or, for a rating scale, score points — ordered from easiest to hardest to endorse, and place a "bookmark" at the point where a borderline case would stop responding favourably. Unlike Angoff, it typically requires real response data first, which some argue grounds the judgement more firmly in how the instrument actually behaves.

Crucially, the cut score is not a truth to be uncovered. Kane (1994), in Review of Educational Research, reframed standard-setting as an argument to be validated: a passing score is a policy judgement, and its defensibility rests on procedural evidence (was the method sound and followed?), internal evidence (do panellists agree?), and external evidence (does the cut classify cases sensibly?). This is the intellectual core you import — it tells a QA committee what would make its 3.5 legitimate.

And method choice is not innocuous. Empirical comparisons on real licensure exams — for example Yim (2018) on the Korean Medical Licensing Examination — find that the modified-Angoff and Bookmark methods can produce different cut scores from the same panel and content. That is not a flaw to hide; it is why standard-setting bodies recommend using more than one method and reconciling them.

Why it matters for course evaluation in practice

Thresholds in course evaluation are consequential and contested. A trigger decides which instructors get scrutinised, which courses get resourced, and sometimes which contracts get renewed. Setting it by feel invites three failures a standard-setting procedure prevents:

  1. Arbitrariness that collapses under appeal. "Why 3.5 and not 3.4?" has no answer if the number was guessed. A documented Angoff panel gives you a reasoned, on-the-record justification.
  2. Norm-referenced drift. Many institutions quietly flag "the bottom 10%," which — as the norm-referenced vs criterion-referenced distinction shows — guarantees that 10% are always "failing" even if every course is excellent. Standard-setting is explicitly criterion-referenced: it asks what is acceptable, not who is last.
  3. Unaccountable panels. Standard-setting formalises who decides and how, so students and junior faculty have a defined voice rather than the standard reflecting one administrator's intuition.

How this differs from Signal Detection Theory thresholds

Koji's knowledge base covers Signal Detection Theory for review thresholds, and the two are complementary, not competing. Signal Detection Theory optimises a cut given the costs of false alarms and misses and the base rate of genuine problems — it answers "given these error costs, where is the statistically optimal line?" Standard-setting answers a prior, more human question: "what does the community judge to be the minimally acceptable standard of teaching?" In practice you use standard-setting to establish the criterion the community will defend, then SDT and funnel plots to reason about the measurement error around it. One supplies legitimacy; the other supplies statistical discipline.

Limitations and honest caveats

  • It is judgement, not measurement. Standard-setting does not remove subjectivity; it organises and documents it. A panel can be systematically lenient or harsh, and the resulting cut inherits that.
  • Panel composition drives the result. Who sits on the panel — senior faculty, students, discipline mix — measurably moves the cut. This must be planned and disclosed, or the standard simply launders one group's preferences.
  • Different methods, different numbers. As Yim (2018) and others show, Angoff and Bookmark can disagree. A single-method cut should be treated as provisional; best practice reconciles at least two.
  • Borderline definitions are hard. The whole procedure hinges on a shared, concrete picture of the "borderline acceptable course." Vague descriptions produce noisy, low-agreement estimates.
  • A cut score is still a threshold on a noisy measure. Even a perfectly-set cut misclassifies courses whose true quality sits near the line, because a single term's rating carries substantial measurement error. Standard-setting fixes where the line is, not the fact that scores scatter around it.
  • High stakes demand more rigour. If the cut informs tenure or dismissal, the validity burden (Kane) is heavy: multiple methods, larger panels, documented validation, and periodic review.

How Koji incorporates this

Koji is built so that thresholds are configured deliberately and transparently, not left as magic numbers.

  • Criterion-referenced, documented thresholds. Koji lets an institution set review triggers as explicit criterion-referenced cut scores, with fields to record the standard-setting method, panel composition, and rationale — so the "why 3.5?" question has an inspectable answer attached to the threshold itself.
  • Support for panel-based standard-setting. Koji's structured-question and reporting tools can host the artefacts a panel needs: item-level distributions for Angoff estimates and difficulty-ordered score points for a Bookmark placement, drawn from the institution's own historical data.
  • Threshold plus uncertainty, together. Because a cut on a noisy score misclassifies borderline cases, Koji pairs any trigger with the measurement-error tools — confidence intervals, empirical-Bayes shrinkage, and funnel-plot context — so a course near the line is flagged for conversation, not automatic judgement.
  • Beyond the number. When a course crosses a threshold, Koji's AI-moderated conversational interviews and quote-anchored thematic analysis surface why, so the standard triggers understanding rather than a bare verdict — the closing-the-loop step a defensible QA process requires.

Koji frames all of this as decision support: standard-setting supplies a defensible line, but the institution owns the judgement, and Koji is designed to keep that judgement honest and documented rather than to automate it. The same configurable-threshold logic underpins prioritisation in product and customer research on Koji's core platform at koji.so.

Related resources

Frequently asked questions

What is the difference between standard-setting and just picking a threshold?

Picking a threshold produces a number with no justification. Standard-setting is a documented procedure in which a panel defines a borderline acceptable case and derives the cut from structured judgements about it. The result is a criterion-referenced cut score with a rationale, panel record, and method on file — defensible under appeal or accreditation review.

How does the Angoff method work for a course-evaluation threshold?

Panellists first agree on a concrete description of a borderline acceptable course. Each panellist then estimates how that borderline course would respond to each evaluation item, and the estimates are averaged into a cut score. It is content-referenced, produces a clear audit trail, and works with panels of any size.

How is this different from Signal Detection Theory thresholds?

Signal Detection Theory finds the statistically optimal cut given the costs of false alarms and misses and the base rate of real problems. Standard-setting answers the prior question of what the community judges to be minimally acceptable teaching. Use standard-setting to establish the criterion, then Signal Detection Theory and measurement-error tools to reason about the uncertainty around it.

Do Angoff and Bookmark give the same cut score?

Not necessarily. Empirical comparisons on licensure examinations (e.g. Yim, 2018) find the two methods can yield different cut scores from the same panel and content. This is why standard-setting bodies recommend applying more than one method and reconciling the results rather than trusting a single number.

Is a cut score objective?

No. A cut score is a policy judgement, and Kane (1994) argues its defensibility rests on validation evidence — that the procedure was sound and followed, that panellists agreed, and that the cut classifies cases sensibly. Standard-setting organises and documents subjectivity; it does not eliminate it.

Should the same threshold be used for low-stakes and high-stakes decisions?

No. The higher the stakes — tenure, non-renewal — the heavier the validity burden. High-stakes use calls for multiple methods, larger and more representative panels, documented validation, and periodic review, whereas a purely developmental trigger can be lighter-weight.

References

  • Angoff, W. H. (1971). Scales, norms, and equivalent scores. In R. L. Thorndike (Ed.), Educational Measurement (2nd ed., pp. 508–600). American Council on Education.
  • Mitzel, H. C., Lewis, D. M., Patz, R. J., & Green, D. R. (2001). The Bookmark procedure: Psychological perspectives. In G. J. Cizek (Ed.), Setting Performance Standards: Concepts, Methods, and Perspectives (pp. 249–281). Lawrence Erlbaum Associates.
  • Kane, M. (1994). Validating the performance standards associated with passing scores. Review of Educational Research, 64(3), 425–461. https://doi.org/10.3102/00346543064003425
  • Cizek, G. J., & Bunch, M. B. (2007). Standard Setting: A Guide to Establishing and Evaluating Performance Standards on Tests. SAGE Publications.
  • Yim, M. (2018). Comparison of results between modified-Angoff and Bookmark methods for estimating the cut score of the Korean medical licensing examination. Korean Journal of Medical Education, 30(4), 347–357. https://doi.org/10.3946/kjme.2018.110

Related articles

analysis-reporting

Stop Comparing Raw Averages: Empirical-Bayes Shrinkage for Course-Evaluation Scores

Why raw course-evaluation means mislead — especially for small classes — and how empirical-Bayes shrinkage (partial pooling), the method Kane and Staiger applied to teacher value-added, produces fairer, more accurate instructor estimates.

analysis-reporting

Should You Report an Instructor''s Percentile? Norm-Referenced vs Criterion-Referenced Course-Evaluation Scores

Telling a lecturer they are "in the 40th percentile of the department" is norm-referenced reporting — and it manufactures losers by construction, no matter how good everyone is. Criterion-referenced reporting asks instead whether teaching met a defined standard. Here is the evidence on why the choice matters and how to report responsibly.

analysis-reporting

Statistical Process Control for Course Evaluations: Using Control Charts to Separate Signal from Noise

How Shewhart control charts and statistical process control distinguish real changes in teaching quality from ordinary random variation in course evaluation scores - so QA teams stop reacting to noise. Grounded in Sivena and Nikolaidis (2019, 2022).

analysis-reporting

Which Instructor Scores Are Genuinely Unusual? Funnel Plots for Fair Course-Evaluation Comparison

Ranking instructors by mean evaluation score turns sampling noise into a league table. Funnel plots — Spiegelhalter's method for institutional comparison — show which scores are genuinely unusual and which are just small-sample wobble.