New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods11 min read

Did Crossing the Threshold Change the Rating? Regression Discontinuity for Course-Evaluation Cutoff Effects

When a rule assigns students, instructors, or courses on either side of a sharp cutoff, regression discontinuity turns that arbitrary boundary into near-experimental evidence about what actually moved an evaluation score.

Koji Education Team

Product

The short answer

When a rule sorts students, courses, or instructors by a sharp numeric threshold — a grade boundary, a class-size cap, an honours cut-off, an eligibility score for a teaching intervention — you can estimate a genuinely causal effect on evaluation outcomes by comparing units just above and just below that line. This is regression discontinuity (RD), and its logic is simple: cases that barely miss a cut-off are, on average, almost identical to cases that barely clear it, so any jump in the outcome right at the boundary is attributable to the treatment the boundary triggers, not to the confounds that plague ordinary comparisons. RD does not replace a randomised trial, but where a real institutional rule exists, it is one of the most credible quasi-experimental designs available to a quality-assurance office.

BLUF for busy readers: Regression discontinuity exploits an arbitrary cut-off to approximate random assignment near the threshold. If your institution assigns anything about a course by a sharp rule, RD can tell you whether the rule caused the change in evaluation scores — not merely that the two are correlated. The estimate is local to the cut-off and only as good as the assumption that nobody precisely manipulated their position around it.

What the research says

Regression discontinuity was invented inside education research. Thistlethwaite and Campbell (1960), writing in the Journal of Educational Psychology, wanted to estimate the effect of receiving a National Merit award on students' later academic aspirations. They could not randomise who won. Instead, they noticed that awards were assigned by a qualifying score: students just above the threshold received a Certificate of Merit, students just below received only a letter of commendation. The two groups differed by a hair on the qualifying test but were treated very differently by the award system. By modelling the outcome as a function of the qualifying score and looking for a discontinuity — a vertical jump in the regression line — exactly at the cut-off, they isolated the award's effect from the general tendency of higher-scoring students to do better anyway. Their sample was substantial for the era: 5,126 certificate recipients and 2,848 commended students.

The design lay comparatively dormant for decades before economists rediscovered and formalised it. Imbens and Lemieux (2008) in the Journal of Econometrics and Lee and Lemieux (2010) in the Journal of Economic Literature set out the modern estimation framework: the distinction between sharp RD (crossing the threshold determines treatment with certainty) and fuzzy RD (crossing the threshold only changes the probability of treatment, estimated via an instrumental-variables step); local-linear or local-polynomial estimation inside a bandwidth around the cut-off; and the battery of validity checks — covariate-balance tests, a density test for manipulation of the running variable, and placebo cut-offs — that a credible RD study must pass. Lee and Lemieux's central insight is that, under mild continuity assumptions, "the variation in treatment near the threshold is randomised as though from a randomised experiment." That is a strong claim, and it is what makes RD attractive: it converts an administrative rule into a natural experiment.

Why it matters for course evaluation in practice

Universities are full of sharp rules, and every one of them is a potential RD design sitting in your data warehouse:

  • Class-size caps. Suppose enrolment above 30 triggers a second seminar section, or above 150 moves a course from a seminar room to a lecture theatre. Do evaluation scores drop discontinuously the moment a course crosses the cap? Compare courses that enrolled 28–30 with those that enrolled 31–33. If scores fall at exactly 30, the cap — not some general "big courses are worse" story — is implicated.
  • Grade or GPA thresholds. If students below a GPA line are placed on academic probation, are they systematically harsher (or more grateful) evaluators of the courses that pushed them over? The probation cut-off is the running variable.
  • Intervention eligibility. Many teaching-development programmes, tutoring schemes, or "at-risk course" reviews are triggered when a prior evaluation score falls below a threshold — say, any course averaging below 3.5 gets a mandatory peer-observation. RD around 3.5 tells you whether the intervention actually raised next year's score, cleanly separating the intervention's effect from simple regression to the mean (a low score is partly bad luck and tends to bounce back on its own — the single most common way institutions fool themselves into thinking an intervention worked).
  • Honours and prize thresholds. Where a mark boundary (e.g. a 70 for a first-class classification in the UK system) changes a student's outcome, RD can probe whether students who just missed the boundary rate the course and teaching differently from those who just cleared it.

The probation and "at-risk course review" example deserves emphasis, because it is where course-evaluation offices most often mislead themselves. When you flag every course below a cut-off and require an intervention, next year's scores for those courses will rise even if the intervention does nothing at all — because you selected them at a temporary low point. A naïve before-and-after comparison will credit your intervention for what is really statistical noise reverting. RD is one of the few designs that gets the counterfactual right here: the courses just above the 3.5 line form the comparison, and they regress to the mean by the same amount, so the discontinuity at the cut-off isolates the intervention's true effect.

Limitations and honest caveats

RD is powerful but narrow, and a sophisticated reader will press on all of the following:

  1. The estimate is local. RD identifies the treatment effect at the cut-off, for units near it. A class-size effect estimated at the 30-student cap tells you little about what happens at 200. Do not generalise the local average treatment effect to the whole distribution.
  2. Manipulation of the running variable is fatal. RD's credibility rests on units being unable to precisely control which side of the line they land on. If instructors can lobby to keep an enrolment just under a cap, or if a marker nudges a 69 up to a 70, the groups on either side are no longer comparable. The standard defence is the McCrary (2008) density test: if the histogram of the running variable shows a suspicious pile-up on one side of the cut-off, be very cautious.
  3. Bandwidth and functional form drive the answer. Choose too wide a bandwidth and you import units that differ systematically; too narrow and you have almost no data and enormous uncertainty. Modern practice uses data-driven optimal-bandwidth selection and reports sensitivity to it. A jump that appears at one bandwidth and vanishes at another is not a finding.
  4. You need a genuinely sharp, well-populated cut-off. Many institutional "rules" are actually applied with discretion, or so few cases sit near the boundary that the estimate is hopelessly imprecise. RD is a design you discover in your rules, not one you can impose on any question.
  5. It answers "did the threshold cause a change", not "why". A discontinuity tells you the rule mattered; understanding the mechanism still requires qualitative evidence — which is where structured follow-up matters (see below).

Honest framing: RD is a scalpel, not a general-purpose tool. When a sharp rule exists, it is close to the credibility of an experiment. When it does not, forcing an RD frame produces a confident-looking artefact. For most everyday confound questions, a design such as difference-in-differences or propensity-score matching will fit the data better.

How Koji incorporates this

Koji is built on the premise that a raw evaluation number is the start of an inquiry, not the end — and RD is a discipline for asking better causal questions of that number. Concretely:

  • Threshold-aware reporting. Where an institution configures a review or intervention trigger (for example, "flag any course below 3.5 for peer observation"), Koji records the trigger threshold as structured metadata rather than a buried business rule. That makes the running variable and the cut-off explicit, so an analyst can later run — or a Koji report can surface — a proper regression-discontinuity comparison of courses just above and just below the line, instead of the misleading naïve before-and-after that credits interventions for regression to the mean.
  • Structured question types that supply the running variable and the outcome. Koji's scale and single_choice questions produce the clean numeric outcomes RD needs, while course-level metadata (enrolment, modality, section) supplies candidate running variables. Because the data is captured in a structured schema, not free-text PDFs, assembling an RD dataset is a query rather than a manual transcription project.
  • Probing the "why" the discontinuity can't explain. RD is silent on mechanism. Koji's AI-moderated conversational interviews are designed to fill exactly that gap: when a threshold effect appears in the numbers, follow-up open_ended prompts and adaptive probes ask students what changed about their experience, and automatic thematic analysis surfaces the mechanism the discontinuity only hints at. This is a mitigation of RD's core limitation, not a claim to replace it.
  • Guarding against manipulation. Because Koji timestamps and logs collection, an analyst can inspect whether responses or scores cluster suspiciously around an institutional cut-off — the qualitative analogue of a McCrary density check — before trusting an RD estimate.

Koji is careful not to overclaim: the platform does not compute a peer-reviewed RD estimate on your behalf, and no software converts a bad cut-off into a good design. What it does is keep the ingredients — thresholds, running variables, clean outcomes, and mechanism-probing follow-up — in a structured, analysis-ready form so that a competent institutional-research team can run the design properly. Koji's core research platform at koji.so applies the same AI-moderated interview engine to product and customer research, where threshold and eligibility rules create the very same natural-experiment opportunities.

References

Frequently asked questions

What is the difference between sharp and fuzzy regression discontinuity? In a sharp design, crossing the cut-off determines treatment with certainty — everyone above the line gets it, everyone below does not. In a fuzzy design, crossing the cut-off only changes the probability of treatment (some eligible units decline, some ineligible ones get in anyway). Fuzzy RD is estimated with an instrumental-variables step that scales the jump in the outcome by the jump in treatment probability at the threshold.

How is regression discontinuity different from difference-in-differences? Difference-in-differences compares changes over time between a treated and an untreated group, relying on a parallel-trends assumption. Regression discontinuity compares units on either side of a threshold at essentially one moment, relying on the continuity of everything except treatment at the cut-off. RD needs a sharp rule; difference-in-differences needs a credible comparison group and a clean before/after. They answer causal questions in different data situations.

Why does regression discontinuity handle regression to the mean when a before-and-after comparison does not? If you flag low-scoring courses for an intervention, they will rebound next year partly through luck reverting — regression to the mean — with or without the intervention. A before-and-after design mistakes that rebound for an effect. RD uses courses just above the threshold as the comparison; they regress by the same amount, so the discontinuity at the cut-off isolates the intervention's true incremental effect.

Can Koji run a regression discontinuity analysis for me automatically? No, and any tool that claims to is overselling. Koji keeps the ingredients analysis-ready — explicit thresholds, clean numeric outcomes, course-level running variables, and mechanism-probing follow-up interviews — so your institutional-research team can run a properly specified RD with appropriate bandwidth selection and validity checks. The statistical judgement stays with a competent analyst.

What is the single most important assumption to check? That units cannot precisely manipulate which side of the cut-off they land on. If instructors can keep enrolment just under a cap, or markers nudge borderline grades across a boundary, the groups stop being comparable and the design breaks. Inspect the density of the running variable around the cut-off (a McCrary-style test) before trusting any RD estimate.

Related resources