New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods9 min read

Did Your Adjustment Actually Work? Negative Controls for Detecting Hidden Bias in Course-Evaluation Comparisons

After adjusting for the confounders you could measure, a negative control tests for the ones you missed: a variable that should show no effect if your analysis is unbiased. If it lights up, your comparison is contaminated.

Koji Education Team

Product

In brief

After you adjust a course-evaluation comparison for the confounders you could measure, you still cannot see the confounders you missed. A negative control is a cheap, powerful falsification test: pick a variable that should show no effect if your analysis is unbiased — an outcome the intervention cannot plausibly cause, or an "exposure" that cannot plausibly matter — and check. If the negative control lights up, your adjusted comparison is contaminated by residual confounding or selection bias, and its headline result should not be trusted.

What the research says

Lipsitch, Tchetgen Tchetgen and Cohen (2010, Epidemiology 21(3):383-388, "Negative Controls: A Tool for Detecting Confounding and Bias in Observational Studies") gave the method its modern framework. They distinguish two families. A negative control outcome is a response that shares the study's confounding structure but cannot be caused by the exposure; if the exposure appears to "affect" it, that association can only come from bias. A negative control exposure is a putative cause that shares the confounders but has no real pathway to the outcome; again, any apparent effect is a bias signal. The paper sets out the conditions under which a clean negative control has the same confounders as the real analysis — the property that makes it diagnostic rather than merely reassuring.

Arnold and Ercumen (2016, JAMA 316(24):2597-2598, "Negative Control Outcomes: A Tool to Detect Bias in Randomized Trials and Observational Studies") brought the idea to a broad clinical audience, framing negative controls as a routine, low-cost check that can reveal confounding, selection and measurement bias that the primary analysis hides. Shi, Miao and Tchetgen Tchetgen (2020, Current Epidemiology Reports 7:190-202, "A Selective Review of Negative Control Methods in Epidemiology") review the field's rapid development, including approaches that go beyond detecting bias to correcting for it (proximal causal inference) when a pair of suitable negative controls is available. Their central message for practitioners is that a negative control is only as good as the assumption that it truly shares the confounding of the primary comparison while being disconnected from the causal effect of interest.

Conceptually, the negative control is the observational cousin of the placebo test and the pre-trend check already familiar from quasi-experimental designs: it asks the data to reproduce a result you know should be null, and treats failure as evidence against the design.

Why it matters for course evaluation in practice

Quality offices constantly report adjusted comparisons — this instructor versus the department, this year versus last, online versus in-person — after controlling for the handful of covariates they happen to have. The unavoidable worry is the covariate they don't have: student motivation, cohort ability, the reason a course was scheduled at 8 a.m. A negative control turns that worry into a test.

Two designs travel well into evaluation. A negative control outcome is an item the intervention cannot plausibly move. Suppose a department redesigns its teaching and its overall ratings rise. If, over the same period, students' ratings of the library, the timetable, or campus facilities — things the teaching redesign did not touch — rise by a similar amount, the "improvement" is probably a cohort or response-climate shift, not the redesign. If those control items stay flat while the teaching items move, the causal story is far more credible. A negative control exposure flips the logic: if instructors assigned to a rooming quirk or an alphabetical scheduling accident that has no bearing on teaching quality nonetheless show a rating "effect", your comparison is picking up something other than teaching.

This is a natural complement to the sensitivity analysis in our E-value article: the E-value asks how strong an unmeasured confounder would have to be to explain away your result, while a negative control asks the data whether such confounding is actually present. It also pairs with the design thinking behind difference-in-differences and collider bias: a failed negative control is often the first visible symptom that a comparison is confounded or that conditioning on response has opened a spurious path.

Limitations and honest caveats

A negative control detects bias; it rarely proves the absence of it. Passing the test means this particular pathway of confounding did not show up in this particular control — a confounder that affects the real outcome but not your chosen control will slip through, so a clean negative control is encouraging, not exonerating. The whole method rests on an untestable judgement: that the control genuinely shares the primary analysis's confounders while having no real causal link to the exposure. Choose a control that is too disconnected and it detects nothing; choose one secretly affected by the intervention and you will "detect bias" that is really a true effect.

There are also power limits. With small classes and noisy ratings, a negative control has to move quite a lot before you can distinguish a real bias signal from sampling noise — the same 4.2-vs-4.4 problem that plagues primary comparisons. And a negative control tells you that something is wrong, not how much your headline estimate is off; correcting bias (rather than flagging it) needs the stronger, assumption-heavy proximal-inference machinery reviewed by Shi and colleagues, which most evaluation datasets cannot support.

How Koji incorporates this

Koji makes negative controls practical because it captures the "should-be-null" signals in the same instrument as the primary ones. A conversational evaluation naturally ranges over aspects the teaching intervention did not touch — facilities, scheduling, administrative processes — and Koji's structured question types let you retain a few of these as deliberate control items rather than discarding them as off-topic. When a redesign's ratings rise, Koji's reporting can show whether the untouched control aspects moved in lockstep (a red flag for a cohort or climate shift) or stayed flat (support for a genuine effect).

Because Koji preserves per-respondent structure and metadata, it also supports negative-control exposures: comparing groups defined by an incidental, non-teaching feature to check that no spurious "effect" appears. Koji's bias-aware reporting is designed to present these falsification checks alongside the headline comparison, so a quality committee sees the diagnostic, not just the conclusion — and its AI-moderated probing helps establish whether a control item really is causally untouched by the intervention, which is the assumption the whole method depends on. Koji is built to surface residual bias through such checks, not to claim it can certify a comparison as unconfounded; a passed negative control is reported as reassurance, not proof. Research teams using Koji's core platform at koji.so apply the same trick to product experiments, checking that a feature launch did not "improve" a metric it could not possibly have touched.

Frequently asked questions

What is a negative control in plain terms?

It is a variable chosen so that, if your analysis were unbiased, it would show no association — an outcome the intervention cannot cause, or an exposure that cannot affect the outcome. Observing an association there is evidence that residual confounding, selection or measurement bias is present.

What is the difference between a negative control outcome and a negative control exposure?

A negative control outcome is a response that shares the confounding but cannot be caused by the exposure. A negative control exposure is a putative cause that shares the confounding but has no real path to the outcome. Both should be null under an unbiased analysis.

How is this different from the E-value?

The E-value quantifies how strong an unmeasured confounder would need to be to nullify your result; it is a "what if" calculation. A negative control is an empirical test that asks whether such confounding is actually showing up in your data. They are complementary.

If the negative control passes, is my result confirmed?

No. Passing means one plausible bias pathway did not appear in one chosen control. A confounder that affects the true outcome but not your control will not be caught. A clean negative control is reassuring but never proof of an unbiased estimate.

How do I pick a good negative control for course evaluation?

Choose something that shares the same students, cohort and response climate as your primary outcome but that the intervention could not plausibly affect — ratings of facilities, timetabling or administration are common choices for a teaching-focused intervention.

Can negative controls correct bias, not just detect it?

In principle yes, with a matched pair of controls and the proximal-causal-inference framework, but that requires strong additional assumptions and richer data than most evaluation datasets provide. For everyday practice, treat negative controls as a detection and falsification tool.

References

Related resources

Related articles

analysis-reporting

Propensity Score Matching for Course-Evaluation Confounds

How propensity score matching (PSM) adjusts student-evaluation-of-teaching comparisons for confounds like class size, course level, discipline, and grading — and where it honestly falls short.

analysis-reporting

Collider Bias: How Selecting Who Gets Evaluated Invents Correlations That Are Not There

Berkson''s paradox and collider bias explain why analysing only the students who respond — or only the courses that survive — can manufacture correlations that do not exist in the population. What the causal-inference literature says, and why controlling for a collider makes things worse.

analysis-reporting

Did the Teaching Change Cause the Score to Move? Difference-in-Differences for Course Evaluation

Difference-in-differences lets you estimate whether a course redesign actually moved evaluation scores by comparing a treated course against a similar untouched one over time. Here is how the design works, when its parallel-trends assumption holds, and how to use it honestly in quality assurance.

analysis-reporting

One Instructor, Two Hundred Defensible Scores: Specification-Curve Analysis for Course Evaluation

Every course-evaluation "score" is the product of dozens of defensible analytic choices — which items count, how to weight, mean or median, whether to adjust for class size. Specification-curve and multiverse analysis compute the result across all of them, so a personnel decision rests on the pattern, not on one analyst's arbitrary path.