How Strong Would the Hidden Confounder Have to Be? The E-Value for Adjusted Course-Evaluation Comparisons
Every adjusted course-evaluation claim invites the objection "but you did not control for X". The E-value, from epidemiology, quantifies exactly how strong that unmeasured X would have to be to explain away your finding — turning a vague worry into a number.
Koji Education Team
Product
In brief
Whenever you report an adjusted course-evaluation comparison — "after controlling for class size and prior grades, online sections still score 0.3 points lower", or "women receive lower ratings net of discipline and cohort" — a critic can always reply, "but you did not control for X". The E-value, introduced by VanderWeele and Ding (2017), converts that open-ended worry into a single number: it is the minimum strength of association an unmeasured confounder would need to have with both the grouping variable and the rating, over and above the covariates you already adjusted for, to fully explain away the observed effect. A large E-value means only an implausibly strong hidden confounder could overturn your finding; a small E-value means a modest, plausible one could. It is a discipline for honest observational reporting, not a certificate of causation.
What the research says
The E-value generalises a classic argument. Cornfield and colleagues (1959), defending the causal reading of the smoking–lung-cancer link, reasoned that for an unmeasured confounder to explain a risk ratio of roughly nine, it would itself have to be about nine times more common in smokers — an implausibility that made confounding an unconvincing alternative. This is the seed of modern sensitivity analysis: rather than asserting no confounding, quantify how much would be required.
Ding and VanderWeele (2016), in Epidemiology, put this on rigorous footing with a "sensitivity analysis without assumptions". They derived a bounding factor and a sharp inequality that make no assumptions about the distribution or direction of the unmeasured confounder: given how strongly a confounder is associated with the exposure and with the outcome, their inequality bounds the maximum bias it could produce. Building on that, VanderWeele and Ding (2017), in Annals of Internal Medicine, defined the E-value as the smallest joint association (on the risk-ratio scale) an unmeasured confounder would need with both exposure and outcome to reduce an observed association to the null. For a risk ratio RR at or above 1, the E-value has a simple closed form: E-value = RR + √(RR × (RR − 1)). So a risk ratio of 2 has an E-value of about 3.4 (a confounder would need risk-ratio associations of at least 3.4 with both the grouping and the outcome to explain it away); a risk ratio of 1.5 has an E-value of about 2.4. You can compute a second E-value for the confidence-interval limit nearest the null, which describes how robust the statistical significance is, not just the point estimate.
Because course-evaluation outcomes are usually continuous mean ratings rather than risk ratios, VanderWeele and Ding also give an approximate conversion from a standardised mean difference d to an approximate risk ratio (roughly RR ≈ exp(0.91 × d)), so an adjusted gap in mean ratings can be placed on the same scale and given an E-value. The conversion is approximate and should be reported as such.
Why it matters for course evaluation in practice
Course-evaluation analysis is saturated with observational, covariate-adjusted claims: adjusted "league tables" of instructors, bias studies that control for discipline and grades, before-after comparisons of a teaching change. None of these is a randomised experiment, so every one is vulnerable to the objection that some unmeasured factor — a simultaneous timetable change, a cohort's unusual composition, an unrecorded room problem — drove the result. Reviewers and committees know this, which is why observational findings are so often waved away with "correlation is not causation" and no further thought.
The E-value replaces that stalemate with a quantified conversation. Instead of arguing about whether confounding might exist, you report: "to explain away this 0.3-point adjusted gap, an unmeasured confounder would need associations of at least 1.9 with both being an online section and the rating — is any plausible factor that strong, given we already adjusted for class size, discipline, and prior grades?" That reframes the debate onto plausibility of a specific magnitude, which is a far more productive question for a QA committee than an abstract worry. It is equally valuable defensively: if your instructor comparison has an E-value barely above 1, you should not be making high-stakes personnel decisions on it, because a trivial unmeasured confounder could reverse it. Reporting E-values alongside adjusted effects is a low-cost way to make observational course-evaluation evidence both more credible when it is strong and more honestly caveated when it is weak.
Limitations and honest caveats
The E-value is a humility statistic, and treating it as a green light is the main way it is abused. It addresses unmeasured confounding only — not measurement error, selection or collider bias, reverse causation, or model misspecification; a study can have a huge E-value and still be wrong for one of those other reasons. A large E-value does not prove no confounder exists; it only says a weak one is not enough. You must still argue, substantively, whether a confounder of the required strength is plausible — an E-value reported without that argument is empty ritual, a critique several methodologists have pressed since 2017. The bound is a worst case: it assumes the confounder is arranged to do maximum damage, which some argue makes the E-value conservative. The continuous-outcome conversion is approximate, so E-values for mean-rating differences should be read as order-of-magnitude, not exact. And the whole exercise presupposes your measured adjustment was done sensibly — E-values sit on top of a covariate model, not in place of one. Finally, none of this rescues the deeper problem that ratings are not learning: a confounding-robust effect on scores is still an effect on scores.
How Koji incorporates this
Koji cannot make an observational comparison causal — nothing can — but it can make adjustments better-specified and their fragility visible, which is where the E-value earns its place:
- Richer measured covariates, fewer things left unmeasured. Koji captures structured context alongside ratings — cohort size, delivery mode, discipline, perceived workload — so the covariate model behind any adjusted comparison is stronger to begin with, shrinking the room for a hidden confounder.
- E-values attached to adjusted reports. Koji's bias-aware reporting is designed to report an adjusted gap with its E-value (and the E-value for the confidence limit), so a committee sees the robustness alongside the estimate rather than the estimate alone.
- Surfacing candidate confounders to judge plausibility. Because the hard part is arguing whether a confounder of the required strength is plausible, Koji's AI-moderated conversational interviews and automatic thematic analysis of open text can surface factors students actually raise — a room that was freezing, a clashing deadline — giving reviewers concrete candidates to weigh against the E-value rather than an abstract unknown.
- Framed as humility, not a certificate. Koji presents the E-value as a caveat on an observational claim, never as evidence of causation, and pairs high-stakes comparisons with the reminder that scores are not learning.
The same measurement-and-interview stack powers Koji's core research platform at koji.so, where product teams face the identical challenge of defending observational effects of a launch against unmeasured confounding.
Frequently asked questions
What exactly is an E-value? It is the minimum strength of association, on the risk-ratio scale, that an unmeasured confounder would need with both the grouping variable and the outcome — beyond the covariates already adjusted for — to fully explain away the observed effect. Larger means more robust.
How do I get an E-value for a difference in mean ratings? Convert the adjusted standardised mean difference to an approximate risk ratio (roughly RR ≈ exp(0.91 × d)) and apply the E-value formula. Report it as approximate, because the continuous-to-risk-ratio conversion is not exact.
Does a large E-value prove my finding is causal? No. It only shows that a weak unmeasured confounder could not overturn it. You still have to argue whether a confounder of the required strength plausibly exists, and the E-value says nothing about measurement error, selection bias, or reverse causation.
Should I compute the E-value for the point estimate or the confidence interval? Both are useful. The point-estimate E-value describes robustness of the effect size; the E-value for the confidence-interval limit nearest the null describes how robust the statistical significance is. Reporting both is best practice.
Is the E-value a substitute for adjusting for confounders? No. It sits on top of a sensible covariate model, quantifying residual vulnerability to what you could not measure. It is not a reason to skip measuring and adjusting for the confounders you can observe.
Related resources
- Collider Bias and Berkson's Paradox in Course-Evaluation Selection
- Propensity Score Matching for Course-Evaluation Confounds
- Selection Bias in Course Evaluations: Goos and Salomons
- Should You Statistically Adjust Course Evaluations? The IDEA Approach
- Value-Added Evidence vs Student Ratings
- Difference-in-Differences for Course Evaluation
References
- Cornfield, J., Haenszel, W., Hammond, E. C., Lilienfeld, A. M., Shimkin, M. B., & Wynder, E. L. (1959). Smoking and Lung Cancer: Recent Evidence and a Discussion of Some Questions. Journal of the National Cancer Institute, 22(1), 173–203. https://doi.org/10.1093/jnci/22.1.173
- Ding, P., & VanderWeele, T. J. (2016). Sensitivity Analysis Without Assumptions. Epidemiology, 27(3), 368–377. https://doi.org/10.1097/EDE.0000000000000457
- VanderWeele, T. J., & Ding, P. (2017). Sensitivity Analysis in Observational Research: Introducing the E-Value. Annals of Internal Medicine, 167(4), 268–274. https://doi.org/10.7326/M16-2607
Related articles
Propensity Score Matching for Course-Evaluation Confounds
How propensity score matching (PSM) adjusts student-evaluation-of-teaching comparisons for confounds like class size, course level, discipline, and grading — and where it honestly falls short.
Should You Statistically Adjust Course Evaluations for Class Size and Difficulty? The IDEA Approach to Adjusted Scores
Some course-evaluation systems report "adjusted" scores that statistically correct for class size, discipline difficulty and student motivation. We examine what the IDEA system actually adjusts for, whether the practice is defensible, and how to contextualise scores without over-correcting.
Collider Bias: How Selecting Who Gets Evaluated Invents Correlations That Are Not There
Berkson''s paradox and collider bias explain why analysing only the students who respond — or only the courses that survive — can manufacture correlations that do not exist in the population. What the causal-inference literature says, and why controlling for a collider makes things worse.
Did the Teaching Change Cause the Score to Move? Difference-in-Differences for Course Evaluation
Difference-in-differences lets you estimate whether a course redesign actually moved evaluation scores by comparing a treated course against a similar untouched one over time. Here is how the design works, when its parallel-trends assumption holds, and how to use it honestly in quality assurance.