Propensity Score Matching for Course-Evaluation Confounds
How propensity score matching (PSM) adjusts student-evaluation-of-teaching comparisons for confounds like class size, course level, discipline, and grading — and where it honestly falls short.
Koji Education Team
Product
Propensity Score Matching for Course-Evaluation Confounds
Can propensity score matching make course-evaluation comparisons fairer? Partly, and only under a strict condition. Propensity score matching (PSM) and related weighting methods can build a "like-for-like" comparison between courses or instructors that differ systematically in class size, level, discipline, prior student interest, and grading — removing bias from the confounders you have measured. But it adjusts for observed confounders only; anything you did not record still contaminates the comparison. Treat it as a disciplined improvement on naive raw averages, not as a causal guarantee.
What the research says
Student-evaluation-of-teaching (SET) scores are almost never comparable "as is." A large required first-year statistics course and a small third-year elective seminar draw different students, meet different expectations, and grade on different scales. When you rank instructors on raw mean SET, you are partly ranking the courses they were assigned, not the teaching they delivered. The statistical machinery for correcting this kind of imbalance in non-experimental data is the propensity score.
The propensity score and the balancing property
The foundational result is Rosenbaum and Rubin (1983). They defined the propensity score as the conditional probability of receiving the "treatment" given a vector of observed covariates — formally, e(x) = Pr(Z = 1 | X = x), where Z is the treatment indicator and X the observed covariates. Their central theorem is the balancing property: conditional on the propensity score, the distribution of the observed covariates is the same for treated and control units. In other words, a single scalar summarising many covariates is sufficient to remove the bias those covariates would otherwise introduce. They also established that if treatment assignment is strongly ignorable given the covariates — meaning assignment is independent of the potential outcomes once you condition on X, and every unit has a non-zero probability of either condition — then adjusting for the propensity score yields unbiased estimates of the average treatment effect (Rosenbaum & Rubin, 1983, Biometrika, 70(1), 41–55).
Crucially, that guarantee is conditional on having measured the right covariates. The propensity score removes bias due to observed covariates; it says nothing about variables you failed to record. This caveat is not a footnote — it is the whole basis on which such analyses stand or fall, and we return to it below.
The workflow
Applied propensity-score analysis follows a standard sequence, described in the methodological reviews by Stuart (2010) and Austin (2011):
- Estimate the propensity score. Regress the "treatment" indicator on the observed covariates (commonly logistic regression, sometimes machine-learning classifiers) to obtain each unit's estimated probability of treatment.
- Condition on it via one of several strategies. Stuart (2010) and Austin (2011) catalogue the main options: matching (pair each treated unit with one or more controls of similar propensity), stratification/subclassification (group units into strata of similar propensity and compare within strata), inverse-probability-of-treatment weighting (IPTW, reweight units by the inverse of their probability of the condition they received), and covariate adjustment using the score.
- Check balance — do not skip this. Both reviews stress that the point of the exercise is covariate balance, not the propensity model's own fit. After matching or weighting, compare covariate distributions across groups (standardised mean differences are the standard diagnostic; Austin recommends a threshold well below 0.1). If balance is poor, revise the model and repeat. This is an iterative design stage that, importantly, is carried out without looking at outcomes, preserving the objectivity of a designed experiment.
- Estimate the effect on the balanced sample, and report an appropriate estimand (average treatment effect, or average treatment effect on the treated).
Stuart (2010), in Statistical Science (25(1), 1–21), frames the entire enterprise as an attempt to "replicate a randomized experiment as closely as possible" by constructing treated and control groups with similar covariate distributions. Austin (2011), in Multivariate Behavioral Research (46(3), 399–424), provides the applied clinician's-eye comparison of the four methods and the diagnostics that separate a credible analysis from a decorative one.
Why raw instructor/course comparisons are confounded
The methodological apparatus matters here because the SET literature has documented, repeatedly, that scores are driven by course features that have little to do with teaching quality. The most pointed evidence is Braga, Paccagnella, and Pellizzari (2014), Economics of Education Review (41, 71–88). Using administrative data from Bocconi University, where students were effectively randomly assigned to teaching sections, they built an outcome-based measure of teacher effectiveness (subsequent exam performance in follow-on courses) and compared it to the same teachers' SET scores. Their striking result: teacher effectiveness was negatively correlated with student evaluations — the instructors whose students went on to perform better tended to receive lower ratings. Their random-assignment design is exactly what most institutions lack, which is why observational adjustment methods like PSM are the realistic fallback: when you cannot randomise students to instructors, you try to reconstruct comparability statistically instead.
Why it matters for course evaluation in practice
Consider the concrete comparison a quality-assurance office actually faces. Instructor A teaches a mandatory 300-student first-year quantitative methods course. Instructor B teaches a 15-student final-year elective seminar in their specialism. On raw mean SET, B almost always "wins." But the two are apples and oranges: class size, course level, whether the course is required or elected (a proxy for prior interest), discipline, and grading distribution all differ, and all are known correlates of SET.
Propensity-score methods let you construct a defensible like-for-like comparison. Treat the characteristic of interest — say, "taught a large required quantitative course" — as the "treatment," model the probability of a cohort having that profile from its metadata (size, level, modality, discipline, grading distribution, share of required-vs-elective enrolment), and then weight or match so that you are comparing cohorts with similar profiles. The resulting benchmark answers a fairer question: how does this instructor's evaluation compare to what we would expect from comparable courses taught to comparable students under comparable conditions? That is the difference between a raw-average league table — which quietly penalises anyone assigned a hard, large, required course — and a covariate-adjusted benchmark that a critical reader can defend in a promotion or programme-review meeting.
This is not exotic. It is the same reasoning hospitals use for risk-adjusted outcome reporting, and the same reasoning economists use when they cannot run the experiment they wish they could. Applied to SET, it shifts the conversation from "who scored highest" to "who scored higher than comparable peers in comparable conditions," which is the only version of the question that is fair to ask of individuals.
Limitations and honest caveats
A PhD reader will — correctly — raise the following objections, and an honest treatment names them up front.
- Unobserved confounding is the hard ceiling. The entire method rests on strong ignorability: that treatment assignment is independent of potential outcomes once you condition on the measured covariates. If an unmeasured variable — student motivation, a charismatic co-instructor, time-of-day effects, a cohort's baseline ability — drives both course type and ratings, PSM cannot touch it. Unlike randomisation, matching offers no protection against confounders you did not measure (Stuart, 2010). Sensitivity analysis (e.g., Rosenbaum bounds) can quantify how strong a hidden confounder would have to be to overturn a conclusion, but it cannot resurrect a variable you never collected.
- Common support / overlap. Matching only works where treated and control cohorts actually overlap in covariate space. If no small elective ever resembles a large required course, there is simply no fair comparator, and honest analysis must restrict the comparison to the region of common support and say so — rather than extrapolating a comparison that the data cannot support.
- Model dependence. Results can shift with the choice of propensity model, matching algorithm (nearest-neighbour, caliper, kernel), caliper width, and whether you weight or stratify. A conclusion that survives only one specification is fragile. Reporting should show that balance was achieved and, ideally, that findings are stable across reasonable specifications.
- Small-sample instability. Departments and cohorts are often small. With few courses per stratum, matched estimates become noisy and weights can blow up (extreme IPTW weights are a well-known failure mode). PSM does not manufacture information that a thin dataset lacks; it can make a small-sample comparison look precise while remaining unstable.
- It is observational, not experimental. Even done perfectly, PSM approximates a randomised experiment; it does not perform one. The output is a more-comparable comparison, not a proven causal effect of an instructor on ratings. Framing it as the latter would overclaim.
The intellectually honest summary is Rosenbaum and Rubin's own: the propensity score removes bias from the covariates you observe, and only those.
How Koji incorporates this
Koji for Education is designed to support covariate-adjusted, bias-aware evaluation reporting — it does not, and cannot, "eliminate confounding." The point is to give an institution the raw materials that credible adjustment requires, and to frame results so that the adjustment's limits stay visible.
- Structured metadata per study and cohort. Koji captures the covariates that PSM depends on — class size, course level, delivery modality, discipline, and cohort descriptors — as structured fields attached to each study rather than as free text buried in a report. This is precisely the covariate vector
Xthat any propensity model needs. Because the metadata is structured and consistent across cohorts, downstream analysis can construct like-for-like comparisons and covariate-adjusted benchmarks instead of naive raw-average league tables. - Triangulation over ranking. Koji is oriented toward comparing a cohort against comparable cohorts and toward triangulating multiple signals, rather than emitting a single decontextualised score. That design maps directly onto the "compare like with like" logic of propensity methods, and onto the caution that any single adjusted number is an estimate, not a verdict.
- Capturing the qualitative "why." Adjustment tells you whether a gap remains after controlling for observables; it does not tell you why. Koji's AI-moderated conversational interviews probe the reasoning behind a rating in a respondent's own words, and its structured question types —
open_ended,scale,single_choice,multiple_choice,ranking, andyes_no— let an institution collect both the quantitative signal an adjustment model consumes and the qualitative context that explains a residual difference no covariate captured. - Closing the loop. Koji supports action tracking so that findings translate into documented follow-up, keeping evaluation a mechanism for improvement rather than a one-off league table.
The same underlying engine powers Koji's core platform at koji.so, where it is applied to product and customer research — the analytical discipline of comparing like-for-like cohorts and capturing the qualitative "why" transfers cleanly from students to customers.
None of this makes an observational comparison experimental. It makes the covariates explicit, the adjustment auditable, and the residual uncertainty legible — which is the most an honest evaluation platform should claim.
References
- Austin, P. C. (2011). An introduction to propensity score methods for reducing the effects of confounding in observational studies. Multivariate Behavioral Research, 46(3), 399–424. https://doi.org/10.1080/00273171.2011.568786
- Braga, M., Paccagnella, M., & Pellizzari, M. (2014). Evaluating students' evaluations of professors. Economics of Education Review, 41, 71–88. https://doi.org/10.1016/j.econedurev.2014.04.002
- Rosenbaum, P. R., & Rubin, D. B. (1983). The central role of the propensity score in observational studies for causal effects. Biometrika, 70(1), 41–55. https://doi.org/10.1093/biomet/70.1.41
- Stuart, E. A. (2010). Matching methods for causal inference: A review and a look forward. Statistical Science, 25(1), 1–21. https://doi.org/10.1214/09-STS313
Related Resources
- Multilevel models for nested course-evaluation data
- Empirical Bayes shrinkage for instructor scores
- Small mean differences and confidence intervals in course evaluation
- Ranking instructors on evaluation scores fairly
- Misclassification risk when ranking instructors for personnel decisions
- Class-size effects on student evaluations
- Grading leniency and student evaluations of teaching
- Course difficulty, workload, and student evaluations
Related articles
Can You Fairly Rank Instructors by Their Course-Evaluation Scores? What Esarey & Valdes (2020) Proved
Even if course evaluations were unbiased, reliable, and valid, ranking instructors by their scores would still misclassify many good teachers. A walk through the Esarey & Valdes (2020) simulation and what it means for how you report and use evaluation data.
The 4.2 vs 4.4 Trap: Why Small Differences in Course-Evaluation Means Are Usually Noise
Faculty and administrators routinely read meaning into tiny gaps between course-evaluation means. Boysen (2015) and Boysen et al. (2014) show this happens even when confidence intervals say the difference is nothing — and that warnings barely help. How to report uncertainty honestly.
Even Fair Evaluations Can Rank Unfairly: Why Course-Evaluation Scores Misclassify Instructors
A simulation-based finding that even unbiased, reliable, and valid course-evaluation scores produce high misclassification rates when used to rank or compare instructors — and what it means for tenure, promotion, and merit decisions.
Stop Comparing Raw Averages: Empirical-Bayes Shrinkage for Course-Evaluation Scores
Why raw course-evaluation means mislead — especially for small classes — and how empirical-Bayes shrinkage (partial pooling), the method Kane and Staiger applied to teacher value-added, produces fairer, more accurate instructor estimates.