How Much of the Rating Gap Is Bias? The Oaxaca-Blinder Decomposition for Course Evaluation
When two groups of instructors get different average ratings, the Oaxaca-Blinder decomposition splits the gap into a part explained by measurable circumstances and an unexplained residual. This article shows how to use it responsibly — and why the unexplained part is a bound on bias, not a measurement of it.
Koji Education Team
Product
In brief
When two groups of instructors receive different average ratings — men versus women, large classes versus small, one discipline versus another — the gap has two parts: a share explained by legitimate differences in circumstances (the groups teach different course types, sizes, or levels) and a residual share that those circumstances cannot explain. The Oaxaca-Blinder decomposition (Oaxaca, 1973; Blinder, 1973) formally splits a mean gap into an "explained" (endowments) component and an "unexplained" (coefficients) component. It is the standard tool for asking how much of a rating gap survives after you account for what you can measure — but the unexplained part is not proof of bias, only its upper bound plus everything you failed to measure.
What the research says
Ronald Oaxaca (1973) and Alan Blinder (1973), working independently on the male-female wage gap, developed a decomposition that has since become one of the most-used tools in applied social science. The idea is disarmingly simple. Fit a regression of the outcome on explanatory variables separately for each group. The difference in the two groups' average outcomes can then be algebraically split into:
- an explained (endowments) component — the part of the gap due to the two groups differing in their average characteristics (e.g., women teach more small seminars, men more large lectures), evaluated at a common set of returns; and
- an unexplained (coefficients) component — the part due to the two groups receiving different returns to the same characteristics (e.g., a large class costs a man fewer rating points than it costs a woman), plus any effect of variables you never measured.
In wage research the unexplained component is often labelled "discrimination," but the more careful literature is explicit that it is residual — it captures both genuinely differential treatment and the influence of any relevant characteristic left out of the model. Jann (2008) provided the canonical software implementation and, importantly, catalogued the technical hazards: how to compute correct standard errors that account for the estimated coefficients, and the identification problem in the "detailed" decomposition (attributing the unexplained gap to individual categorical predictors depends arbitrarily on the omitted reference category). Fortin, Lemieux, and Firpo (2011), in the Handbook of Labor Economics, situate Oaxaca-Blinder within the broader family of decomposition methods and stress its two key assumptions: a correctly specified, linear-additive outcome model, and the "index number" ambiguity of which group's coefficients to use as the non-discriminatory benchmark — a choice that changes the answer.
Why it matters for course evaluation in practice
Raw rating gaps between instructor groups are a recurring flashpoint in quality assurance and personnel decisions. A dean sees that women in the department average 4.1 while men average 4.4 and must decide what, if anything, it means. The naive readings are both wrong: "the men teach better" ignores that the groups may teach systematically different courses, and "this is 0.3 points of bias" ignores exactly the same thing in the other direction.
Oaxaca-Blinder gives the QA analyst a disciplined middle path. Suppose women in the department disproportionately teach small, discussion-heavy, required first-year courses — course types that the broader evidence links to lower ratings independent of teaching quality. The decomposition estimates how much of the 0.3-point gap is explained by that unequal distribution of course types, class sizes, and levels, and how much remains unexplained after adjustment. If the explained component accounts for 0.25 of the 0.3, the story is largely about course allocation — a workload-equity question. If only 0.05 is explained and 0.25 is unexplained, the residual is large enough to warrant serious scrutiny for bias, consistent with the experimental evidence on gender bias in ratings.
Crucially, the decomposition reframes the conversation from a single contested number to a structured question: what would this gap be if the two groups taught the same portfolio of courses? That is far more actionable than a raw average, and it names the specific measured factors driving the explained part.
Limitations and honest caveats
The Oaxaca-Blinder decomposition is often over-interpreted, and a careful reader will insist on the following caveats.
- The unexplained component is not "bias." It is a residual that also absorbs every rating-relevant variable you did not include — student prior interest, time of day, marking severity, cohort composition. Omit an important legitimate factor and it masquerades as "unexplained," inflating the apparent bias. The unexplained gap is an upper bound on discrimination given your model, not a measurement of it.
- It rests on a correctly specified, additive-linear model. Interactions and non-linearities that are mis-modelled leak into the unexplained component. Extensions (e.g., non-linear and quantile-based decompositions in Fortin et al., 2011) relax some of this but add assumptions.
- The index-number problem. The size of each component depends on which group's coefficients serve as the "no-difference" benchmark. Reasonable analysts can report different splits from the same data; the choice must be stated and justified.
- The detailed decomposition is fragile. As Jann (2008) shows, attributing the unexplained gap to specific categorical variables is sensitive to the arbitrary reference category. Report aggregate components with more confidence than variable-level ones.
- It describes; it does not identify causes. Decomposition is an accounting exercise on a fitted model, not a causal design. It cannot, by itself, tell you why returns differ — only that they do, conditional on your specification.
Used honestly, the method's value is precisely that it quantifies your own ignorance: it tells you how much of a gap your measured factors can and cannot account for, which is the right input to a human judgement, not a substitute for one.
How Koji incorporates this
A credible decomposition needs two things a course-evaluation system must supply: rich, comparable covariates to populate the "explained" side, and clean group-level linkage. Koji is designed to furnish both while keeping the interpretation honest.
- Rich covariates for the explained component. The quality of a decomposition is capped by what you can put on the explained side. Koji's structured question types and linked course metadata (level, size, modality, electivity, discipline) give analysts the measured factors needed to account for legitimate circumstance before anything is called "unexplained" - directly shrinking the omitted-variable problem that inflates apparent bias.
- Bias-aware, context-rich reporting rather than bare averages. Koji's reporting layer is built to present group comparisons with their context, discouraging the raw-average readings the decomposition is meant to replace, and complementing formal adjustment approaches.
- Open-text evidence to interrogate the residual. When a decomposition leaves a large unexplained gap, the next question is qualitative: is the residual patterned like bias? Koji's AI-moderated conversational interviews and automatic thematic analysis surface whether, say, women's evaluations contain more comments about warmth and personality and fewer about competence - the gendered-language signature that helps distinguish genuine differential treatment from an omitted covariate.
- Cross-context consistency. The same engine at koji.so is used to decompose satisfaction gaps across customer segments, where the identical caution applies: an unexplained gap is a prompt for investigation, not a verdict.
Koji is designed to make decomposition possible and responsible - supplying the covariates that keep the explained component honest and the qualitative evidence that keeps the unexplained component from being misread as proof of bias.
Frequently asked questions
What does the Oaxaca-Blinder decomposition actually tell me?
It splits the average rating gap between two groups into an "explained" part - due to the groups differing in measurable characteristics like course size, level, and type - and an "unexplained" part that those measured characteristics cannot account for.
Is the "unexplained" component the same as bias?
No. It is a residual that captures both genuinely differential treatment and the effect of any rating-relevant factor you failed to measure. It is best read as an upper bound on discrimination given your model, not a measurement of it.
How is this different from just adding controls to a regression?
A single pooled regression with a group dummy assumes both groups get the same returns to every characteristic and reports one adjusted gap. The decomposition fits each group separately, so it can reveal that the groups receive different returns to the same characteristics - which is exactly the "unexplained" component of interest.
What is the index-number problem?
Each component's size depends on which group's regression coefficients you treat as the non-discriminatory benchmark. Different reasonable choices give different splits, so the benchmark must be stated and justified rather than left implicit.
Can I trust the breakdown by individual variables?
Be cautious. Jann (2008) showed the "detailed" decomposition attributing the unexplained gap to specific categorical predictors is sensitive to the arbitrary omitted reference category. Report the aggregate explained/unexplained split with more confidence than variable-level attributions.
How many observations do I need?
Because it fits a separate regression per group, you need enough observations in each group to estimate those models stably - this is a programme- or institution-level analysis across many courses, not something to run on a single class.
References
- Oaxaca, R. (1973). Male-female wage differentials in urban labor markets. International Economic Review, 14(3), 693-709. https://doi.org/10.2307/2525981
- Blinder, A. S. (1973). Wage discrimination: Reduced form and structural estimates. Journal of Human Resources, 8(4), 436-455. https://doi.org/10.2307/144855
- Jann, B. (2008). The Blinder-Oaxaca decomposition for linear regression models. The Stata Journal, 8(4), 453-479. https://doi.org/10.1177/1536867X0800800401
- Fortin, N., Lemieux, T., & Firpo, S. (2011). Decomposition methods in economics. In O. Ashenfelter & D. Card (Eds.), Handbook of Labor Economics (Vol. 4A, pp. 1-102). Elsevier. https://doi.org/10.1016/S0169-7218(11)00407-2
Related resources
- How Do We Know It Is Bias? Natural-Experiment Evidence on Gender in Course Evaluations
- Gendered Language in Student Comments: Why Men Are Brilliant and Women Are Caring
- Should You Statistically Adjust Course Evaluations for Class Size and Difficulty? The IDEA Approach
- Why Was the Course Rated Well, Not Just Whether? Mediation Analysis and the Mechanism
- Does the Effect Depend on Who or What? Moderation and Interaction Effects
- Racial and Ethnic Bias in Student Evaluations of Teaching: What the Evidence Shows
Related articles
Gendered Language in Student Comments: Why Men Are ''Brilliant'' and Women Are ''Caring''
Bias in course evaluations is not only in the numbers — it is in the words. What Mitchell & Martin (2018) and Storage et al. (2016) found about systematically different language applied to men and women, and why open-text analysis must account for it.
Should You Statistically Adjust Course Evaluations for Class Size and Difficulty? The IDEA Approach to Adjusted Scores
Some course-evaluation systems report "adjusted" scores that statistically correct for class size, discipline difficulty and student motivation. We examine what the IDEA system actually adjusts for, whether the practice is defensible, and how to contextualise scores without over-correcting.
Does the Effect Depend on Who or What? Moderation and Interaction Effects in Course-Evaluation Analysis
A bias or a teaching effect on evaluations often holds only for some students, courses or conditions. Moderation analysis tests those "it depends" claims properly, and it is far harder than it looks.
Racial and Ethnic Bias in Student Evaluations of Teaching: What the Evidence Shows
Peer-reviewed evidence shows students rate instructors of colour systematically lower than White peers on the same teaching. We review Reid (2010), Chávez & Mitchell (2020) and Bavishi et al. (2010), the limitations, and how to design evaluation so race-linked bias does not contaminate quality decisions.