New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods11 min read

Which Item Is Biased? Differential Item Functioning Detection for Course-Evaluation Questions

Measurement invariance tests the whole scale; differential item functioning (DIF) pinpoints the single biased item. A practical guide to Mantel-Haenszel and logistic-regression DIF, uniform vs non-uniform bias, and what to do when a course-evaluation item behaves differently across groups.

Koji Education Team

Product

In brief: A test of measurement invariance tells you whether a course-evaluation scale, taken as a whole, means the same thing to two groups of students. Differential item functioning (DIF) goes one level deeper and pinpoints the specific item that does not. An item shows DIF when two students who are equally satisfied overall — the same standing on the underlying trait — nonetheless have systematically different expected answers to that one question purely because of their group membership (the instructor is a woman, the course is quantitative, the class is online, the respondent is a non-native speaker). The workhorse detection methods are the Mantel-Haenszel procedure (Holland & Thayer, 1988) and logistic regression (Swaminathan & Rogers, 1990), the latter separating uniform from non-uniform bias. DIF analysis is how a quality-assurance office moves from "our evaluations might be biased" to "this item is the problem — revise it, flag it, or drop it."

Why item-level analysis matters

Most bias discussions in higher education treat a course-evaluation form as a single black box: "student ratings are biased against women," "quantitative courses score lower." Those statements are usually true at the aggregate level, but they are not actionable. A dean cannot fix "the form"; they can only fix items. The value of DIF is that it decomposes an aggregate bias signal into its constituent questions, so you learn where the unfairness enters.

This is a genuinely different question from the one answered by scale-level measurement invariance. Invariance testing, usually run as multi-group confirmatory factor analysis, asks whether the factor structure, loadings, and intercepts of the whole instrument are equal across groups. When invariance fails, you know something is non-comparable — but not which item. DIF is the item-level microscope you reach for next: it identifies the offending question and quantifies the size of the effect for that question alone.

What the research says

The methods. Holland and Thayer (1988) adapted the Mantel-Haenszel (MH) statistic — originally a tool from epidemiology for stratified 2×2 tables — into the standard DIF detection procedure. The logic is elegant: instead of comparing raw item scores between, say, students of male and female instructors, you first stratify respondents into bands of equal total score (a proxy for the latent trait), then ask whether, within each matched band, the two groups still answer the target item differently. Matching on total score is what makes DIF a fair-comparison method rather than a naive group-difference test: it holds overall satisfaction constant and isolates the item-specific residual.

Swaminathan and Rogers (1990), writing in the Journal of Educational Measurement, introduced the logistic-regression approach that has since become dominant. They model the probability of endorsing an item as a function of (1) the matching variable (total score), (2) group membership, and (3) the group-by-score interaction. This decomposition is the method's key contribution: a significant group term signals uniform DIF (one group is consistently advantaged at every ability level — the curves are parallel but shifted), while a significant interaction term signals non-uniform DIF (the direction of the advantage flips across the trait continuum — the curves cross). In simulation, they showed logistic regression is more powerful than MH for detecting non-uniform DIF and comparable for uniform DIF. Rogers and Swaminathan (1993) confirmed the comparison directly.

The application to student ratings. The idea that student-evaluation items should be checked for cross-group equivalence is not new. Marsh and Hocevar (1984) examined the factorial invariance of the widely used SEEQ (Students' Evaluations of Educational Quality) instrument and found its multidimensional structure held reasonably well across settings — an early demonstration that evaluation instruments can, and should, be interrogated for measurement equivalence rather than assumed comparable. More recent multi-group confirmatory factor analyses of student perceptions of teaching (for example, van de Grift and colleagues' six-country study of teaching-behaviour perceptions) apply exactly the invariance logic that DIF then refines to the item.

Why it changes the interpretation of bias findings. If the gap between how students rate male and female instructors is driven by one or two items — say, an item about "authority" or "approachability" that carries gendered expectations — then the fair response is to revise those items, not to abandon the instrument or apply a blanket statistical correction. Conversely, if DIF is absent and the gap persists across every matched band, the difference is either a real difference in the measured trait or a pervasive response bias that no single-item fix will cure.

Why it matters for course evaluation in practice

A DIF audit turns an abstract fairness worry into a concrete editorial task list. Concretely, a QA office can:

  1. Screen the standard form. Run MH and logistic-regression DIF across the groups that matter for your context — instructor gender, course discipline (STEM vs humanities), delivery mode (online vs in-person), and student first-language status are the usual suspects. Flag any item with statistically significant DIF and a non-trivial effect size (the ETS classification — negligible/A, moderate/B, large/C — based on the MH delta or the change in Nagelkerke R² is the conventional threshold; significance alone over-flags in large samples).

  2. Diagnose the type. Uniform DIF (a consistent shift) often points to an item whose wording invokes a group-linked stereotype. Non-uniform DIF (a crossing) is subtler and usually signals that the item measures something different for the two groups at different satisfaction levels.

  3. Act. Revise the wording, replace the item with a low-inference behavioural item that is harder to answer stereotypically, or — if the item cannot be salvaged — drop it and document the decision as part of your validity argument.

  4. Re-test. DIF detection is iterative. After removing a flagged item, the matching variable (total score) is now cleaner, so re-running the analysis (a step called purification) can reveal items masked in the first pass.

Done well, this is the difference between a form that survives an equality-impact review and one that collapses the first time a union or an ombudsman asks "how do you know your evaluation questions are fair to women?"

Limitations and honest caveats

DIF is powerful but easy to misuse, and a PhD reader will raise the following objections — rightly.

  • Sample size cuts both ways. DIF tests need reasonable numbers in both groups within each matching band. Many individual courses are far too small; DIF is a form-level or department-level analysis run on pooled data across many sections, not something you compute for one seminar of twelve students. In very large samples the opposite problem appears: trivial DIF becomes statistically significant, which is why effect-size classification, not the p-value, must govern action.

  • DIF is not bias. This is the cardinal caveat. DIF is a statistical finding of non-equivalence; item bias is a substantive judgement that the non-equivalence is irrelevant to the construct and unfair. An item can show DIF for a legitimate reason — if online and in-person students genuinely differ on an item about "the physical learning environment," that is a real difference, not bias. Every flagged item requires human, content-expert review before you conclude anything is wrong.

  • The matching criterion can be contaminated. If the total score used for matching itself contains biased items, the DIF test inherits that bias. Purification (iterative removal of flagged items from the matching total) mitigates but does not eliminate this circularity.

  • Group definitions are researcher choices. DIF is always relative to a chosen grouping. You will only find gender DIF if you look for it; an unexamined axis (disability, mature students, commuter vs residential) stays invisible. DIF disciplines your fairness claims to the groups you actually tested.

  • Ordinal data needs the right model. Course-evaluation items are ordinal Likert responses, so the cleanest DIF analyses use ordinal logistic regression or IRT-based methods rather than treating a 1–5 scale as continuous — the same argument made for cumulative-link models in ordinary reporting.

How Koji incorporates this

Koji is designed to make item-level fairness auditing tractable rather than aspirational.

  • Structured, reusable item banks. Because Koji collects responses against defined question objects (scale, single_choice, yes_no, open_ended) with stable identifiers across cohorts and semesters, the pooled, matched datasets that DIF analysis requires are available by design — not reconstructed by hand from PDFs. An item that is flagged in one term can be tracked, revised, and re-tested in the next.

  • Bias-aware reporting rather than raw league tables. Koji's reporting layer is built to surface matched comparisons and to caveat small-sample and cross-group differences, discouraging the naive between-group averages that DIF exists to correct. This complements the platform's broader approach to fair comparison — see empirical-Bayes shrinkage and funnel plots.

  • AI-moderated conversational interviews to explain the why. DIF tells you an item behaves differently; it does not tell you why. Koji's AI-moderated interviews probe beyond the Likert number with adaptive follow-up questions, so when an item is flagged you can review the open-text and conversational context that reveals how different groups are interpreting it — the qualitative evidence a content-expert review needs to decide whether the DIF is genuine bias or a real difference.

  • Automatic thematic analysis of that open text helps a QA officer see, at a glance, whether (for example) non-native speakers systematically read a "clarity" item as a comment on accent rather than on teaching structure.

Koji is careful not to overclaim here: no platform eliminates item bias. What Koji is designed to do is make the detect-diagnose-revise-retest loop routine, so biased items are caught and fixed rather than quietly shaping personnel decisions for years. (Koji's core research platform at koji.so applies the same AI-moderated interview engine to product and customer research, where DIF-style equivalence questions arise whenever you compare segments.)

Related Resources

References

  • Swaminathan, H., & Rogers, H. J. (1990). Detecting differential item functioning using logistic regression procedures. Journal of Educational Measurement, 27(4), 361–370. https://doi.org/10.1111/j.1745-3984.1990.tb00754.x
  • Holland, P. W., & Thayer, D. T. (1988). Differential item performance and the Mantel-Haenszel procedure. In H. Wainer & H. I. Braun (Eds.), Test Validity (pp. 129–145). Lawrence Erlbaum.
  • Rogers, H. J., & Swaminathan, H. (1993). A comparison of logistic regression and Mantel-Haenszel procedures for detecting differential item functioning. Applied Psychological Measurement, 17(2), 105–116. https://doi.org/10.1177/014662169301700201
  • Marsh, H. W., & Hocevar, D. (1984). The factorial invariance of student evaluations of college teaching. American Educational Research Journal, 21(2), 341–366. https://doi.org/10.3102/00028312021002341
  • Zumbo, B. D. (1999). A handbook on the theory and methods of differential item functioning (DIF). Directorate of Human Resources Research and Evaluation, Department of National Defense, Ottawa.

Related articles