New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Methodology9 min read

How to Run an Internal Bias Audit of Your Course Evaluation Data

Stop arguing about whether SET bias findings apply to your institution — measure it. A step-by-step methodology for auditing your own course evaluation data for demographic bias, with the statistical traps marked.

Koji Education Team

Product ·

The student-evaluation bias literature is large, contested, and mostly about other people's universities. Whether the penalties documented in North American experiments transfer to your institution is ultimately an empirical question about your data — and most institutions already hold everything needed to answer it: evaluation records, course metadata, and instructor characteristics from HR. This is a step-by-step methodology for an internal bias audit — descriptive gaps, multilevel models with the right controls, item-level checks, comment analysis, and robustness tests — along with the statistical and governance traps that turn well-meaning audits into misleading ones.

Why audit at all?

Three reasons. First, the external evidence is genuinely mixed and context-dependent: causal studies such as MacNell, Driscoll and Hunt (2015), Boring (2017) and Mengel, Sauermann and Zölitz (2019) found gender penalties in specific settings, while other samples show small or null effects — and transferability to European contexts is an open question. Second, if scores feed reappointment, tenure or promotion, the institution owns the fairness risk regardless of what the literature says. Third, several European equality bodies and courts have begun treating unexamined evaluation systems as a liability; "we never checked" is a poor defense.

Step 0: Governance before statistics

Decide before touching data who commissions the audit, who sees results at what granularity, and what decision rules follow from findings. An audit whose consequences are undefined becomes either a drawer report or a witch hunt. Pre-specify the analysis plan — outcomes, comparisons, controls, thresholds — and log it internally. This is the same discipline that specification-curve analysis formalizes: decide the garden of forking paths in advance, or wander it unconsciously.

Under GDPR, instructor gender, age band and nationality from HR records are personal data being repurposed; run this through your DPO with a documented legitimate-interest or public-task basis. Aggregate reporting must respect small-cell disclosure rules — a "bias dashboard" that lets colleagues reverse-engineer one colleague's scores is its own harm.

Step 1: Assemble the analytic dataset

One row per evaluation response (not per course), with:

  • Outcome variables: overall rating; each item rating; response-level text features (length, specificity) if available.
  • Instructor attributes: gender, age band, career stage/contract type, nationality or language background where lawful and available.
  • Course controls: discipline, level, class size, required vs elective, modality, grading distribution, time of day.
  • Student attributes (if linkable): gender, year, programme — many biases are interactions between rater and rated, as the student-rater gender evidence shows.
  • Response metadata: response rate per course, timing of completion.

Three years of data is a reasonable minimum; you need repeated observations per instructor for the models below.

Step 2: Descriptives first — but do not stop there

Plot full rating distributions (not means) by instructor group, within discipline. Raw gaps are where every internal conversation will start, so report them — but label them clearly as unadjusted. A raw gap can reflect bias, or the fact that women are overrepresented in large required first-year courses that everyone rates lower; required-course penalties and discipline norms confound naive comparisons in both directions.

Step 3: Model the gap properly

The workhorse is a multilevel model — responses nested in course offerings nested in instructors — with course-type controls. Key design choices:

  • Control for course characteristics, not for "instructor quality." Class size, level, electivity, discipline: yes. Anything downstream of the bias you are testing (e.g., prior-year scores): no.
  • Test interactions. Bias findings are frequently conditional: junior women but not senior women in Mengel et al.; rater-gender × instructor-gender cells; minority instructors in some disciplines only. Main effects alone can average a real penalty into invisibility.
  • Report uncertainty honestly. With hundreds of instructors, trivially small gaps become "significant." Report effect sizes against a meaningful benchmark — for instance, what a 0.3 difference actually means on your scale — not asterisks.

If your instrument has multiple items, add measurement-invariance / differential item functioning checks: a gap concentrated in "enthusiasm" and "authority" items but absent in "returned feedback on time" is a fingerprint of stereotype-driven rating, and far more informative than a global gap.

Step 4: Audit the comments, not just the numbers

Numeric gaps understate what open-text feedback reveals. Analyze comment corpora for: volume and length by instructor group; theme distribution (are women's comments disproportionately about persona and appearance rather than course substance — the pattern found in gendered-language studies?); and incidence of abusive or identity-referencing remarks. Manual coding does not scale past a few hundred comments; LLM-assisted thematic analysis does, provided you validate a sample against human coders and watch for automation bias in how committees consume the summaries.

Step 5: Stress-test your own findings

Before any result leaves the analyst's laptop:

  • Multiverse it. Re-run across reasonable alternative specifications (controls in/out, samples, outcome codings). A gap that appears in 9 of 200 specifications is noise; one that survives most of them is signal.
  • Interrogate non-response. With typical online response rates, who responds is not random; differential non-response across student groups can manufacture or mask gaps.
  • Remember the ecological trap. Institution-level nulls can hide department-level penalties and vice versa — aggregation is not innocent.

Step 6: Decide what follows

An audit is only as valuable as its consequence rules. Reasonable responses scale with findings: adding uncertainty bands and context notes to all score reports; removing global items from personnel files in favor of behaviorally specific ones; comparative-context statements for committees; instrument redesign; or — the option too few institutions consider — demoting numeric averages from personnel evidence altogether and rebuilding evaluation around substantiated qualitative evidence.

But isn't an observational audit unable to prove bias?

Correct — and worth stating plainly. Without random assignment you cannot cleanly separate "students penalize women" from "women systematically teach less-rateable courses in unmeasured ways." Randomized and natural experiments (MacNell's perceived-gender design; Boring's and Mengel's random assignment of students to sections) exist precisely because observational gaps are ambiguous. But the audit's purpose is not to publish a causal paper; it is institutional risk assessment. A robust adjusted gap in your own data — whatever its ultimate cause — means your measurement system produces systematically different career evidence for different demographic groups, which is a problem on any causal story. And a null, honestly obtained, is genuinely reassuring information you currently do not have.

The second honest objection: audits can become theater — an annual PDF nobody acts on. That is a governance failure, not a measurement one, and Step 0 exists because of it.

Where Koji fits

An audit tells you where your current instrument leaks; Koji for Education is built to leak less and to make the ongoing monitoring routine.

  • Evidence-anchored data collection. Koji's AI-moderated conversational interviews probe ratings for the specific episode behind them, producing feedback that committees can weigh on substance — and its six structured question types support the low-inference, behaviorally specific item design that DIF-style audits consistently recommend.
  • Thematic analysis at audit scale. The comment-audit of Step 4 is a built-in output rather than a bespoke NLP project: themes, quality scores and abusive-content flags across every course, every cycle.
  • Programme- and institution-level reporting with appropriate aggregation makes the monitoring loop — not just the one-off audit — sustainable, and GDPR/AVG-compliant EU data handling keeps the DPO conversation short.

Institutions running similar audits on employee or user research data will find the same engine on the main Koji platform.

The takeaway

You do not need to referee the SET bias literature; you need to know what your own instrument is doing. Pre-specify, model the nesting, test interactions and item-level patterns, audit the comments, stress-test the findings, and attach consequences before you start. Whatever the result, you will make better personnel and quality decisions than an institution still arguing from other people's data.

Koji for Education gives quality teams evaluation data that is auditable by design — probed, thematically analyzed, and reported with the uncertainty visible. Book a demo.