New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Evaluation bias10 min read

Your Gender Equality Plan vs Your Student Evaluations

To take Horizon Europe money, your institution signed a Gender Equality Plan committing it to remove structural bias against women. Then it kept using student evaluations — one of the best-documented gender-biased instruments in higher education — as evidence in promotion. Those two facts are in tension.

Koji Education Team

Product · August 12, 2026

Bottom line up front: Since 2022, having a Gender Equality Plan (GEP) has been an eligibility criterion for Horizon Europe funding for public bodies and higher education institutions across the EU and associated countries. A GEP is a signed, published commitment to identify and remove structural gender bias in your institution — including in recruitment and career progression. Meanwhile, most of those same institutions still use student evaluations of teaching (SET) as evidence in promotion and probation decisions. SET is one of the most thoroughly documented gender-biased instruments in higher education. Using it as career-shaping evidence sits in direct tension with the GEP you signed to get the grant money — and the GEP's own monitoring requirement is what will expose the contradiction.

The commitment you already made

The Horizon Europe GEP eligibility criterion applies to all public bodies, higher education institutions and research organisations in Member States and associated countries applying to calls with deadlines from 2022 onwards. To comply, an institution's GEP must meet four mandatory process requirements, confirmed by the European Institute for Gender Equality:

  1. Publication — a formal document, signed by top management and published on the institution's website.
  2. Dedicated resources — expertise and funding to implement it.
  3. Data collection and monitoring — sex/gender-disaggregated data on staff and students, with annual reporting.
  4. Training — awareness-raising and capacity-building on gender equality and unconscious bias.

Recruitment and career progression are among the recommended thematic areas a GEP should address. So this is not a document about abstract research culture; it is a commitment that reaches directly into how your institution hires, promotes and tenures. And promotion is exactly where SET does its damage.

What the evidence on SET gender bias actually says

The finding is not a single contested study; it is a convergent literature built on strong designs.

  • MacNell, Driscoll and Hunt (2015) ran an online course where instructors presented under both a male and a female identity to otherwise-equivalent groups. The same person received significantly lower ratings when students believed they were female — the cleanest possible demonstration that the bias is about perceived gender, not teaching.
  • Boring (2017), analysing over 20,000 evaluations at Sciences Po where students were effectively randomly assigned across sections, found male instructors received higher ratings, with much of the effect driven by male students, despite no corresponding advantage in actual learning outcomes.
  • Mengel, Sauermann and Zölitz (2019), a field experiment at Maastricht University with random assignment to instructors, found women received systematically lower evaluations — again most strongly from male students, and unrelated to the students' subsequent performance.

We have unpacked this body of work — including who gives the lower scores, the core gender-bias evidence, how strong that evidence really is, and whether the US findings transfer to Europe. The short version: the bias is real, it is measurable, and it compounds with race, accent and other characteristics in intersectional ways.

The contradiction, stated plainly

Put the two facts side by side. Your GEP commits you, in a document signed by your rector, to identify and remove structural bias against women in career progression. Your promotion process, in many institutions, still treats SET averages as evidence of teaching quality. SET averages carry a documented gender penalty that has nothing to do with teaching quality. Therefore your promotion process imports, into a career-defining decision, precisely the kind of structural bias your GEP promises to remove. We have argued separately that leaning on SET in tenure and promotion is hard to defend on the evidence; the GEP turns that from a methodological concern into an institutional-commitment problem.

And here is the sharp edge: requirement 3 obliges you to look. A compliant GEP must collect sex-disaggregated data and monitor. The moment an institution disaggregates its SET scores by instructor gender — which its own GEP tells it to do — it is likely to find the gap the literature predicts. The GEP is simultaneously the exposure and the remedy.

But isn't this overstated? Three honest objections

"The bias is small or contested." Some studies find smaller effects, and effect sizes vary by discipline and context; the literature is not perfectly unanimous. But two things hold. First, the strongest designs — random assignment, identity-swap experiments — are the ones that find the effect most clearly, which is the opposite of what you would expect if it were an artefact. Second, promotion decisions are made at the margin, and a small systematic penalty applied at a threshold changes who gets promoted. "Small on average" is not "harmless at the cut-score."

"Just statistically adjust the scores for gender." Tempting, and we have examined whether you should correct SET scores for bias. The problem is that adjustment requires you to model a bias whose size you do not know precisely, it can mask other confounds, and it still treats a fundamentally weak proxy as if it were sound once "corrected." Adjustment is a patch on an instrument that should not be load-bearing in the first place.

"The GEP is about research careers, not teaching evaluation." This misreads the criterion. Career progression is explicitly in scope, teaching evaluation is a documented gendered input to progression, and requirement 3's monitoring duty does not exempt SET. If anything, ignoring a known gendered input while claiming to monitor gender equality is the weaker position to defend — including under the public-sector equality duties that already expose biased evaluation to legal risk.

Monitoring is the criterion that closes the trap

It is worth dwelling on requirement 3, because it is the part institutions treat as paperwork and it is actually the sharpest instrument in the GEP. "Collect sex-disaggregated data and monitor annually" is not a request to count how many women you employ. Read seriously, it is a duty to look for structural bias in the mechanisms that shape careers — and teaching evaluation is one of those mechanisms. An institution that disaggregates every input to promotion except the one the literature specifically flags as gender-biased has not really complied; it has monitored around the problem. Conversely, an institution that does disaggregate its evaluation data will have, in hand, exactly the evidence its GEP action plan needs to justify changing how SET is weighted. The monitoring requirement is what turns a vague commitment into an obligation with teeth, and it is why "we didn''t know" stops being available as a defence once the first disaggregated report is run.

What Koji can and cannot do about it

Let us be scrupulous, because this is a domain where overclaiming is its own failure. No evaluation instrument can remove the bias that lives in students' own heads. A student who rates women more harshly will carry that bias into any channel. Koji does not "eliminate bias," and any vendor who says theirs does is not being honest.

What Koji for Education can do is reduce how much a biased signal is allowed to decide a career, and help you meet the GEP's monitoring duty:

  • Shift weight off the single global rating. SET bias concentrates in the overall "how good was this instructor" number. Koji's AI-moderated interviews generate specific, behaviourally-grounded evidence — what actually happened in the course — so committees can weigh described teaching practice rather than a gender-loaded gestalt score.
  • Standardised, bias-aware AI moderation. A consistent AI moderator removes one source of variation that human-moderator setups add on top of student bias. It does not remove student bias, but it does not compound it either.
  • Thematic analysis that makes bias visible. Automatic analysis of open-text feedback can surface gendered language patterns — the warmth-versus-competence framing the literature describes — turning your feedback corpus into monitoring evidence rather than an unexamined average. That directly serves GEP requirement 3.
  • Programme- and institution-level reporting disaggregated as your GEP requires, so monitoring is a standing capability rather than a one-off audit.
  • GDPR-compliant, EU-appropriate handling of what is, when disaggregated by gender, sensitive personnel data.

Framed honestly: Koji mitigates the institutional over-reliance on a biased number and surfaces the bias for monitoring; it does not eliminate the underlying student bias. That is exactly the distinction a PhD-level equality committee will want to hear, and exactly the one your GEP's language of monitoring and structural change already uses.

Institutions doing this work often run the same conversational engine — shared with koji.so — on staff-experience and inclusion research more broadly. If your GEP is signed and your promotion process still runs on SET averages, that gap is worth closing before your next monitoring report. See how Koji for Education approaches it.

Frequently asked questions

Is a Gender Equality Plan mandatory for Horizon Europe funding?

Yes. Having a Gender Equality Plan is an eligibility criterion for public bodies, higher education institutions and research organisations in EU Member States and associated countries applying to Horizon Europe calls with deadlines from 2022 onwards. The plan must be a formal document signed by top management and published on the institution's website, backed by dedicated resources, sex-disaggregated data collection and monitoring, and training on gender equality and unconscious bias. Private companies are exempt.

Are student evaluations of teaching really gender-biased?

The evidence is convergent and rests on strong designs. MacNell et al. (2015) found the same online instructor was rated lower under a female identity. Boring (2017), analysing over 20,000 evaluations at Sciences Po, found male instructors rated higher with no learning advantage. Mengel et al. (2019), a randomised field experiment at Maastricht, found women systematically rated lower, most strongly by male students and unrelated to student performance. The penalty is real and unrelated to teaching quality.

Why does the Gender Equality Plan matter for how we use course evaluations?

A GEP commits an institution to identify and remove structural gender bias, including in recruitment and career progression. Using gender-biased SET scores as promotion evidence imports exactly that bias into career decisions. The GEP's mandatory data-monitoring requirement also obliges institutions to collect sex-disaggregated data — so disaggregating SET by instructor gender, as the GEP implies, is likely to reveal the gap the literature predicts.

Can't we just statistically adjust evaluation scores for gender bias?

Adjustment is problematic. It requires modelling a bias whose exact size is unknown, can mask other confounds, and still treats a weak proxy for teaching quality as sound once corrected. It is a patch on an instrument that arguably should not be load-bearing in promotion at all. Reducing reliance on the single biased global score is more defensible than correcting it.

Can Koji eliminate gender bias in evaluations?

No, and it does not claim to. No instrument can remove bias that lives in students' own perceptions; a student who rates women more harshly carries that into any channel. What Koji can do is reduce how much a biased global score decides a career — by generating specific, behaviourally-grounded evidence committees can weigh instead — and surface gendered language patterns through thematic analysis to support the monitoring a Gender Equality Plan requires. It mitigates and surfaces; it does not eliminate.