Gendered Language in Student Comments: Why Men Are ''Brilliant'' and Women Are ''Caring''
Bias in course evaluations is not only in the numbers — it is in the words. What Mitchell & Martin (2018) and Storage et al. (2016) found about systematically different language applied to men and women, and why open-text analysis must account for it.
Koji Education Team
Product
In brief: Even when numeric ratings look similar, the language students use in open-text comments differs systematically by instructor gender. Mitchell and Martin (2018) found men described as "brilliant," "intelligent" and "knowledgeable," while women drew "caring," "nice," "helpful," "bossy" and were more often called "teacher" rather than "professor." Storage et al. (2016) showed the words "brilliant" and "genius" appear far more often for men across every field studied. This means qualitative analysis of comments — including automated analysis — can launder gender bias into apparently neutral themes unless it is explicitly designed to detect it.
What the research says
The quantitative gender-bias literature (MacNell, Driscoll and Hunt, 2015; Boring, 2017) established that students can assign systematically lower numeric scores to women instructors in controlled comparisons. A newer strand asks a different question: what do students write, and does the language itself encode bias even when the numbers do not?
Mitchell and Martin (2018), in PS: Political Science & Politics, combined two methods. First, they analysed the language of comments about male and female instructors and found students describe the two groups with different vocabularies: men attract competence-and-brilliance words ("brilliant," "intelligent," "knowledgeable," "funny"), while women attract warmth, personality and appearance words ("caring," "nice," "helpful") and disproportionately negative-personality terms ("bossy," "mean"). Women were also more likely to be labelled "teacher" and men "professor" — a status asymmetry embedded in word choice. Second, in a quasi-experiment in identical online courses, students rated the instructor presented as male more highly than the one presented as female, replicating the MacNell paradigm.
Storage, Horne, Cimpian and Leslie (2016), in PLOS ONE, scaled this up dramatically. Mining over 14 million RateMyProfessors.com reviews, they counted how often "brilliant" and "genius" appeared by field and instructor gender. The words were applied to men markedly more often than women in every one of the fields examined, and — the striking structural finding — fields whose evaluations emphasised "brilliance" had fewer women and African American PhDs. The language of raw intellectual talent tracks, and may help reproduce, who is seen as belonging in a discipline.
A wider corroborating literature reinforces the pattern. Work on role-congruity and the "double bind" (Eagly and Karau, 2002) explains why agentic, authoritative behaviour that earns men "strong" earns women "bossy." And large-scale text analyses of evaluations (e.g., the widely discussed gendered-language visualisations of RateMyProfessors data) repeatedly surface the same competence-for-men / warmth-for-women split. The convergence of an experiment (Mitchell and Martin), a 14-million-review corpus (Storage et al.) and social-psychological theory (Eagly and Karau) makes this one of the better-triangulated findings in the SET field.
Why it matters for course evaluation in practice
If your quality process reads open-text comments — and most do — gendered language has concrete, sometimes invisible, consequences.
1. "Clean" numbers do not mean clean evaluations. A woman instructor whose mean rating matches a male colleague''s may still be receiving comments that frame her as warm-but-not-brilliant, or authoritative-as-bossy. A committee skimming verbatim comments absorbs that framing even when the scores are equal. The bias has simply moved from the number to the narrative.
2. Automated text analysis can amplify, not remove, the bias. A naïve sentiment or theme model trained to surface "enthusiastic," "caring," or "authoritative" descriptors will faithfully reproduce the gendered distribution of those words — handing committees a tidy, scientific-looking summary that encodes the very stereotype it should flag. This is a serious risk as institutions adopt AI to summarise feedback (see interpreting and reporting student ratings responsibly).
3. Status words shape high-stakes readings. Whether an instructor is consistently called "professor" or "teacher," "brilliant" or "nice," subtly cues seniority and competence judgements in promotion and renewal decisions. Because "brilliance" language also tracks who is under-represented in a field (Storage et al.), uncritically rewarding it in evaluation summaries can entrench existing demographic skews.
The practical imperative is to analyse comment language for differential framing, not just comment topic or sentiment — and to keep a human, bias-aware reader between the raw text and any personnel decision.
Limitations and honest caveats
A careful reader should hold several qualifications.
- Much of the corpus is RateMyProfessors, not institutional data. Storage et al. and many language analyses draw on a self-selected, anonymous, US, public-rating platform whose commenters and norms differ from a confidential institutional survey. The direction of the effect is robust; the magnitudes may not transfer to a European, in-course evaluation.
- Correlation between "brilliance" language and demographics is not causation. Storage et al. show fields emphasising genius have fewer women and Black PhDs; they cannot prove the language causes under-representation rather than co-evolving with disciplinary culture. The authors are explicit about this.
- Quasi-experiments have constraints. The MacNell-style identical-course designs that Mitchell and Martin replicate use small samples and online settings; their internal validity is strong but their generalisability to large in-person cohorts is debated, and some replications find weaker or context-dependent effects.
- Language is culturally specific. Gendered connotations of "bossy," "caring" or "brilliant" are English-language and Anglo-American. In multilingual European evaluation, the form of bias may differ, so importing an English word-list wholesale risks both false positives and missed bias.
- Not every difference is bias. Men and women may, on average, teach differently or different subjects, so some lexical difference could reflect real variation in pedagogy or discipline rather than pure stereotype. Disentangling the two requires careful controls that headline figures rarely include.
The defensible reading: gendered framing in comment language is real, well-triangulated and consequential, but its precise size is context-dependent and it should prompt scrutiny of your own data rather than a mechanical correction.
How Koji incorporates this
Koji treats open-text analysis as a place where bias can hide, and designs its feedback collection and analysis to surface rather than launder differential framing.
- Thematic analysis that separates content from characterisation. Rather than collapsing comments into a single sentiment score, Koji''s automatic thematic analysis groups feedback by what it is about (clarity, assessment, organisation) and is designed to distinguish substantive teaching feedback from personality- and appearance-focused characterisation — exactly the warmth-vs-competence split Mitchell and Martin documented. The intent is to help QA staff see when an instructor is being judged on different criteria, not to silently fold that language into a tidy theme.
- Probing for specifics over impressions. Koji''s AI-moderated conversational interview is designed to push beyond a thin descriptor ("she''s nice" / "he''s brilliant") toward concrete, behaviour-anchored examples — the low-inference, observable evidence that is harder to contaminate with stereotype than an unprompted adjective (see low-inference teaching behaviours).
- Bias-aware reporting and human-in-the-loop review. In line with the automation-bias risk, Koji frames AI-generated summaries as assistive drafts for human review, not verdicts, and reporting is designed to flag differential language and score patterns across cohorts so reviewers can investigate rather than rubber-stamp. We describe these as mechanisms designed to mitigate gendered framing — they do not eliminate a bias rooted in the wider culture.
- Disclosure and accountability. Because the evidence shows automated analysis can reproduce bias, Koji''s approach emphasises transparency about where AI is used in moderation and analysis, consistent with emerging expectations under the EU AI Act for evaluation tools that inform personnel decisions.
The same conversational and thematic-analysis engine underpins Koji''s core research platform at koji.so, where detecting how language about a product or person differs across respondent groups is a routine analytic task.
The bottom line for practice
The uncomfortable implication of this literature is that anonymising scores and equalising means does not, on its own, produce a fair evaluation. As long as committees read verbatim comments — and they should — the framing embedded in everyday words travels with the text. "Brilliant" and "caring" are not neutral synonyms for "good"; they sort instructors into competence and warmth in gendered ways. A defensible process therefore audits the language, not just the numbers: it asks whether men and women, and majority and minority instructors, are being praised and criticised on the same terms, and it keeps a human reader accountable for any inference that feeds a personnel decision.
Related resources
- Gender bias in student evaluations of teaching
- Text analytics for open-ended student comments
- Racial and ethnic bias in student evaluations
- Abusive and non-constructive comments and faculty wellbeing
- Interpreting and reporting student ratings responsibly
- Low-inference teaching behaviours in evaluation items
References
- Mitchell, K. M. W., & Martin, J. (2018). Gender bias in student evaluations. PS: Political Science & Politics, 51(3), 648–652. https://doi.org/10.1017/S104909651800001X
- Storage, D., Horne, Z., Cimpian, A., & Leslie, S.-J. (2016). The frequency of "brilliant" and "genius" in teaching evaluations predicts the representation of women and African Americans across fields. PLOS ONE, 11(3), e0150194. https://doi.org/10.1371/journal.pone.0150194
- MacNell, L., Driscoll, A., & Hunt, A. N. (2015). What''s in a name: Exposing gender bias in student ratings of teaching. Innovative Higher Education, 40(4), 291–303. https://doi.org/10.1007/s10755-014-9313-4
- Eagly, A. H., & Karau, S. J. (2002). Role congruity theory of prejudice toward female leaders. Psychological Review, 109(3), 573–598. https://doi.org/10.1037/0033-295X.109.3.573
- Boring, A. (2017). Gender biases in student evaluations of teaching. Journal of Public Economics, 145, 27–41. https://doi.org/10.1016/j.jpubeco.2016.11.006
Related articles
Text Analytics for Open-Ended Student Comments: What NLP Can and Cannot Tell You
Natural-language processing can turn thousands of free-text course-evaluation comments into themes and sentiment at scale — but the research (Cunningham-Nelson 2019; Sunar & Khalid 2023) shows where automated analysis is reliable and where human judgement is still required.
Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees
The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.
Abusive Open-Text Comments in Course Evaluations: The Evidence and the Duty of Care
Lakeman et al. (2022) found that 91% of surveyed Australian academics received non-constructive anonymous comments — insults, threats, remarks on appearance. A research-grounded look at abusive student feedback, its effect on staff wellbeing, and how to moderate open text responsibly.
Gender Bias in Student Evaluations of Teaching: Evidence and Mitigation
A research-grounded synthesis of gender bias in student evaluations of teaching (SET), anchored to MacNell, Driscoll & Hunt (2015) and corroborated by Boring (2017) and Mengel, Sauermann & Zölitz (2019), with practical mitigations for European quality-assurance teams.