New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Sector trends9 min read

Does AI-Moderated Course Evaluation Just Introduce Algorithmic Bias?

Large language models demonstrably reproduce racial, dialect, and gender stereotypes. So does using AI to moderate and analyse student feedback simply swap human bias for machine bias? An honest look at the evidence — and the guardrails that decide the answer.

Koji Education Team

Product · June 11, 2026

The short answer: Large language models do carry measurable bias, so an AI evaluation tool that is built carelessly can absolutely import it. But the relevant comparison is not "AI versus a neutral ideal" — it is "AI versus the human-moderated, average-a-Likert-score status quo," which is itself riddled with well-documented bias. Whether AI helps or harms depends entirely on where the model is allowed to act, what it is allowed to see, and how its outputs are audited. This piece sets out the evidence on both sides and the specific design choices that separate a bias-amplifying tool from a bias-aware one.

The objection, stated at full strength

Anyone serious about course evaluation should take this concern seriously, because the evidence behind it is real. In a 2024 study published in Nature, Hofmann and colleagues showed that leading language models make covertly racist decisions about people based on dialect alone — judging speakers of African American English as less employable and assigning them more menial jobs and harsher criminal sentences, even while the same models endorsed positive explicit stereotypes about the same group (Hofmann et al., 2024, Nature / PMC). A large multi-model annotation study found that the same sentence was rated significantly less professional (mean gap −0.774 across 19 models) and more "angry" when written in African American Vernacular English rather than Standard American English (LLMs Reproduce Racial Stereotypes in Text Annotation, 2026). And work in PNAS in 2025 demonstrated that even explicitly unbiased LLMs still form biased implicit associations — the surface politeness produced by human-feedback training does not remove the underlying patterns (Bai et al., PNAS 2025).

Apply that to course evaluation and the worry is obvious. If a model summarises open-text comments, scores "quality," or clusters themes, it could systematically under-weight feedback written by international students in non-standard English, or echo the gendered language ("warm," "bossy," "easy") that already distorts how students rate women and minoritised faculty. We have written before about how strong that human bias already is in gender and racial and ethnic ratings. The fear is that AI laminates a new, harder-to-see layer on top.

But compare it to the right baseline

The rhetorical trick in "AI just adds bias" is the implied alternative: a clean, neutral traditional survey. That alternative does not exist. The conventional student-evaluation-of-teaching (SET) instrument — a five-point Likert form, mean-aggregated and compared across instructors — is one of the most thoroughly impeached measures in higher education. Boring, Ottoboni and Stark (2016) showed SET scores track instructor gender and perceived warmth more than teaching effectiveness; Uttl, White and Gonzalez's 2017 meta-analysis found SET ratings explain at most 1% of the variance in actual student learning (Uttl et al., 2017, Studies in Educational Evaluation). Human-moderated focus groups add interviewer effects, social desirability, and the simple fact that no two human moderators probe the same way.

So the honest question is not "is AI perfectly fair?" — nothing here is — but "does a well-governed AI process surface a more representative, less idiosyncratic picture than the status quo it replaces?" On several dimensions it plausibly does: an AI moderator asks every student the same calibrated follow-up rather than warming to some and not others; it can be instructed to focus thematic analysis on what was said about the course and away from comments about an instructor's appearance or accent; and unlike a Likert mean, it leaves an auditable trail of exactly which comments drove which theme.

What separates a bias-aware tool from a bias-amplifying one

The difference is not the word "AI." It is governance. Five design choices do most of the work:

  1. Constrain the decision surface. The most dangerous use of an LLM is letting it assign a high-stakes score to a person. The safest is using it to organise and surface what humans said, leaving judgement to humans. A tool that converts comments into searchable themes carries far less risk than one that outputs "instructor quality: 6.2/10."
  2. Standardise the probe, not the person. AI moderation should be calibrated to ask the same evidence-seeking follow-ups of every respondent ("Can you give a specific example?"), which reduces the human-moderator inconsistency that itself is a bias source — without inferring anything from the respondent's demographics.
  3. Keep demographic cues out of the analysis path. Much LLM bias is triggered by names, dialect markers, or identity cues in the text. Thematic analysis should be run on de-identified content and instructed to ignore commentary on protected characteristics.
  4. Audit outputs against known bias patterns. Because bias is measurable, it can be tested for: run matched-text audits (the same comment in standard vs. non-standard English) and check that theme assignment and quality scores do not move. This is exactly the audit-style evaluation the research literature now recommends.
  5. Preserve the raw evidence. Aggregation hides bias; traceability exposes it. Every summary should link back to verbatim source comments so a quality officer can check whether the machine read them fairly.

How Koji approaches it

Koji for Education is built around that governance, and we are deliberately precise about what it does and does not claim. Koji does not assign teaching-quality scores to individuals or make automated personnel decisions — a stance that also keeps it clear of the EU AI Act's high-risk provisions for educational evaluation. What it does is run AI-moderated conversational interviews that ask every student the same calibrated, evidence-seeking follow-ups; apply automatic thematic analysis to open-text feedback so that themes are tied back to the verbatim comments that produced them; and standardise moderation so the inconsistency of human focus-group facilitators is removed. Analysis runs on the substance of what students say about the course, and the platform is designed for GDPR/AVG-compliant, EU-appropriate data handling. We say Koji mitigates and surfaces bias and inconsistency — not that it eliminates bias, which no system can.

The same conversational interview engine powers the main Koji platform for general user and customer research, where the identical governance questions — standardised probing, de-identified analysis, traceable themes — apply.

A due-diligence checklist for buyers

Because the answer to "does this tool add bias?" lives in the governance rather than the brochure, the burden is on buyers to interrogate it. A quality-assurance officer evaluating any AI-assisted evaluation product should ask, and expect documented answers to, the following:

  • Where does the model make a decision? Ask for an explicit map of every point where the AI scores, ranks, or classifies. Be most wary anywhere it outputs a judgement about an individual instructor.
  • What does the model see? Is analysis run on de-identified text, and are demographic or dialect cues stripped before processing? If the model ingests names and identity markers, it can act on them.
  • Has it been audited for bias, and can you see the results? A credible vendor will have run matched-text tests (the same comment in standard and non-standard English) and will share what moved and what did not.
  • Is there a human in the loop for anything consequential? No high-stakes decision about a person should rest on an unreviewed model output.
  • Can every summary be traced to source comments? If you cannot click from a theme back to the verbatim feedback behind it, you cannot check the machine's reading — and neither can the faculty member affected.
  • How is drift monitored? Models change; a tool audited once is not audited forever. Ask how bias is re-tested over time.

A vendor who cannot answer these is asking for trust it has not earned. A vendor who can has done the work that separates a bias-aware instrument from a bias-amplifying one.

The bottom line

"AI just introduces algorithmic bias" is a fair warning and a lazy verdict. The warning is correct: unconstrained models reproduce documented stereotypes, and any vendor who waves that away should not be trusted. The verdict is lazy because it measures AI against a fictional neutral baseline instead of the demonstrably biased instrument it replaces. A carelessly built AI tool can be worse than a Likert form. A carefully governed one — narrow decision surface, standardised probing, de-identified and audited analysis, full traceability — can surface a fairer, richer, more representative picture of the student experience than averaging a number ever could. The technology is not the answer to the bias question. The guardrails are.