New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods9 min read

Why Quantitative Courses Get Lower Evaluations: The Discipline-Bias Problem

Uttl & Smibert (2017) show that instructors of quantitative courses receive systematically lower student evaluations than those teaching qualitative subjects — a bias with real career consequences. What the evidence says and how to compare ratings fairly across disciplines.

Koji Education Team

Product

In brief: Student evaluations of teaching are not comparable across disciplines. Uttl and Smibert (2017), analysing nearly 15,000 class evaluations from over 325,000 students, found that instructors of quantitative courses (mathematics, statistics, hard sciences) receive systematically lower ratings than instructors of qualitative courses (English, humanities), even when teaching equally well. The gap is large enough to flip a professor from "satisfactory" to "unsatisfactory" under common cut-off rules — with consequences for tenure, promotion and pay. The practical implication is unambiguous: never rank or threshold instructors on raw cross-disciplinary scores. Benchmark within discipline, and read scores alongside qualitative evidence.

The question, and why it matters

Most institutions compute a single satisfaction number per course and then, explicitly or implicitly, compare it against a campus-wide benchmark: "the faculty average is 4.1; this instructor is at 3.6, so there is a problem." That comparison silently assumes that a 3.6 in calculus means the same thing as a 3.6 in creative writing. If it does not — if some disciplines simply attract lower ratings for reasons unrelated to teaching quality — then every cross-disciplinary comparison, every league table, every "below-average flag" is contaminated. Because evaluations feed real personnel decisions, the question is not academic: it determines whether a competent mathematics lecturer is unfairly marked as failing.

What the research says

The anchor study is Bob Uttl and Dylan Smibert's 2017 paper, "Student evaluations of teaching: teaching quantitative courses can be hazardous to one's career" (PeerJ, 5, e3299). Using 14,872 publicly posted class evaluations representing more than 325,000 students, the authors examined how class subject relates to SET ratings. Their findings are stark:

  • Subject area is strongly associated with ratings. Quantitative classes (e.g., mathematics) received systematically lower SETs than non-quantitative classes (e.g., English), independent of teaching quality.
  • The effect is large enough to change categorical labels. When professors are classified as "satisfactory" vs. "unsatisfactory" or "excellent" vs. "non-excellent" using typical thresholds, those teaching quantitative courses are at substantially higher risk of falling below the line.
  • The consequence is career risk. Because thresholded SETs feed reappointment, promotion, tenure and merit-pay decisions, instructors of quantitative subjects face a higher chance of being labelled deficient — hence the paper's title.

Importantly, the authors stress that the criterion used to classify professors matters enormously: the same underlying ratings produce very different "failure" rates depending on where the cut-off sits, which means the discipline penalty interacts with arbitrary administrative thresholds to amplify unfairness.

This is not an isolated result. A 2022 Spanish replication in PeerJ — "Spain is not different: teaching quantitative courses can also be hazardous to one's career (at least in undergraduate courses)" (e13456) — reproduced the core pattern in a different national and linguistic context, strengthening the claim that the discipline gap is not merely an artefact of one country's rating culture. More broadly, Centra (2009) documented systematic differences in Student Instructional Report responses across disciplines, and Troy Heffernan's 2022 synthesis, "Sexism, racism, prejudice, and bias" (Assessment & Evaluation in Higher Education, 47(1), 144–154), reviewing 136+ publications, lists subject and discipline among the well-established sources of SET bias alongside gender and race. The convergence across methods and countries is what gives the finding its weight.

Why it matters for course evaluation in practice

The operational consequence is simple to state and frequently ignored: raw SET scores cannot be compared across disciplines, and certainly cannot be thresholded against a single institutional benchmark. A quality-assurance process that flags any instructor below the campus mean will disproportionately flag those teaching demanding quantitative material — not because they teach worse, but because the instrument is biased against their field.

Fair practice requires, at minimum:

  1. Within-discipline benchmarking. Compare a statistics lecturer against statistics norms, not against the whole university.
  2. Abandoning single hard thresholds for personnel decisions, or at least setting them per discipline and treating them as one input among many.
  3. Triangulation with peer observation, learning outcomes and qualitative evidence, so that no career decision rests on a discipline-contaminated number.

There is also a cultural cost to ignoring this. If able quantitative teachers learn that rigour is punished by the evaluation system, the rational response is to dilute content — exactly the perverse incentive that undermines the disciplines students most need.

Limitations and honest caveats

A critical reader should hold several reservations.

  • Source of the data. The Uttl & Smibert dataset is drawn from publicly posted, self-selected online ratings, which are subject to selection bias — students who choose to post may differ from a full cohort. The headline magnitudes may not transfer to mandatory institutional evaluations.
  • Confounding. Discipline correlates with class size, gender composition of faculty, grading distributions and workload, all of which independently affect ratings. Disentangling a pure "subject" effect from these correlates is difficult, and some of the apparent discipline penalty may operate through those channels.
  • What "lower" means. A lower mean does not by itself prove the teaching is equally good; in principle quantitative courses could be taught less well on average. The bias interpretation rests on the assumption — well supported but not proven — that teaching quality is comparable across fields.
  • Generalisability of cut-offs. The career-risk claim depends on institutions actually using rigid thresholds. Where SETs are used formatively and holistically, the harm is smaller.

These caveats narrow the claim but do not dissolve it: across multiple datasets and two countries, the direction of the effect is consistent, and prudent governance should assume it is real.

How Koji incorporates this

Koji is built around the premise that a bare cross-disciplinary average is an unsafe basis for judgement — which is exactly what the discipline-bias literature demonstrates. Several mechanisms map directly to the problem:

  • Within-cohort and within-discipline reporting. Koji is designed to present results against relevant comparison groups rather than a single institutional mean, so a quantitative course is read against its peers, not against the humanities. This directly counters the "below campus average" trap.
  • Bias-aware reporting that flags when a comparison crosses disciplines, discouraging the apples-to-oranges thresholding that the research shows penalises quantitative teaching.
  • AI-moderated conversational interviews that gather reasons, not just scores. When a calculus cohort rates a course lower, Koji's interview can probe whether the dissatisfaction concerns teaching, pace, the inherent difficulty of the material, or assessment — separating signal from the structural discipline penalty.
  • Triangulation support through structured question types (scale, single_choice, ranking) combined with open_ended probing and automatic thematic analysis, so programme leaders can corroborate or contextualise a number rather than act on it alone.

The honest framing is that Koji is designed to mitigate the misuse of discipline-contaminated scores — by reframing comparisons and enriching them with qualitative evidence — not to magically equalise the underlying rating tendencies of different student populations. No instrument can make a raw quantitative-course mean directly comparable to a humanities mean; what Koji can do is stop your governance process from pretending they already are. The same AI-moderated engine powers general user and customer research at koji.so, where comparing scores across unlike segments raises the identical methodological hazard.

A worked example of within-discipline benchmarking

Consider two lecturers at the same faculty. Lecturer A teaches second-year statistics and scores 3.7; Lecturer B teaches a literature seminar and scores 4.3. Against a single institutional mean of 4.1, A is flagged "below average" and B is praised — and if a rigid 3.8 threshold governs reappointment, A is now formally at risk. The Uttl & Smibert evidence says this comparison is invalid: the half-point gap is well within the range the discipline penalty alone can produce. The defensible procedure is to rank A against the distribution of other quantitative courses and B against other seminar courses. Done that way, A may sit comfortably in the upper half of statistics teaching while B is merely typical for seminars — the opposite of the naive reading. Institutions that cannot compute within-discipline norms because of small numbers should, at minimum, suppress cross-disciplinary thresholds entirely and route any borderline case to peer review rather than to an automated flag.

Related Resources

References

  • Uttl, B., & Smibert, D. (2017). Student evaluations of teaching: teaching quantitative courses can be hazardous to one's career. PeerJ, 5, e3299. https://doi.org/10.7717/peerj.3299
  • Morales-Vives, F., et al. (2022). Spain is not different: teaching quantitative courses can also be hazardous to one's career (at least in undergraduate courses). PeerJ, 10, e13456. https://doi.org/10.7717/peerj.13456
  • Centra, J. A. (2009). Differences in responses to the Student Instructional Report: Is it bias? Educational Testing Service / Research in Higher Education context. https://doi.org/10.1007/s11162-008-9107-6
  • Heffernan, T. (2022). Sexism, racism, prejudice, and bias: a literature review and synthesis of research surrounding student evaluations of courses and teaching. Assessment & Evaluation in Higher Education, 47(1), 144–154. https://doi.org/10.1080/02602938.2021.1888075