New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods9 min read

Do Student Evaluations Measure Teaching or the Instructor's Personality?

Research consistently finds that the Big Five personality traits of an instructor explain substantial variance in student evaluation of teaching (SET) scores, over and above grades and perceived learning. What that means for fair, valid course evaluation.

Koji Education Team

Product

In brief: A consistent body of research finds that a substantial share of the variance in student evaluation of teaching (SET) scores is explained by the instructor's personality — particularly extraversion, openness, agreeableness and conscientiousness — over and above the grades students expect and how much they believe they learned. Patrick (2011) found that personality predicted ratings beyond grades and perceived learning, and Clayson & Sheffet (2006) showed that personality impressions formed in under five minutes of contact correlated with end-of-term ratings strongly enough that, in their words, a rating form "could be replaced with a personality inventory of the instructor with little change in outcome." This does not make SET worthless — but it means a raw mean is partly a personality score, and it should never be read as a pure measure of instructional quality.

What the research says

For a course evaluation to be valid, the score has to move because teaching quality moves — not because of something irrelevant to teaching. In measurement terms (Messick, 1989), anything that systematically pushes scores around without reflecting the construct you intend to measure is construct-irrelevant variance. A growing literature identifies the instructor's personality as exactly this kind of contaminant.

Patrick (2011) examined whether the Big Five personality traits and expected grades relate to student ratings of teachers and courses at the college level. Writing in Assessment & Evaluation in Higher Education, she reported that extraversion, openness, agreeableness and conscientiousness were the traits favoured in instructors, whereas neuroticism was not. A significant correlation appeared between students' expected grades and their evaluation of the course, but not of the instructor; and once perceived amount of learning was taken into account, the grade effect on teacher ratings disappeared. Crucially, personality explained variance in both teacher and course evaluations over and above grades and perceived learning — i.e., personality was not merely a proxy for "I got a good grade so I rate highly."

Clayson & Sheffet (2006), in the Journal of Marketing Education, pushed the point further. They had students rate the instructor's agreeableness, creativity, conscientiousness, stability and extroversion. These personality ratings — some formed on the basis of only a few minutes of classroom contact — correlated significantly with conventional end-of-semester evaluations. Their provocative conclusion was that a student rating instrument "could be replaced with a personality inventory of the instructor with little change in outcome," and that students appear to use a contaminated measure when judging instruction.

Kim & MacCann (2018), in the British Journal of Educational Psychology, replicated and extended this with evidence from two subject areas, confirming that instructor personality matters for student evaluations. Their work shows the effect is not an artefact of a single discipline or a single national context.

Set against the broader literature, this is consistent rather than fringe. Spooren, Brockx & Mortelmans (2013), in their state-of-the-art review in Review of Educational Research, conclude that SET validity is "a multifaceted and complex issue" and that scores reflect more than teaching effectiveness alone. Instructor personality is one well-documented strand of that "more."

The mechanism is intuitive. Extraverted, warm, expressive instructors are more engaging in the lecture theatre; that engagement is genuinely pleasant and is easy to convert into a high number on a Likert item. The problem is not that engagement is bad — it is that expressiveness can inflate ratings independently of how much students actually learn (the classic "Dr Fox effect"). Personality is the dispositional version of that same confound.

Why it matters for course evaluation in practice

If a meaningful chunk of a SET mean is personality, three practical consequences follow for any university quality-assurance (QA) process.

  1. Cross-instructor comparison is hazardous. Ranking instructors by raw means, or applying fixed thresholds in promotion and tenure cases, partly ranks them by extraversion and warmth. A reserved but rigorous lecturer can be systematically under-rated relative to a charismatic but average one. This compounds the well-known small-sample and confidence-interval problems with point estimates.

  2. It interacts with other biases. Personality perception is not neutral across gender, ethnicity or accent. Warmth and "agreeableness" are judged through stereotyped lenses, so a personality confound can act as a carrier for demographic bias rather than an independent one. Treating the SET mean as objective hides this layering.

  3. Improvement signals get muddied. If a department reads a modest score as "this person cannot teach," when much of the gap is dispositional style, the resulting intervention is mistargeted. The actionable information about what to change in the course lives in the specifics — the open text, the concrete incidents — not in the global number that personality most contaminates.

The constructive response is not to abandon student voice. Students are excellent witnesses to their own experience. The response is to (a) stop treating the global mean as a precise, comparable measure of teaching quality, and (b) collect feedback in a form that separates what happened in the course from how much students liked the person delivering it.

Limitations and honest caveats

A careful reader should hold several caveats.

  • Direction and self-report. Much of this work relies on students rating both personality and teaching, raising common-method and halo concerns: a student who likes a teacher may rate them as both warmer and better, inflating the correlation. Clayson & Sheffet's thin-slice design mitigates but does not eliminate this.
  • "Favoured" traits are partly real teaching. Conscientiousness and openness plausibly cause better course organisation and richer content. Some of the personality-rating link may therefore be valid variance, not contamination. The honest claim is that personality explains variance beyond perceived learning and grades — not that all of it is illegitimate.
  • Effect sizes vary. Studies differ in samples, instruments and disciplines; the magnitude of the personality contribution is not a single fixed number, and some estimates are modest. Generalising from business or psychology cohorts to, say, clinical or studio teaching should be done cautiously.
  • Personality is not a "bias" you can simply subtract. Unlike a measurable covariate such as class size, instructor personality is latent and entangled with legitimate pedagogy, so statistical "correction" is not straightforward.

None of these caveats overturns the core finding. They sharpen it: the prudent reading is that SET means are partly personality, the share is non-trivial, and the appropriate institutional posture is humility about what a single number can tell you.

How Koji incorporates this

Koji is built on the premise that a course evaluation should surface what to improve, not just how much students liked the lecturer. Several design choices target the personality-contamination problem specifically — framed, honestly, as mitigation rather than elimination.

  • Conversational, AI-moderated interviews that probe past the number. Instead of stopping at a 1-to-5 rating, Koji's AI interviewer follows up: what specifically helped you learn? where did you get stuck? what would you change? This moves the centre of gravity from a global affect-laden judgement (where personality bites hardest) toward concrete, behaviour-anchored evidence about the course design.
  • Structured question types that separate constructs. Koji supports open_ended, scale, single_choice, multiple_choice, ranking and yes_no items. QA teams can deliberately pair a global scale item with open-ended and behaviour-specific questions, so the personality-loaded global score is never read in isolation.
  • Automatic thematic analysis of open text. Koji clusters free-text responses into themes (pacing, assessment clarity, feedback timeliness), which are far less contaminated by instructor charm than a single Likert mean and far more actionable for a programme director.
  • Bias-aware, dispersion-sensitive reporting. Koji reports distributions, theme prevalence and uncertainty rather than encouraging a bare mean to be ranked against a threshold — directly addressing the "do not over-read the global score" lesson.
  • Quality scoring of responses helps distinguish a substantive, evidence-rich answer from a one-line affect dump, so committees weight the informative signal.

The aim is to make the interpretation match the science: keep the student voice, but route decisions through the parts of the evidence that personality contaminates least.

Koji's core research platform at koji.so applies the same AI-moderated interview engine to product and customer research, where separating "I liked the person" from "the thing actually worked" is just as important.

Related resources

References

  • Patrick, C. L. (2011). Student evaluations of teaching: effects of the Big Five personality traits, grades and the validity hypothesis. Assessment & Evaluation in Higher Education, 36(2), 239–249. https://doi.org/10.1080/02602930903308258
  • Clayson, D. E., & Sheffet, M. J. (2006). Personality and the student evaluation of teaching. Journal of Marketing Education, 28(2), 149–160. https://doi.org/10.1177/0273475306288402
  • Kim, L. E., & MacCann, C. (2018). Instructor personality matters for student evaluations: Evidence from two subject areas at university. British Journal of Educational Psychology, 88(4), 584–605. https://doi.org/10.1111/bjep.12205
  • Spooren, P., Brockx, B., & Mortelmans, D. (2013). On the validity of student evaluation of teaching: The state of the art. Review of Educational Research, 83(4), 598–642. https://doi.org/10.3102/0034654313496870
  • Messick, S. (1989). Validity. In R. L. Linn (Ed.), Educational Measurement (3rd ed., pp. 13–103). American Council on Education / Macmillan.

Related articles

research-methods

Beauty Bias in Course Evaluations: Are Attractive Instructors Rated as Better Teachers?

What the research says about physical-attractiveness bias in student evaluations of teaching — from Hamermesh & Parker (2005) to German and laboratory replications — and how to design evaluation so appearance does not masquerade as teaching quality.

research-methods

The Dr. Fox Effect: Does Lecturer Charisma Inflate Course Evaluations?

The 1973 Dr. Fox experiment, its meta-analytic re-interpretation, and the 2014 re-revisitation — what the evidence really says about whether expressive delivery seduces students into rating empty teaching highly, and how to evaluate substance over showmanship.

research-methods

The Halo Effect in Course Evaluations: When One Impression Colours Every Rating

When students like an instructor, that single global impression bleeds into their ratings of unrelated specifics. From Nisbett & Wilson (1977) to Feistauer & Richter (2018) and Cannon & Cipriani (2021), here is how the halo effect distorts itemised course evaluations — and what to do about it.

research-methods

What Do Student Evaluations Actually Measure? Marsh, the SEEQ, and the Case for Multidimensional Feedback

Herbert Marsh spent decades showing that well-built student evaluations are multidimensional, reliable, and stable — and that a single global score throws away most of what they can tell you. What his work establishes, where critics push back, and how to design feedback that is actually usable.