New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods9 min read

The Warmth Halo: How Instructor Immediacy Shapes Course Evaluations More Than Learning Does

What the meta-analytic evidence on teacher nonverbal immediacy and warmth tells us about course evaluations — why warmth strongly predicts how much students like a course and think they learned, but only weakly predicts what they actually learn.

Koji Education Team

Product

In brief

Instructor warmth — eye contact, smiling, movement, vocal variety, approachability, the cluster researchers call immediacy — is one of the strongest correlates of how students rate a course. But the same meta-analytic evidence shows immediacy is tied far more tightly to perceived and affective learning than to actual cognitive learning. A course evaluation dominated by global "overall" items therefore partly measures likability, which is real and consequential but is not the same thing as teaching effectiveness — and which carries documented gendered expectations.

What the research says

The foundational quantitative result is Witt, Wheeless, and Allen's (2004) meta-analysis in Communication Monographs, synthesizing 81 studies and more than 24,000 students. The pattern in their pooled correlations is the heart of the matter. Nonverbal immediacy correlated r ≈ .51 with perceived learning and r ≈ .49 with affective learning (how much students like the subject and instructor), but only r ≈ .17 with cognitive learning — actual performance on tests and assignments. Verbal immediacy showed the same shape even more starkly: r ≈ .49 with perceived learning and affective learning, but just r ≈ .06 with cognitive learning. In other words, warmth powerfully shapes how much students feel they learned and how much they like the experience, while its association with what they can actually demonstrate is slight.

This is the quantitative backbone behind a much older and more theatrical demonstration: the Dr. Fox effect. Naftulin, Ware, and Donnelly (1973) hired a professional actor to deliver an expressive, charismatic, but deliberately content-free lecture (laced with double-talk and non sequiturs) to an audience of professionals, who nonetheless rated "Dr. Fox" favourably — evidence that expressive delivery can "seduce" listeners into believing they have learned. The phenomenon is not pure illusion, though. Marsh and Ware (1982) re-analyzed the Dr. Fox paradigm multidimensionally and showed the picture is more nuanced: expressiveness primarily inflated the instructor enthusiasm dimension of ratings, while content coverage drove the knowledge dimension and exam performance. Warmth, then, does not contaminate every rating equally — it loads onto the rating dimensions logically tied to delivery.

How quickly does this warmth signal form? Astonishingly fast. Ambady and Rosenthal (1993) found that strangers' ratings of teachers based on silent video clips under 30 seconds long — and even clips as short as two to ten seconds — significantly predicted those teachers' actual end-of-semester student evaluations, with physical attractiveness adding further predictive power. Students appear to form a warmth-based impression in the first moments of contact, and end-of-term ratings are in part the elaboration of that thin-slice judgment.

The crucial limitation visible across this literature is that warmth expectations are not gender-neutral. MacNell, Driscoll, and Hunt (2015) ran an online course in which two instructors each taught under both a male and a female identity, holding actual instruction constant. Students rated the perceived-male identity significantly higher — including, in related work, on items that have nothing to do with warmth. Because students often expect and reward warmth more from women instructors while judging competence against a male-professor prototype, the "warmth halo" interacts with bias in ways that disadvantage women who are evaluated as insufficiently nurturing or men's competence ratings that ride on warmth they were never expected to provide.

Why it matters for course evaluation in practice

If warmth correlates around .5 with perceived learning but near .1–.2 with actual learning, then a global evaluation item is, to a substantial degree, a likability measure. For a quality-assurance process, three implications matter.

First, ranking instructors on overall ratings partly ranks interpersonal warmth, not instructional effectiveness — and because warmth is judged against gendered and cultural expectations, those rankings can encode bias rather than quality. Second, warmth is genuinely valuable and should not be dismissed: affective learning predicts motivation, attendance, help-seeking, and persistence, so an instructor whose warmth raises affective learning is doing something that matters for student outcomes over time. The error is not valuing warmth; it is mistaking warmth for cognitive effectiveness. Third, multidimensional instruments are more defensible than single global items: because expressiveness loads onto enthusiasm-type dimensions rather than knowledge or workload dimensions (per Marsh and Ware), separating these dimensions in reporting lets a committee see which aspect of teaching a score reflects, instead of blending warmth and substance into one contaminated number.

Limitations and honest caveats

Several caveats temper a strong reading. The immediacy literature is largely correlational: warm instructors may also prepare more, design better, or teach motivated cohorts, so the warmth–perceived-learning correlation is not cleanly causal. Many immediacy studies rely on student-reported measures of both immediacy and learning, inflating shared-method correlations; the small cognitive-learning correlation is in fact the more method-independent estimate, which is partly why it is smaller. The Dr. Fox studies used short, artificial lectures and small samples, and Marsh and Ware's re-analysis shows the original "pure seduction" interpretation was overstated. Ambady and Rosenthal's thin-slice finding is robust but describes prediction, not proof that the snap judgment is wrong. And affective outcomes, again, are not noise — they are part of what a university reasonably wants from teaching. The honest conclusion is that warmth is a real, valuable, but distinct construct that global evaluation items over-weight relative to learning, with a gendered confound that makes naive cross-instructor comparison risky.

How Koji incorporates this

Koji for Education is built to keep warmth visible as warmth rather than letting it silently inflate a single effectiveness number.

  • Dimension-aware questioning. Rather than collapsing everything into one "overall" item, Koji supports structured scale, single_choice, and ranking questions that distinguish enthusiasm and approachability from clarity, challenge, workload, and learning — mirroring the multidimensional structure Marsh and Ware showed is necessary to keep delivery from contaminating substance.
  • Conversational probing of the "why". Koji's AI-moderated interviews follow a warm reaction with a substantive probe: a student who says "she was so engaging" is asked what they can now do that they could not before. This converts a diffuse warmth impression into specific, learning-anchored evidence.
  • Thematic analysis that separates affect from learning. Koji automatically codes open text so a quality office can distinguish "I enjoyed it / she was approachable" themes from "I can now apply X" themes, preventing affective warmth from being read as cognitive effectiveness.
  • Bias-aware reporting. Because warmth expectations are gendered, Koji's reporting is designed to discourage raw cross-instructor ranking on global items and to foreground the dimensional and qualitative evidence that is harder to reduce to a popularity contest.
  • Triangulation by design. Koji positions student feedback as one source among several (peer review, assessment outcomes), consistent with the consensus that warmth-laden ratings should never stand alone in high-stakes decisions.

These mechanisms are framed as mitigations, not cures: they make the warmth halo legible and reduce its weight in interpretation, but no instrument can remove a bias that operates in the first thirty seconds of human contact. Koji's core research platform at koji.so applies the same interview engine to customer and product research, where likability of an interface and its actual usefulness diverge in exactly the way warmth and learning do here.

Designing around the warmth halo

If warmth loads onto enthusiasm-type dimensions and barely onto learning, the instrument should be built to keep those signals separate rather than blending them into one verdict. A defensible design asks distinct questions about approachability and responsiveness, about clarity and structure, and about challenge and workload, and reports them as a profile rather than a single average. This does three things at once: it preserves the genuinely useful affective signal (a warm, approachable instructor really does help students engage and persist), it prevents that signal from masquerading as evidence of cognitive effectiveness, and it makes the gendered confound auditable — because a pattern of equal clarity but lower "authority" or "leadership" ratings for women instructors is visible in a profile and invisible in a mean.

A second design principle follows from the thin-slice finding: because warmth impressions form in seconds and end-of-term ratings partly elaborate them, the most diagnostic feedback is the kind that forces students past the global impression toward specific evidence. Asking "what could you do by the end that you could not do before, and what helped?" pulls a student from "I liked her" toward an account a committee can actually weigh. The goal is not to suppress warmth from the record but to ensure a likable instructor and an effective one are not automatically scored as the same person — and that an effective but reserved instructor is not penalized for the absence of a trait only weakly tied to what students learn.

Related resources

References

  • Witt, P. L., Wheeless, L. R., & Allen, M. (2004). A meta-analytical review of the relationship between teacher immediacy and student learning. Communication Monographs, 71(2), 184–207. https://doi.org/10.1080/036452042000228054
  • Naftulin, D. H., Ware, J. E., & Donnelly, F. A. (1973). The Doctor Fox lecture: a paradigm of educational seduction. Journal of Medical Education, 48(7), 630–635. https://pubmed.ncbi.nlm.nih.gov/4708420/
  • Marsh, H. W., & Ware, J. E. (1982). Effects of expressiveness, content coverage, and incentive on multidimensional student rating scales: new interpretations of the Dr. Fox effect. Journal of Educational Psychology, 74(1), 126–134. https://doi.org/10.1037/0022-0663.74.1.126
  • Ambady, N., & Rosenthal, R. (1993). Half a minute: predicting teacher evaluations from thin slices of nonverbal behavior and physical attractiveness. Journal of Personality and Social Psychology, 64(3), 431–441. https://doi.org/10.1037/0022-3514.64.3.431
  • MacNell, L., Driscoll, A., & Hunt, A. N. (2015). What's in a name: exposing gender bias in student ratings of teaching. Innovative Higher Education, 40(4), 291–303. https://doi.org/10.1007/s10755-014-9313-4