New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
research-methods9 min read

The Halo Effect in Course Evaluations: When One Impression Colours Every Rating

When students like an instructor, that single global impression bleeds into their ratings of unrelated specifics. From Nisbett & Wilson (1977) to Feistauer & Richter (2018) and Cannon & Cipriani (2021), here is how the halo effect distorts itemised course evaluations — and what to do about it.

Koji Education Team

Product

Answer

The halo effect is the tendency for one global impression of an instructor — most often simple likeability — to contaminate students'' answers to specific, supposedly independent questions about clarity, organisation, fairness or workload. The evidence is robust: a single positive (or negative) overall feeling pulls the whole itemised profile in the same direction, which means a multi-item course evaluation often measures one general sentiment dressed up as several distinct dimensions. The practical consequence is that fine-grained item scores are less diagnostic than they appear; the fix is not to abolish itemised evaluation but to design and interpret it so that the halo is detected and the specific, actionable signal is recovered.

What the research says

The canonical demonstration is Nisbett and Wilson (1977). They showed students one of two videotaped interviews with the same college instructor, who spoke English with a European accent. In one version he was warm and friendly; in the other, cold and distant. Students who saw the warm instructor rated his appearance, mannerisms and accent as appealing, while those who saw the cold version rated those very same attributes as irritating — even though the attributes themselves were identical. A global evaluation (likeable / not likeable) reached backwards and altered ratings of concrete, observable features. Tellingly, subjects were unaware of the influence and, when asked, often denied it. This is the halo effect in its purest experimental form, and it used precisely the object course evaluations rely on: students judging a teacher.

Four decades later, Feistauer and Richter (2018) quantified the same mechanism in real student evaluations of teaching. Examining likeability and prior subject interest as potential biasing factors, they found the instructor''s perceived likeability alone explained roughly 36.5% of the variance in the overall course rating — an enormous share for a variable that is conceptually distinct from teaching quality. When one third of the variance in your headline number is "how much we liked them," the itemised structure underneath it is heavily shared rather than independent.

Cannon and Cipriani (2021) approached the problem with an ingenious identification strategy. They combined a subjective questionnaire item (students'' rating of lecture-room capacity) with an objective fact (the actual size of the room). Because the true answer was known, any systematic drift in the subjective rating that tracked the student''s overall satisfaction — rather than the real room size — exposed a halo. They confirmed the presence of a halo effect, but added an important nuance: the contaminated item still retained some informative content. Their reading is comparatively sanguine — halo distortion is real but does not render evaluation questionnaires useless. That qualification matters and keeps the literature honest: halo is a contaminant to manage, not a reason to discard student feedback.

Together these studies sketch a coherent picture. The halo is (a) demonstrable under controlled manipulation (Nisbett & Wilson), (b) large in field data when the global impression is likeability (Feistauer & Richter), and (c) present but not totally destructive of information (Cannon & Cipriani).

Why it matters for course evaluation in practice

1. Itemised profiles are less independent than they look. If a department compares an instructor''s "organisation" score with their "fairness" score and reads the small gap as meaningful, they may be over-interpreting noise around a single underlying sentiment. High inter-item correlations are a halo signature, not proof of uniformly excellent teaching.

2. Global items contaminate specific items — order and framing matter. Asking "Overall, how satisfied are you?" first can prime the halo for every subsequent specific item. Survey structure can either amplify or dampen the effect.

3. Likeability is not teaching quality — but it is not irrelevant either. Rapport genuinely supports learning, so the goal is not to strip likeability out entirely but to separate "I enjoyed this instructor" from "this instructor helped me learn X." That separation is exactly what a single Likert grid struggles to achieve, and where richer evidence helps.

4. High-stakes use is the danger zone. For formative feedback, a halo-tinged profile is still useful directional information. For summative personnel decisions, mistaking a likeability halo for evidence of pedagogical skill (or its absence) is a real fairness risk — one that compounds the documented demographic biases in student ratings.

Limitations and honest caveats

  • Halo is not always bias. Some positive correlation across items is legitimate: genuinely good teachers tend to be clear and fair and organised. The methodological challenge is separating a true common cause (real teaching quality) from a spurious one (mere likeability). Cannon and Cipriani''s design is clever precisely because it pins down distortion against an objective benchmark; most field studies cannot.
  • Effect sizes are context-bound. The 36.5% figure (Feistauer & Richter, 2018) is specific to their instrument, discipline and sample; it should be read as evidence that the halo can be large, not as a universal constant.
  • Lab realism. Nisbett and Wilson''s manipulation is powerful but staged; real evaluations are completed weeks later, in bulk, with different motivation. The mechanism generalises; the magnitude may not.
  • The "still informative" finding cuts both ways. Cannon and Cipriani show contaminated items keep some signal — reassuring — but that does not tell you how much signal survives in your instrument, which depends on your items and population.

How Koji incorporates this

Koji for Education is designed to surface the specific, actionable signal that a halo tends to flatten — without pretending the halo can be eliminated.

  • Probing beyond the global impression. A single Likert grid is where the halo does its worst, because every number is anchored to the same overall feeling. Koji''s AI-moderated conversational interview instead asks students to substantiate — "you said the course was well organised; can you give a specific example?" Concrete, behaviour-anchored answers are harder to generate from a diffuse "I liked them" sentiment, which helps separate likeability from teaching specifics.
  • Thematic analysis that decomposes the global feeling. Koji''s automatic thematic analysis of open text can show why a cohort is positive or negative — pulling apart rapport/likeability themes from clarity, pace, assessment-fairness, and workload themes. This gives a reviewer the decomposition a correlated Likert profile hides.
  • Bias-aware reporting. Koji is designed to flag when itemised scores move together suspiciously and to frame likeability-laden signals cautiously, so a programme director does not read a halo as independent corroboration across dimensions.
  • Right tool for the stakes. Because halo distortion is most dangerous in summative use, Koji supports triangulation across cohorts and mid-cycle formative collection, encouraging halo-tinged single-cycle data to be corroborated before it informs decisions about people.

These mechanisms are designed to mitigate the halo effect, not to claim they remove it; the underlying cognitive bias (as Nisbett and Wilson showed) operates below students'' awareness. Koji''s core research platform at koji.so applies the same probing, theme-decomposing interview engine to product and customer research, where a single "overall satisfaction" halo distorts feedback just as readily.

Detecting halo in your own data

A quality-assurance team does not need an experimental design to look for halo signatures in routine evaluation data; a few diagnostics go a long way.

  • Inspect the inter-item correlation matrix. If "organisation", "fairness", "clarity" and "workload appropriateness" all correlate above roughly 0.7 with one another and with the overall item, you are likely looking at one general sentiment rather than four independent judgements. Genuinely differentiated teaching profiles show some spread across dimensions.
  • Watch the overall-first ordering. Placing a global satisfaction item before the specifics can prime the halo for everything that follows. Where possible, ask for concrete, behaviour-anchored judgements before the global summary.
  • Triangulate against non-student evidence. Peer observation, learning-outcome attainment, and curriculum review are not subject to the same likeability halo. Convergence across sources is the strongest defence against mistaking "well-liked" for "effective".
  • Separate enjoyment from learning explicitly. Ask, in words, both "what did you enjoy?" and "what helped you learn, even if it was difficult?" The gap between the two answers is often where the halo hides — and where the most useful, least-flattering feedback lives.

None of these steps eliminate the halo; as Nisbett and Wilson showed, it operates below conscious awareness and cannot be willed away by either students or analysts. But surfacing it stops a department from reading a single diffuse impression as if it were independent corroboration across many dimensions — which is precisely the error that makes halo-tinged ratings dangerous in summative use.

Related Resources

References

  • Nisbett, R. E., & Wilson, T. D. (1977). The halo effect: Evidence for unconscious alteration of judgments. Journal of Personality and Social Psychology, 35(4), 250–256. https://doi.org/10.1037/0022-3514.35.4.250
  • Feistauer, D., & Richter, T. (2018). Validity of students'' evaluations of teaching: Biasing effects of likability and prior subject interest. Studies in Educational Evaluation, 59, 168–178. https://doi.org/10.1016/j.stueduc.2018.07.009
  • Cannon, E., & Cipriani, G. P. (2022). Quantifying halo effects in students'' evaluation of teaching. Assessment & Evaluation in Higher Education, 47(1), 1–14. https://doi.org/10.1080/02602938.2021.1888868