New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Evaluation bias9 min read

The Course Before Yours Is Grading Your Course: Contrast Effects in Student Evaluation

A student who rated a brilliant seminar an hour before yours is a harsher judge of your course than one who sat through a dull lecture first. Contrast effects mean the comparison set, not just your teaching, shapes the number. Here is the evidence and what to do about it.

Koji Education Team

Product ·

The short answer: A course-evaluation score is never formed in a vacuum. Decades of judgment research show that a rating is anchored to whatever the respondent was recently exposed to — the course they evaluated a screen earlier, the instructor they had the previous hour, the standard set by the rest of their timetable. When the comparison is favourable, your course looks worse by contrast; when it is unfavourable, your course looks better. This is the contrast effect, and it means two instructors of identical quality can receive systematically different scores purely because of what surrounded them. It cannot be averaged away, and it is invisible in the clean-looking mean on the dashboard.

What the contrast effect is

Human beings do not judge on absolute scales. We judge relative to a reference point, and that reference point is heavily influenced by whatever is salient at the moment of judgment. The classic demonstration comes from the employment-interview literature. Wexley, Yukl, Kovacs and Sanders (1972), in Importance of contrast effects in employment interviews (Journal of Applied Psychology, 57[1]), had raters watch videotaped candidates in sequence. An average candidate was rated substantially lower when preceded by two strong candidates, and higher when preceded by two weak ones. In some conditions the preceding candidates accounted for a large share of the variance in ratings of the target — the person being judged had not changed at all; only the comparison set had.

The same mechanism has been shown in essay grading. Studies of Contrast Effects in Evaluating Essays find that a mediocre essay is marked more harshly after a run of excellent ones, and more generously after a run of poor ones. The grader believes they are applying a fixed standard; the data show the standard drifts with the sequence.

Why does this happen? The dominant account is the Inclusion/Exclusion Model of Norbert Schwarz and Herbert Bless (1992; elaborated 2010). When a comparison stimulus is included in the mental representation of the target, you get assimilation — the target is pulled toward the context. When the comparison is excluded and used instead as a standard, you get contrast — the target is pushed away from it. Sequential judgments of the kind students make when rating one course after another are precisely the conditions under which contrast dominates: each course is a distinct object, so the previous one becomes a yardstick rather than part of the target.

Why this matters specifically for course evaluation

Traditional student evaluation of teaching (SET) is delivered in exactly the format that maximises contrast. Consider the modern reality:

  • Batch end-of-term surveys. Students are prompted to evaluate all their modules in one sitting, often back-to-back in the same portal. The second course is judged in the shadow of the first. A student who rates a genuinely outstanding module first will carry a heightened standard into the next module's form.
  • Timetable adjacency. The lecture a student attends immediately before yours sets an affective and quality baseline. A dull 9 a.m. makes a competent 11 a.m. feel like a relief; an electrifying 9 a.m. makes the same 11 a.m. feel flat.
  • Cohort comparison. Students implicitly rank their teachers against each other within a programme. The best instructor in a cohort raises the bar that every colleague is then measured against — a structural penalty for teaching alongside a star.

None of this reflects the quality of your course. It reflects the reference set the respondent happened to assemble. Yet institutions routinely compare a 4.1 to a 4.4 as though the difference were a property of the teaching. As our analysis of benchmarking course-evaluation scores argues, comparison across contexts is already fraught; contrast effects add a further, largely uncontrolled source of non-comparability at the level of the individual respondent.

Contrast is a cousin of two biases we have covered before. It shares machinery with question-order and context effects, where an earlier item reframes a later one, and with the peak-end and recency effects that distort what students remember at the moment they fill in the form. Where those concern the internal structure of one survey, contrast concerns the external sequence of what the student evaluated before arriving at yours.

How large is the effect, and how worried should you be?

Honesty requires proportion. In the Wexley et al. paradigm, contrast effects were statistically robust but modest in many conditions — accounting for a minority of variance when a target followed a single strong or weak predecessor, and becoming large only in the more extreme designs (an average target sandwiched by two very strong or two very weak predecessors). Real timetables are messier and noisier than a controlled videotape study, which cuts both ways: the effect is diluted by many uncorrelated comparisons, but it is also never controlled for.

The right conclusion is not "all scores are meaningless." It is more specific: contrast is a source of construct-irrelevant variance that is systematically larger for some instructors than others — precisely those teaching adjacent to unusually strong or weak colleagues, or later in a batch-evaluation sequence. For high-stakes decisions built on small score differences, that unmodelled variance is exactly the kind that turns noise into a personnel judgement. This is the same over-interpretation problem we examine in is a 0.3-point difference real: the mechanism here supplies one more reason the third decimal place is not signal.

"But doesn't every measurement have context effects? Isn't this just an excuse?"

This is the strongest objection, and it deserves a direct answer. Yes — every judgment is contextual, and no instrument eliminates that. If contrast were random, it would inflate error but not bias any particular instructor, and the fix would simply be "collect more data." The problem is that contrast is not random with respect to the things institutions care about. It correlates with timetabling, with programme composition, and with survey design — all of which are stable features that attach to specific instructors term after term. A lecturer permanently scheduled after the programme's most charismatic colleague carries a permanent handicap that no amount of averaging within their own data will remove, because the bias is in the comparison set, not in the sampling of respondents.

A second objection: "students aren't interviewers rating three candidates; the analogy is loose." Fair. The transfer from lab to lecture hall is not one-to-one, and we should not pretend a 12%-of-variance figure from a videotape study is the number for a French seminar in Utrecht. But the direction of the effect is one of the most replicated findings in judgment research, spanning interviews, essays, attractiveness ratings (Kenrick & Gutierres, 1980) and performance appraisal. The burden of proof runs the other way: there is no reason to believe course ratings are the one domain immune to a mechanism that appears everywhere else humans make sequential evaluative judgments.

What actually reduces contrast — and what Koji does about it

You cannot abolish comparison; it is how cognition works. But you can change the conditions that let it dominate, and you can stop treating the resulting number as a context-free measurement.

  1. Break the batch. Contrast is worst when courses are rated back-to-back against each other. Collecting feedback per module, closer to the experience rather than in an end-of-term data-entry marathon, reduces the number of adjacent comparisons a single sitting generates. This is one practical reason we advocate formative, mid-cycle collection over a single summative sweep.

  2. Ask about specifics, not global impressions. Contrast bites hardest on holistic, "overall satisfaction" judgments, which are exactly the judgments that float free of any fixed anchor. Questions tied to concrete, verifiable events — what happened when you asked for help; describe how feedback on the second assignment was returned — give the respondent an internal standard that is harder to displace with an external one.

Koji is built around both of these. Instead of a static Likert form completed in a batch, Koji runs an AI-moderated conversational interview that probes for concrete instances rather than a summary rating, using six structured question types (open_ended, scale, single_choice, multiple_choice, ranking and yes_no) so that a global number is never the only signal. Because the AI moderator is standardized, every student is guided to the same evidence-anchored questions regardless of which course they evaluated first — removing the human-moderator inconsistency that would otherwise add its own contrast. Koji's automatic thematic analysis then surfaces what students actually described, so a course that scored slightly lower can be read against the substance of the feedback rather than an unadjusted mean. For institutions comparing across a programme, that shift — from "whose average is higher" to "what evidence supports a concern" — is the only defensible response to a bias that lives in the comparison itself.

Koji does not claim to eliminate contrast; no instrument can. It claims to reduce its grip by moving away from the global, batch-rated Likert score that contrast most easily distorts, and to make the residual comparability problem visible rather than hidden inside a clean number. The same conversational engine underpins the main Koji platform for customer and user research, where sequential-context effects distort product feedback in exactly the same way.

The bottom line

The number on your evaluation dashboard is partly a report on the courses your students happened to evaluate before yours. That is not a reason to discard student feedback — it is a reason to stop reading small differences in the mean as facts about teaching, and to collect feedback in a form where the substance survives the sequence. Contrast is a feature of human judgment, not a flaw in your course. The flaw is an evaluation system that pretends it isn't there.