New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Methodology9 min read

The Expectation Gap: Why Course Evaluations Measure Surprise, Not Quality

Student satisfaction is the gap between what students expected and what they got. That means a rigorous course can score low simply for defying expectations — and an easy one can score high for flattering them.

Koji Education Team

Product ·

Bottom line up front: When you ask students how satisfied they are with a course, you are not measuring teaching quality in the abstract. You are measuring the gap between what they expected and what they experienced — what consumer-behaviour researchers call disconfirmation. This has an uncomfortable implication: a demanding, well-taught course that defies students' expectations of an easy ride can score lower than a shallow course that comfortably meets a low bar. If your evaluation reports a raw satisfaction mean and stops there, you may be rewarding expectation-management over education.

Satisfaction is a comparison, not a measurement

The dominant theory of how satisfaction forms is expectancy-disconfirmation theory (EDT), introduced by Richard Oliver in 1980. Its claim is simple and robust across decades of research: satisfaction is the result of comparing a prior expectation against a perceived outcome. When the experience exceeds expectation (positive disconfirmation), satisfaction rises; when it falls short (negative disconfirmation), satisfaction drops — even if the absolute quality of the experience was high. Satisfaction is a relative judgement, anchored to whatever the person walked in expecting.

Applied to course evaluation, this is not a metaphor; it has been tested directly. A frequently cited study had business students evaluated using both a conventional satisfaction question and a disconfirmation framing. The results diverged sharply: on the traditional question 89.3% said they were satisfied, but when the same students were asked through the expectation-comparison lens, 93.1% were classified as dissatisfied — their experience had not lived up to what they believed an excellent education should be (see Disconfirmation Theory and student satisfaction, ERIC). The two framings, asked of the same people about the same courses, produced almost mirror-image conclusions. A single satisfaction mean hides all of that.

The rigour penalty

Here is why this matters for anyone using evaluations to judge teaching. If satisfaction tracks the expectation gap, then anything that raises expectations or withholds the easy rewards students anticipated will depress scores — regardless of learning. Research consistently finds that perceived workload and grading leniency move ratings: students who expected a light, high-grade course and met a rigorous one experience negative disconfirmation and rate accordingly. This is the same mechanism behind Wolfgang Stroebe's argument that evaluations, when consequential, encourage poor teaching and contribute to grade inflation: the path of least resistance to a high score is to meet or exceed the expectation of an easy, entertaining, generously graded experience.

Prior interest compounds the effect. A student who arrives expecting to love a subject — and does — records positive disconfirmation; a student compelled into a required course they expected to dislike starts from a deficit no amount of teaching skill fully erases. None of this is the instructor's doing, yet all of it lands in the satisfaction number. The same study tradition that established the multidimensional nature of student ratings (Herbert Marsh's SEEQ work) also showed prior subject interest is among the stronger correlates of overall ratings — a correlate that has nothing to do with how well the course was taught.

But isn't a met expectation exactly what we want?

The strongest counterargument deserves a fair hearing. Surely, a critic says, satisfying students is a legitimate goal — meeting expectations is what a service should do, and dismissing satisfaction as "mere surprise" is academic snobbery that ignores the student's real, paid-for experience.

This is partly right, and worth conceding. Student satisfaction is a genuine outcome; a programme that routinely leaves students feeling cheated has a real problem, and disconfirmation data can surface it usefully. But two things follow. First, satisfaction is a measure of expectation management, and expectation management is only sometimes aligned with good teaching — it is perfectly possible to satisfy students by demanding less of them, which is precisely the failure mode quality assurance exists to catch. Second, the fix is not to discard the signal but to decompose it: separate "did this meet your expectations?" from "what did you learn and how was it taught?" so that a low score can be read correctly — as unmet expectation, as genuine teaching weakness, or as the predictable friction of a rigorous course. A raw mean fuses all three and lets you mistake any one for the others.

Reading evaluations through the expectation lens

For quality officers and programme directors, the practical implications are concrete:

  • Stop treating a satisfaction mean as a teaching-quality score. Interpret it as an expectation-gap signal and ask what expectations the cohort carried in.
  • Measure expectations explicitly where you can. A brief sense of what students anticipated — in workload, difficulty, relevance — turns an uninterpretable number into a readable one.
  • Probe the why behind a low score. "Disappointing" can mean "I expected an easy A" or "the teaching genuinely failed me." Only open, follow-up questioning tells them apart.
  • Watch for the rigour penalty in high-stakes use. Before a demanding course's low score counts against an instructor, ask whether you are penalising teaching or rewarding leniency.

Where expectations come from — and why you can shape them

If satisfaction is a comparison against expectation, then expectations are not fixed weather; they are formed, and partly formable. Students arrive with priors set by the syllabus and module descriptor, by what previous cohorts told them, by public platforms like RateMyProfessors, by the marketing language of the programme, and by the difficulty of adjacent courses. A module advertised as a gentle introduction that turns out to be demanding generates negative disconfirmation almost regardless of how well it is taught; the same content, framed honestly as rigorous and supported, can produce satisfaction because the experience meets a calibrated expectation.

This has a constructive implication that is easy to miss in the gloom about bias: expectation management is a legitimate, teachable lever for both teaching quality and satisfaction at once. Setting accurate expectations early — being explicit about workload, the purpose of difficulty, and the support available — narrows the gap not by lowering standards but by aligning the prior with the reality. Research on student satisfaction consistently finds that unmet expectations, more than absolute difficulty, drive dissatisfaction; calibrate the expectation and a rigorous course need not pay a rating penalty.

For evaluation design, the lesson is to capture the prior, not just the outcome. Asking students what they expected — and, retrospectively, whether the course was framed accurately — converts an uninterpretable satisfaction number into a diagnosis: was a low score unmet expectation, miscommunication, or genuine teaching weakness? Without that, you are reading the difference of two quantities while only ever measuring one of them.

Where Koji fits

Decomposing satisfaction into expectation, experience and learning is hard with a static form — a Likert grid captures the fused number and nothing underneath it. This is the gap an AI-native approach is built to close. Koji for Education uses AI-moderated conversational interviews that do exactly the probing the expectation lens requires: when a student gives a low rating, the AI can ask why — distinguishing "I expected less work" from "I did not understand the material" — rather than leaving you with an ambiguous 2.8. Its six structured question types let you ask about expectations and learning alongside satisfaction, and its automatic thematic analysis surfaces, across a whole cohort, whether low scores cluster around unmet expectations or genuine teaching issues. Standardised, bias-aware AI moderation means every student is probed consistently, without the variability of human interviewers. The same conversational engine powers expectation-versus-experience research on the main Koji platform for teams studying customers rather than courses.

Koji does not claim to eliminate the expectation gap — it is a real feature of human judgement, not a bug to be removed. What it does is make the gap visible and interpretable, so a low satisfaction score becomes a question you can answer rather than a verdict you must accept.

Stop mistaking surprise for quality. Explore Koji for Education to see what your satisfaction scores are actually measuring.