What Services Marketing Already Knew About Course Evaluation: SERVQUAL, SERVPERF and HEdPERF
If your course evaluation treats students as customers rating a service, you have adopted a measurement paradigm that services marketing spent decades stress-testing. Here is what SERVQUAL, SERVPERF and the higher-education-specific HEdPERF can lend course evaluation — and the assumptions you should refuse to import.
Koji Education Team
Product · July 31, 2026
Bottom line up front: The moment your course evaluation asks students to rate "responsiveness", "the learning environment" or "how well the course met your expectations", you have borrowed — usually without knowing it — from services marketing. That field built three instruments to measure service quality: SERVQUAL (Parasuraman, Zeithaml & Berry, 1988), its perceptions-only successor SERVPERF (Cronin & Taylor, 1992), and the higher-education-specific HEdPERF (Abdullah, 2006). They offer course evaluation two things most home-grown forms lack — a properly multidimensional structure and a diagnostic logic — while smuggling in two things you should refuse: the "expectations gap" and the full student-as-consumer frame. This piece separates the parts worth borrowing from the parts worth leaving.
The gap model, in one paragraph
SERVQUAL grew out of Parasuraman, Zeithaml and Berry's work in the mid-1980s and was formalised in their 1988 Journal of Retailing paper. It measures service quality across five dimensions — tangibles, reliability, responsiveness, assurance and empathy — using 22 paired items. For each item, the respondent rates both their expectation of an excellent service and their perception of the service they received; quality is scored as the gap between the two (perceptions minus expectations). The intuition is seductive: a service is "good" not in the abstract but relative to what you were promised and what you expected (Parasuraman, Zeithaml & Berry, 1988).
That intuition should sound familiar. It is the same disconfirmation logic that makes course evaluations measure surprise rather than quality — a problem we unpack in the expectation gap. SERVQUAL did not invent the problem; it institutionalised it.
SERVPERF: the same idea, minus the expectations
Cronin and Taylor (1992) argued the expectations half of SERVQUAL was dead weight. Difference scores are notoriously unreliable (subtract one noisy measure from another and you get a noisier one), respondents' "expectations" tend to pile up near the ceiling, and asking every question twice doubles survey length for little gain. Their SERVPERF scale keeps the same five dimensions but measures perceptions only. A large body of subsequent work has found SERVPERF's perceptions-only approach at least as valid as SERVQUAL and considerably more efficient. For a sector already drowning in survey fatigue, halving the item count without losing validity is not a footnote.
HEdPERF: built for universities, not banks
Here is the part most quality officers have never been told. SERVQUAL and SERVPERF were validated in retail banking, dry cleaning, fast food and telecoms — not classrooms. Firdaus Abdullah's HEdPERF (Higher Education PERFormance) was developed specifically for higher education and proposes six dimensions: non-academic aspects, academic aspects, reputation, access, programme issues, and understanding (Abdullah, 2006, International Journal of Consumer Studies). In a head-to-head comparison, Abdullah reported that HEdPERF was the relatively more reliable and valid instrument for higher-education institutions than SERVPERF (Abdullah, 2006, Marketing Intelligence & Planning).
The lesson is not "adopt HEdPERF". It is that borrowing a generic satisfaction scale and bolting it onto a course is exactly the kind of construct mismatch course evaluation keeps making — a validity problem we describe as construct underrepresentation. An instrument built for a bank will faithfully measure a bank-shaped construct.
What course evaluation should actually borrow
1. Real multidimensionality — done, not asserted. Services marketing does not average ten items into one "quality" score and call it done. It models distinct dimensions and validates that they hold together. Most course-evaluation forms claim to measure several things but, because of the halo effect, collapse into one global impression. Borrowing the discipline of dimensional validation — not the specific dimensions — is the upgrade.
2. Diagnosis, not just a score. The gap logic, stripped of its unreliable difference-score arithmetic, is genuinely useful in one form: comparing how much a dimension matters against how well it performed. That is precisely importance-performance analysis, and services marketing has 40 years of practice pointing scarce improvement effort at the quadrant that matters.
3. A validated instrument instead of a committee's favourite questions. The strongest argument HEdPERF makes is simply use something that was validated. It is the same case for adopting SEEQ or IDEA instead of a home-grown form.
But doesn't this just double down on the student-as-consumer model?
This is the strongest objection, and it deserves a straight answer: partly, yes. Importing a services-marketing instrument wholesale risks entrenching the very frame that distorts course evaluation — the idea that a student is a customer and a degree is a purchased service. We have argued at length that the student-as-consumer model misprices most of what a university does. A satisfied customer and an educated graduate are not the same person, and a course that maximises the first can shortchange the second.
Three specific cautions follow. First, the expectations-gap machinery is empirically shaky — difference scores are unreliable and the SERVQUAL-vs-SERVPERF debate has never been settled in SERVQUAL's favour. Second, the dimensional structure is unstable across contexts and cultures; a "responsiveness" factor found in Malaysian universities may not be the same latent variable in a German Fachhochschule, which is exactly the measurement-invariance problem that undermines cross-group comparison. Third, and most fundamentally, service quality is an input-and-experience construct, not a learning construct. It can tell you the seminar room was cold and the feedback was late; it cannot tell you whether anyone learned. For that you need direct measures of learning, not a satisfaction scale, however well validated.
So the honest position is: borrow the engineering (multidimensionality, validation, diagnostic gap analysis), refuse the ideology (the customer frame as the whole truth), and never let a service-quality score stand in for evidence of learning.
Where Koji fits
A five-point grid of "assurance" and "empathy" items has a hard ceiling: it can tell you that responsiveness scored 3.4, never why, and never for whom. That gap between a dimension score and an actionable cause is exactly where a static form stops and where Koji for Education begins.
Koji treats the service-quality dimensions as a scaffold for question design, not a scoring ritual. Its AI-moderated conversational interview can open on a structured item — "How responsive was the teaching team when you got stuck?" — and then probe: what happened, when, what would have helped. Across six structured question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no), you keep the comparability of a validated scale and recover the causal detail a gap score throws away. Automatic thematic analysis then clusters those open responses into the dimensions that actually drove the number, with the underlying quotes attached — so "responsiveness is low" becomes "responsiveness is low because assignment feedback consistently arrives after the next assignment is due." Because the moderation is standardised and bias-aware, you avoid the human-interviewer inconsistency that makes qualitative service-quality research so hard to scale, and reporting rolls up cleanly to programme and institution level.
Many of the same teams run general customer and user research too; the main Koji platform uses the same AI interview engine for exactly that, so the service-quality thinking you build for course evaluation transfers directly.
Borrow what services marketing engineered. Leave the customer myth at the door. And when a dimension score raises a question, use a tool that can actually ask the follow-up.
A short checklist for borrowing responsibly
If you want the engineering without the ideology, four rules keep you honest. Validate before you trust a dimension. Do not assume your "assurance" and "empathy" items hold together as factors just because SERVQUAL's did in a bank; run the factor analysis on your own data or adopt an instrument already validated in higher education. Measure perceptions, not gaps. Follow SERVPERF and drop the expectations half; a single well-worded perception item is more reliable than a difference score and halves respondent burden. Map every dimension to an action owner. A service-quality dimension nobody can act on is a vanity metric; if "access" scores low, someone must own timetabling or the item should not be asked. Never let the composite stand alone. Report the dimensions, their spread and their importance separately, because a single averaged "quality" score reintroduces exactly the halo problem the multidimensional model was meant to solve. Borrowed carefully, service-quality thinking sharpens question design; borrowed lazily, it just dresses a satisfaction survey in marketing vocabulary.
Ready to move from rating dimensions to understanding them? See how Koji for Education works.