Are Students Customers? The SERVQUAL Gap Model, HEdPERF, and What Service-Quality Thinking Adds to Course Evaluation
The SERVQUAL gap model and its higher-education variant HEdPERF measure the distance between what students expect and what they perceive they received. We assess the evidence, the sharp limitations of treating students as customers, and how Koji uses expectation framing without collapsing learning into satisfaction.
Koji Education Team
Product
In brief: SERVQUAL, the service-quality instrument from Parasuraman, Zeithaml and Berry (1988), measures quality as the gap between customer expectations and perceptions across five dimensions (tangibles, reliability, responsiveness, assurance, empathy). Applied to universities, it reframes course evaluation as expectation-disconfirmation, and a higher-education-specific variant, HEdPERF (Abdullah, 2006), outperformed generic SERVQUAL on reliability and validity in HE settings. The model is genuinely useful for diagnosing where a programme falls short of student expectations — but it is dangerous if taken literally, because a student is not a customer and "satisfying expectations" is not the same as "producing learning." The disciplined use is to treat the gap model as one lens among several, never as the definition of teaching quality.
What the research says
The service-quality tradition begins with Parasuraman, Zeithaml and Berry's (1988) SERVQUAL: A multiple-item scale for measuring consumer perceptions of service quality (Journal of Retailing, 64(1), 12–40). Its central idea is that quality is not an absolute property of a service but a gap: perceived quality equals the customer's perception of what was delivered minus their prior expectation of what should be delivered. SERVQUAL operationalises this across five dimensions — tangibles (facilities, materials), reliability (delivering what was promised, dependably), responsiveness (willingness to help promptly), assurance (knowledge and credibility that inspire trust), and empathy (individualised attention). Each is measured twice: once for expectations, once for perceptions. A negative gap (perception below expectation) signals a quality shortfall.
Higher education adopted the model enthusiastically, because universities increasingly operate in competitive, fee-paying markets where student expectations are explicit. But generic SERVQUAL did not always fit the classroom. Abdullah (2006), Measuring service quality in higher education: HEdPERF versus SERVPERF (Marketing Intelligence & Planning, 24(1), 31–47), tested a purpose-built higher-education instrument (HEdPERF) against the performance-only SERVPERF scale and a hybrid, using 381 usable responses (a 68% response rate) from public and private institutions in Malaysia. He reported that HEdPERF performed better on unidimensionality, reliability, validity, and explained variance within the higher-education setting — arguing that authentic determinants of HE service quality (academic aspects, non-academic aspects, reputation, access, and programme issues) are not fully captured by an instrument built for retail and banking.
The recurring empirical finding across the HE service-quality literature is the negative quality gap: students' perceptions consistently fall below their expectations on most or all dimensions. More recent work continues to confirm that perceived service quality is strongly associated with satisfaction. Seitova et al. (2024), Perceived service quality and student satisfaction (Frontiers in Education, 9, 1492432), surveyed 350 students at a Kazakh-Turkish university and, using structural equation modelling, found a large positive relationship between perceived service quality and satisfaction (standardised coefficient ~0.87, p < 0.01), with reliability the strongest dimension and empathy the weakest. The performance-only critique — that measuring expectations separately adds noise and that perceptions alone predict satisfaction well — traces back to Cronin and Taylor (1992), Measuring service quality: a reexamination and extension (Journal of Marketing, 56(3), 55–68), whose SERVPERF instrument dropped the expectations half of the equation.
So the evidence base says three things: (1) the gap framing reliably surfaces a structured map of where expectations are unmet; (2) HE-specific instruments fit universities better than generic ones; and (3) service quality robustly predicts satisfaction — which is precisely where the caution begins.
Why it matters for course evaluation in practice
Service-quality thinking adds something most SET instruments lack: an explicit standard of comparison. A 4.0/5 on "the course was well organised" is uninterpretable in isolation; a gap score that says "students expected 4.6 and perceived 3.9" is diagnostic. For a quality-assurance office, this has practical value:
- It prioritises. Gap scores rank dimensions by the size of the shortfall, helping a programme decide where remediation has the greatest return rather than chasing every below-average item.
- It captures the non-teaching envelope. Much of what shapes a student's course experience — timetabling, feedback turnaround, administrative responsiveness, learning resources — is not the lecturer's delivery. SERVQUAL/HEdPERF dimensions (responsiveness, tangibles, access) make these visible, which a lecturer-centric SET form hides.
- It speaks the language of strategy. In a marketised European HE sector with the National Student Survey, fee competition, and reputation pressures, expectation-management is a real lever: gaps can be closed by improving delivery or by setting accurate expectations up front.
Limitations and honest caveats
The "students are customers" frame is where a critical reader should push hardest, and the methodological objections are serious.
-
Satisfaction is not learning. The single most important caveat: service-quality models optimise satisfaction, and the SET bias literature (Uttl and colleagues, the active-learning "feeling of learning" studies) shows satisfaction can be negatively related to effortful, effective teaching. A demanding, cognitively challenging course may produce a large negative gap (students expected ease, perceived difficulty) while producing excellent learning. Treating the gap as the goal can actively reward grade inflation and reduced rigour.
-
The student-as-customer metaphor is contested. Unlike a retail customer, a student co-produces the outcome; expectations may be uninformed about what good pedagogy requires; and the "product" (a credential and competence) is partly the student's own work. Bunce and colleagues' research on the consumer mindset suggests stronger consumer orientation is associated with lower academic performance.
-
Measuring expectations is psychometrically fraught. Difference scores (perception minus expectation) have well-known reliability problems, expectations are often inflated and unstable, and post-hoc expectation ratings can be contaminated by the experience just had. This is exactly why Cronin and Taylor's SERVPERF dropped expectations — and why some argue the "gap" is statistical theatre.
-
Dimensional fit varies. Even HEdPERF's structure does not replicate identically across cultures and institution types; measurement invariance cannot be assumed when comparing gap scores across countries or programmes.
-
Common-method and halo contamination remain. A single satisfied or dissatisfied student rates every dimension through one affective lens, so high inter-dimension correlations may reflect a halo, not five distinct quality facets.
The honest position: service-quality thinking is a useful diagnostic supplement for the administrative and resource envelope around teaching, not a definition of educational quality.
How Koji incorporates this
Koji for Education borrows the useful part of the service-quality tradition — the expectation/perception framing and the structured, multi-dimensional diagnosis — while refusing the part that collapses learning into satisfaction.
-
Expectation framing without naive gap scores. Koji can field expectation and experience items using its
scaleandsingle_choicetypes, but rather than reporting raw difference scores (with their known reliability problems), Koji's AI-moderated interview probes why a perception fell short of an expectation: was the expectation unrealistic, was delivery genuinely poor, or was the course appropriately hard? This is designed to separate "the course failed me" from "the course challenged me," which a bare gap score conflates. -
Separating the envelope from the learning. Koji's structured question banks let an institution evaluate the service envelope (responsiveness, resources, feedback turnaround) and learning-oriented constructs in the same instrument, then report them distinctly — so a low responsiveness score does not contaminate the read on whether learning happened, and a demanding course is not penalised as a "service failure."
-
Thematic analysis against service dimensions. Open-text comments are automatically themed, and those themes can be mapped to dimensions such as reliability, responsiveness, and access — turning "I never heard back about my assignment" into a quantified responsiveness signal rather than an unread sentence.
-
Bias-aware, anti-satisficing reporting. Because service-quality measures are vulnerable to halo and common-method bias, Koji is designed to triangulate across cohorts and flag affect-driven straight-lining, rather than presenting five dimension means as five independent facts.
-
Closing the loop on expectations. Gaps can be closed by setting accurate expectations, not only by changing delivery; Koji's action tracking records whether a programme adjusted communication or delivery and whether the next cohort's perceptions moved.
These are mechanisms designed to mitigate the satisfaction-equals-quality trap; none of them eliminate it, and Koji's reporting is explicit that satisfaction and learning are different constructs. Koji's core research platform at koji.so uses the same expectation-versus-experience interviewing approach for product and customer research, where the gap framing is on firmer ground because the customer relationship is genuine.
Related Resources
- Should You Use Net Promoter Score for Courses?
- The Student-as-Consumer Effect on Course Evaluations
- Is One "Overall" Question Enough? Single-Item vs Multi-Item Measures
- What Do Student Evaluations Actually Measure? Marsh and Multidimensionality
- Interpreting and Reporting Student Ratings Responsibly
- Does Closing the Feedback Loop Actually Matter?
References
- Parasuraman, A., Zeithaml, V. A., & Berry, L. L. (1988). SERVQUAL: A multiple-item scale for measuring consumer perceptions of service quality. Journal of Retailing, 64(1), 12–40.
- Abdullah, F. (2006). Measuring service quality in higher education: HEdPERF versus SERVPERF. Marketing Intelligence & Planning, 24(1), 31–47. https://doi.org/10.1108/02634500610641543
- Cronin, J. J., & Taylor, S. A. (1992). Measuring service quality: a reexamination and extension. Journal of Marketing, 56(3), 55–68. https://doi.org/10.1177/002224299205600304
- Seitova, M., Temirbekova, Z., Kazykhankyzy, L., Khalmatova, Z., & Celik, H. E. (2024). Perceived service quality and student satisfaction: a case study at Khoja Akhmet Yassawi University, Kazakhstan. Frontiers in Education, 9, 1492432. https://doi.org/10.3389/feduc.2024.1492432
Related articles
Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees
The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.
Should You Use Net Promoter Score for Courses? The "Would You Recommend" Question in Higher Education
Net Promoter Score is migrating from customer experience into student feedback. What Reichheld (2003) actually claimed, why Keiningham et al. (2007) failed to replicate its superiority, and whether a single recommend-question belongs in course evaluation.
Does Closing the Feedback Loop Actually Matter? The Evidence on Acting on Student Evaluations
Universities are good at collecting student feedback and bad at acting on it visibly. The research — Watson (2003), Leckey & Neill (2001), Shah et al. (2017) — shows that failing to close the loop drives the scepticism and declining response rates that quietly destroy your evaluation data.
What Do Student Evaluations Actually Measure? Marsh, the SEEQ, and the Case for Multidimensional Feedback
Herbert Marsh spent decades showing that well-built student evaluations are multidimensional, reliable, and stable — and that a single global score throws away most of what they can tell you. What his work establishes, where critics push back, and how to design feedback that is actually usable.