Importance-Performance Analysis: Turning Course-Evaluation Scores into a Priority Map
A course-evaluation report that lists twenty item means tells you nothing about where to act first. Importance-Performance Analysis (IPA) plots each attribute by how much it matters to students against how well you did, producing a four-quadrant map that separates urgent fixes from wasted effort.
Koji Education Team
Product
In brief
A typical course-evaluation report gives you a column of item means and leaves you to guess which weakness matters. Importance-Performance Analysis (IPA) solves this by plotting every teaching attribute on two axes — how important it is to students, and how well the course performed — and reading the result as four quadrants: "Concentrate here", "Keep up the good work", "Low priority", and "Possible overkill" (Martilla & James, 1977; Cladera, 2021). The payoff is a defensible priority map: a low score on a low-importance attribute is not the same emergency as a low score on a high-importance one, and IPA makes that distinction visible. Its main hazard is deciding how to measure importance — stated importance and statistically-derived importance frequently disagree.
What the research says
Importance-Performance Analysis was introduced by Martilla and James (1977) in the Journal of Marketing as a simple, visual way to convert satisfaction data into managerial priorities. The logic is that customer satisfaction is a function of two things: how much an attribute matters, and how well the provider delivers on it. Plotting attributes on an importance (y-axis) × performance (x-axis) grid, with the axes crossed at their means, yields four quadrants:
- Quadrant I — Concentrate here (high importance, low performance): the urgent problems. Fix these first.
- Quadrant II — Keep up the good work (high importance, high performance): your genuine strengths; protect them.
- Quadrant III — Low priority (low importance, low performance): weak, but nobody much cares; do not spend scarce effort here.
- Quadrant IV — Possible overkill (low importance, high performance): you are over-delivering on something students do not value; consider reallocating resources.
The method migrated into higher education because end-of-course questionnaires are exactly the kind of multi-attribute satisfaction data IPA was designed for. Cladera (2021), in Educational Assessment, Evaluation and Accountability, demonstrates the application directly: IPA lets an instructor "obtain a visual representation of what teaching attributes are important for their students, how important each attribute is, and how well the instructors performed on each attribute" (Cladera, 2021). Instead of a flat ranking of means, the lecturer sees which weak attributes are worth acting on and which strengths are being under-recognised. Wohlfart and Hovemann (2019) apply the same framework at programme level to align curricula with graduate-employer expectations, showing IPA scales from a single module up to whole-programme quality improvement.
A crucial methodological fork is how importance is obtained. There are two families:
- Stated (direct) importance — ask students how important each attribute is, on its own scale. Simple, but prone to ceiling effects: students tend to rate almost everything "very important."
- Derived (implicit) importance — estimate importance statistically as the correlation (or regression weight) between each attribute and an overall satisfaction measure. This avoids the "everything matters" problem but assumes a linear, symmetric relationship between attribute performance and overall satisfaction — an assumption the Kano and penalty-reward literatures show is often false.
The two methods routinely place attributes in different quadrants, which is the single most important caveat for anyone using IPA.
Why it matters for course evaluation in practice
Quality-assurance officers drown in numbers and starve for priorities. A programme with fifteen modules, each reporting a dozen item means, generates hundreds of data points per cycle. IPA is valuable precisely because it is a triage tool.
It stops you fixing the wrong thing. The natural human response to an evaluation report is to attack the lowest score. IPA shows why that is often a mistake: if the lowest score is on an attribute students rate as unimportant (Quadrant III), fixing it buys almost no satisfaction. The lowest high-importance score (Quadrant I) is the real emergency.
It reveals over-investment. Quadrant IV — "possible overkill" — is invisible in a standard report. An instructor pouring effort into an elaborate attribute students barely value is a resource-allocation problem IPA surfaces and a mean-ranking never will.
It communicates. A four-quadrant chart is legible to a programme director, a dean, and an external reviewer in seconds, in a way a table of decimals is not. For closing the feedback loop, being able to say "here are the two Quadrant-I attributes we are acting on this year" is far more credible than "we noted the feedback."
It works at multiple altitudes. The same grid works for a single course (which teaching behaviours to improve) and for a programme (which curriculum areas matter to students and employers), making it a natural spine for a continuous-improvement cycle.
Limitations and honest caveats
IPA is intuitive, which is exactly why it is often misused. A critical reader should hold several objections.
The importance measure decides the answer. Because attributes are classified relative to the mean importance and mean performance, and because stated and derived importance disagree, the same data can produce different maps. If you use stated importance and everything clusters near the ceiling, the y-axis becomes noise and the quadrants are arbitrary. Report which importance method you used and, ideally, check robustness across both.
The axes are relative, not absolute. Crossing the axes at the sample means means an attribute's quadrant depends on the other attributes in the set. Add or drop items and things move quadrant. IPA tells you relative priorities within one questionnaire, not absolute quality.
Linearity and symmetry are assumed. Classic IPA treats the importance-performance relationship as linear and symmetric — improving any attribute helps satisfaction equally, in both directions. The Kano model and asymmetric (penalty-reward) analyses show this is frequently untrue: some attributes only hurt when absent (must-be), others only delight when present. Naive IPA can mis-prioritise these. This is why IPA and Kano are complementary, not interchangeable.
Small-sample instability. Quadrant boundaries sit at the means, so with a low response rate the means — and therefore the quadrant assignments near the crosshairs — are unstable. An attribute sitting on a boundary should not drive a decision.
It describes, it does not explain. IPA tells you an attribute is high-importance and low-performance. It does not tell you why performance is low or what to change. That requires qualitative follow-up.
How Koji incorporates this
Koji for Education is designed to feed an IPA workflow end to end rather than leaving a QA officer to reconstruct one from a spreadsheet.
- Both axes, natively. IPA needs a performance rating and an importance signal per attribute. Koji's structured question types cover both: scale items capture performance; ranking and single_choice / multiple_choice items capture stated importance directly, while overall-satisfaction items let derived (correlation-based) importance be estimated. Because both live in one instrument, an importance-performance grid can be assembled without bolting two surveys together.
- Derived importance from real drivers. Where an institution prefers implicit importance, Koji's analysis can relate each attribute to an overall experience measure, mitigating the "everything is very important" ceiling problem that plagues stated-importance IPA.
- The missing "why". IPA's central weakness is that it flags a Quadrant-I problem without explaining it. Koji's AI-moderated conversational interviews close that gap: when a student marks a high-importance attribute low, the moderator probes for the specific cause, and automatic thematic analysis aggregates those probes so each urgent quadrant arrives with evidence about the fix — not just a coordinate.
- Closing the loop, visibly. Because Koji supports action tracking, the two or three Quadrant-I attributes an institution commits to can be recorded and re-measured next cycle, turning the priority map into a documented improvement trail suitable for accreditation evidence.
Koji does not claim IPA is a complete quality system — it is designed to mitigate the "which weakness matters" problem and to supply the qualitative explanation IPA structurally lacks. Teams running product or customer research beyond teaching apply the same importance-versus-performance logic on Koji's core platform at koji.so.
A worked reading
Consider a module whose end-of-course report shows "clarity of explanation" rated 4.4 and marked most important, "assessment feedback" rated 3.1 and important, "use of the virtual learning environment" rated 3.0 and low importance, and "optional enrichment materials" rated 4.5 and low importance. A mean-ranking would send the instructor to attack the two lowest scores — feedback and the VLE — with equal urgency. IPA reads the same four numbers differently: assessment feedback sits in Quadrant I ("Concentrate here") and is the genuine priority; the VLE sits in Quadrant III ("Low priority") and can wait; the enrichment materials sit in Quadrant IV ("Possible overkill"), signalling effort that could be redirected; and clarity sits in Quadrant II, a strength to protect and publicise. Identical data, a completely different action plan — which is the entire reason to plot importance against performance rather than rank performance alone.
Related resources
- Should You Use Net Promoter Score for Courses?
- Best-Worst Scaling (MaxDiff) for Course-Evaluation Priorities
- The SERVQUAL Gap Model, HEdPERF, and Service-Quality Thinking
- Top-Box vs Mean: How to Report Course Evaluation Scores
- Does Closing the Feedback Loop Actually Matter?
- What Can Open-Text Student Comments Tell You That Likert Scores Cannot?
References
- Cladera, M. (2021). An application of importance-performance analysis to students' evaluation of teaching. Educational Assessment, Evaluation and Accountability, 33(4), 701–715. https://doi.org/10.1007/s11092-020-09338-4
- Martilla, J. A., & James, J. C. (1977). Importance-performance analysis. Journal of Marketing, 41(1), 77–79. https://doi.org/10.1177/002224297704100112
- Wohlfart, O., & Hovemann, G. (2019). Using importance-performance analysis to bridge the information gap between industry and higher education. Industry and Higher Education, 33(5), 300–307. https://doi.org/10.1177/0950422219838465
- Abalo, J., Varela, J., & Manzano, V. (2007). Importance values for importance-performance analysis: A formula for spreading out values derived from preference rankings. Journal of Business Research, 60(2), 115–121. https://doi.org/10.1016/j.jbusres.2006.10.009
Related articles
Should You Use Net Promoter Score for Courses? The "Would You Recommend" Question in Higher Education
Net Promoter Score is migrating from customer experience into student feedback. What Reichheld (2003) actually claimed, why Keiningham et al. (2007) failed to replicate its superiority, and whether a single recommend-question belongs in course evaluation.
Top-Box vs Mean: How to Report Course Evaluation Scores Without Throwing Away Information
Reporting the percentage of students who chose the top box feels intuitive, but collapsing a scale to favorable/unfavorable discards information. Here is what the measurement evidence says and how to report responsibly.
Does Closing the Feedback Loop Actually Matter? The Evidence on Acting on Student Evaluations
Universities are good at collecting student feedback and bad at acting on it visibly. The research — Watson (2003), Leckey & Neill (2001), Shah et al. (2017) — shows that failing to close the loop drives the scepticism and declining response rates that quietly destroy your evaluation data.
When Everything Scores 4/5: Best-Worst Scaling (MaxDiff) for Course-Evaluation Priorities
Likert ratings on course evaluations cluster near the top and cannot tell you what matters most to students. Best-worst scaling (Louviere, Flynn & Marley) forces trade-offs that reveal genuine priorities. What the method is, its limits, and how it fits a Koji study.