Your Course Evaluation Says What Scored Low. Importance-Performance Analysis Says What to Fix First
The lowest-scoring item on your dashboard is not automatically your top priority. Importance-Performance Analysis crosses how well a course did on each attribute with how much students actually weight it — turning a ranked bar chart into a defensible action map.
Koji Education Team
Product ·
The short version: the item with the lowest average on your course-evaluation dashboard is not necessarily the thing you should fix first. A weak score on something students barely weight is a poor use of a programme team's finite attention; a merely-mediocre score on something students weight heavily is often the real emergency. Importance-Performance Analysis (IPA) — introduced by Martilla and James in 1977 — is the simplest rigorous way to tell the two apart. It crosses how well a course performed on each attribute with how much that attribute matters to students, and sorts every item into one of four action quadrants. The output is not another number. It is a priority list you can defend to a committee.
The hidden assumption in every ranked bar chart
Open almost any evaluation dashboard and you will find the same artefact: a list of items sorted from lowest mean to highest. "Assessment and feedback: 3.4. Organisation: 3.9. Enthusiasm: 4.5." The implicit instruction is start at the top of the low end and work down. That instruction encodes an assumption almost nobody states out loud — that every item deserves equal weight, so the lowest score is the most urgent problem.
Students do not weight items equally. A cohort can rate "variety of assessment types" a full point below "clarity of what is expected of me," yet care about the second far more. Fixing the first would move a number on a report; fixing the second would change how students experience the course. Ranking by score alone cannot see the difference, because it has thrown away the one thing that distinguishes a cosmetic problem from a consequential one: importance.
This is not a small caveat. Because most course-evaluation distributions are compressed near the top — the well-documented ceiling effect that makes almost every course score around 4 out of 5 — the gaps between items are often tiny and noisy. Sorting by a 0.2-point difference and calling the loser your "priority" is exactly the kind of over-reading of small differences that good measurement practice warns against. Weighting by importance is what rescues the exercise.
What Importance-Performance Analysis actually does
IPA was first described by John Martilla and John James in the Journal of Marketing in 1977 as a tool for turning satisfaction data into managerial priorities (Martilla & James, 1977). The idea is disarmingly simple. For each attribute you measure, you obtain two coordinates:
- Performance — how well the course did on that attribute (your existing rating).
- Importance — how much that attribute matters to the student's overall judgement.
Plot every attribute on a grid with importance on one axis and performance on the other, draw a crosshair through the two means, and you get four quadrants:
- Concentrate here (high importance, low performance) — your genuine priorities. Improvement here changes how students experience the course.
- Keep up the good work (high importance, high performance) — your strengths. Protect them; do not fix what is not broken.
- Low priority (low importance, low performance) — weak but not worth chasing. The bar-chart trap lives here: these items look like emergencies and are not.
- Possible overkill (low importance, high performance) — you may be over-investing effort students do not value.
The method has been applied to higher education for decades precisely because it disciplines improvement planning: it is described in the literature as "a useful tool for directing continuous quality improvement in higher education" (O'Neill & Palmer; see review in Silva & Fernandes), and continues to appear in recent European studies of student perceptions of their institutions (Importance-Performance study of sustainable universities, 2023).
The clever part: you do not have to ask students what matters
The obvious objection is that measuring importance means adding a second question for every item, doubling the length of an already fatiguing survey. There is a better way, and it is the reason IPA is more than a spreadsheet trick.
Derived importance infers how much each attribute matters from how strongly it correlates with students' overall judgement, rather than asking them to rate importance directly. If "clarity of expectations" moves in lockstep with overall satisfaction across a cohort while "variety of assessment types" barely does, the data itself is telling you which lever matters — no extra questions, and no reliance on students accurately introspecting on their own priorities (which they are famously bad at). Stated importance and derived importance frequently disagree, and the disagreement is itself informative.
A worked example
Suppose a second-year module returns these means (out of 5) and derived-importance weights (correlation with overall satisfaction):
- Feedback usefulness: performance 3.3, importance 0.61 → Concentrate here.
- Workload manageability: performance 3.4, importance 0.28 → Low priority.
- Lecturer enthusiasm: performance 4.6, importance 0.55 → Keep up the good work.
- Variety of media: performance 4.5, importance 0.12 → Possible overkill.
The naïve bar chart says: fix feedback (3.3), then workload (3.4). IPA says: fix feedback, and leave workload alone — it scores low because students do not weight it, and pouring effort into it buys almost nothing. Meanwhile the module is arguably over-investing in production variety no one values. Same data, completely different action list. This is why collecting the feedback is the easy 20% — the hard part is knowing what to do with it.
The strongest counterargument, taken seriously
IPA is not above criticism, and a PhD audience deserves the objections stated plainly.
"Importance and performance are not independent." The most serious methodological critique is that the two axes are entangled. An attribute's measured importance can shift depending on how well it performed — a factor that is failing badly may look important simply because it is salient (a nonlinear, Kano-type relationship). If importance is not stable, the quadrant boundaries move. This is real, and it is why IPA should be read as a decision aid, not an oracle. The honest practitioner reports derived importance alongside the caveat and does not treat a point that sits near the crosshair as firmly classified.
"Where you put the crosshair changes the story." Using scale means, medians, or the scale midpoint to divide the quadrants can reassign items. There is no single correct choice; the defensible move is to state which you used and check whether conclusions survive the alternative.
"It is still a Likert average underneath." Fair. IPA inherits every limitation of the numbers it is built on. If your performance scores are compressed by ceiling effects and your importance weights come from a single closed item, IPA organises noisy inputs more usefully — but it cannot manufacture signal that was never collected. The remedy is not to abandon IPA; it is to feed it richer inputs than a five-point grid.
Where a conversational, AI-native platform changes the inputs
This is the point at which the tooling matters. Traditional survey platforms (EvaSys, Qualtrics exports, paper forms) can compute IPA, but only from the thin data a static form collects: a battery of Likert items and, at best, one overall-satisfaction anchor for the derived-importance correlation. The quadrants are only ever as good as those inputs.
Koji for Education changes the inputs rather than just the arithmetic:
- Derived importance from what students actually dwell on. Because Koji collects feedback through AI-moderated conversational interviews rather than a fixed grid, importance can be inferred not only from correlation with an overall score but from what students spontaneously return to, elaborate, and care enough to explain. That is a far richer importance signal than a checkbox.
- Thematic analysis that populates the grid automatically. Koji's automatic thematic analysis of open-text responses surfaces the attributes students raise in their own words — so the "Concentrate here" quadrant is not limited to the items an administrator thought to ask about. It goes beyond a sentiment score, which on its own tells you almost nothing actionable.
- Quality scoring to weight substance over throwaway. Koji scores the quality of each response, so a considered paragraph counts for more than a one-word placeholder when the priorities are assembled.
- Closing the loop on the priority quadrant. IPA is worthless if the "Concentrate here" list is never acted on. Koji's action-tracking and programme-level reporting turn each quadrant into a tracked commitment rather than a slide that is forgotten by week three.
Teams that also run general customer or user research beyond the classroom can point the same AI interview engine on the main Koji platform at those studies — the derived-importance logic is identical whether the respondent is a student or a customer.
IPA does not eliminate judgement, and Koji does not eliminate IPA's limitations. What both do is move you from "which number is lowest?" to "which improvement will students actually feel?" — which is the only question a quality-improvement cycle should be asking.
Try it on your next cycle
If your evaluation reporting still ranks items by mean alone, you are almost certainly mis-prioritising. Rebuild one module's report as an importance-performance grid and watch how the action list changes. If you want that grid generated automatically from conversational feedback — with the importance weights, the thematic quadrants, and the action tracking built in — see how Koji for Education turns evaluation data into priorities.
Koji surfaces and prioritises — it does not eliminate the need for academic judgement about what to change.