New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Methodology8 min read

Should Universities Use Net Promoter Score for Course Evaluation? A Methodological Critique

NPS is migrating from corporate dashboards into student-experience surveys. But the metric's own academic literature shows it predicts no better than ordinary satisfaction measures — and the way it collapses a distribution into one number is exactly wrong for course evaluation. A clear-eyed look at when, if ever, NPS belongs in higher education.

Koji Education Team

Product ·

Bottom line up front: The Net Promoter Score (NPS) is increasingly pitched to universities as a simple "one number" for student experience. Before adopting it for course or programme evaluation, institutions should know that NPS has never demonstrated superior predictive power even in the commercial setting it was built for — a peer-reviewed Journal of Marketing study found it "performs no better than" conventional satisfaction measures — and that its core mechanic of collapsing an 11-point distribution into a single net figure discards exactly the information a teaching-and-learning centre needs. NPS can be a useful conversation-starter for whole-institution sentiment. It is a poor instrument for evaluating a course.

What NPS actually is

NPS was introduced by Fred Reichheld in a 2003 Harvard Business Review article titled, with no false modesty, The One Number You Need to Grow. Respondents answer a single question — "How likely are you to recommend X to a friend or colleague?" — on a 0–10 scale. They are then bucketed:

  • Promoters (9–10)
  • Passives (7–8)
  • Detractors (0–6)

The score is % Promoters − % Detractors, yielding a number from −100 to +100. Its appeal is obvious: one question, one figure, easy to put on a dashboard and track over time.

The problem the marketing literature already found

Reichheld's central claim was that this single number was the best predictor of growth. That claim did not survive replication. In a longitudinal study published in the Journal of Marketing, Keiningham, Cooil, Andreassen and Aksoy demonstrated that "the Net Promoter metric … performs no better than other measures of customer satisfaction and loyalty in predicting company growth." The American Customer Satisfaction Index and ordinary satisfaction items correlated with revenue at least as strongly. Subsequent replication efforts, summarised by research practitioners such as MeasuringU, have repeatedly failed to reproduce NPS's claimed superiority.

In other words: even on its home turf, NPS is not magic. It is a satisfaction question with an unusual scoring rule. That matters, because the entire reason to import a corporate metric into higher education is the promise that it carries special validity. It does not.

Why the scoring rule is wrong for course evaluation

Set aside predictive validity and look at the mechanics. NPS does three things that are actively harmful for evaluating teaching:

1. It throws away the distribution. A class where every student is a lukewarm passive (all 7s and 8s) and a class split between enthusiasts and the alienated can produce the same NPS. For improving a course these are completely different situations — yet the metric erases the difference. This is the same fallacy we describe in why averaging Likert scores misleads, in an even more lossy form: NPS does not just average an ordinal scale, it bins it into three crude categories first.

2. The cut-points are arbitrary and punitive. Treating a 6 as a "detractor" identical to a 0, while a 7 is a neutral "passive," imposes a steep, unjustified discontinuity. A student who rates a course 6/10 is mildly positive in most people's reading, not a detractor. On the ordinal scales used in evaluation, these thresholds have no measurement-theoretic basis.

3. "Would you recommend this course?" is the wrong question. Course evaluation exists to improve teaching and assure quality, not to measure word-of-mouth marketing. Recommendation intent conflates the instructor, the timetable, the difficulty, the cohort and the student's own goals. It is a satisfaction-and-marketing construct, not a diagnostic of learning or teaching quality. It tells a programme director almost nothing actionable about what to change next term.

"But isn't a simple, trackable number useful?"

This is the honest counterargument, and there is something to it.

A single longitudinal number is easy to communicate to senior management and governors, and trend lines do have value — a sudden drop in whole-institution NPS is a legitimate signal that something is wrong somewhere. For institution-level sentiment, tracked over time and never over-interpreted, NPS is defensible as one lightweight indicator among many. It is cheap, fast, and people understand it.

The error is the leap from "useful headline indicator" to "course evaluation instrument." A thermometer is useful; it is not a diagnosis. The moment you need to know why a course is struggling, what to change, or whether an individual instructor is effective, NPS gives you nothing — and its three-bucket scoring will actively mislead anyone who reads a −10 as a verdict on a lecturer. Use it, if at all, as a top-of-funnel sentiment gauge, and keep it well away from personnel and programme-improvement decisions.

A worked example: two classes, one score

Consider two seminars of 20 students. In the first, every student rates the course a 7 or 8 — mildly content, no one alienated, no one evangelical. NPS treats all of them as passives, so the score is 0. In the second, ten students rate the course a 10 and ten rate it a 6. Half are enthusiasts; half are quietly dissatisfied. NPS counts ten promoters and ten detractors, so the score is again 0. Two profoundly different teaching situations — one stable and adequate, one polarised and possibly fracturing in ways worth understanding — are reported as identical.

Now watch the threshold do damage. Move a single student in the second class from 6 to 7 and the score jumps from 0 to +5; move one from 7 to 6 in the first class and it falls from 0 to −5. A one-point shift on an ordinal scale, well within normal response noise, swings the headline figure by five points in either direction. For a metric meant to be tracked term over term and reported to leadership, that volatility is not a rounding detail — it is the difference between a course that looks like it is "improving" and one that looks like it is "in decline," generated by nothing more than the arbitrary placement of the promoter/detractor cut. No teaching-and-learning centre should make decisions on a number that behaves this way.

What good course evaluation measures instead

The constructs that actually help are diagnostic and specific:

  • Multidimensional, behaviour-anchored items rather than a single global recommendation — clarity of explanation, usefulness and timeliness of feedback, alignment of assessment with learning outcomes.
  • The full distribution and its uncertainty, so a polarised cohort is visible rather than averaged or netted away.
  • Rich qualitative feedback that explains the numbers — the single most valuable and most underused part of any evaluation.
  • Triangulation with peer observation and learning evidence, because no student-reported number is sufficient alone (see triangulation in teaching evaluation).

How Koji approaches it

Koji for Education is built around the opposite premise to NPS: that the reasons behind a rating are more valuable than the rating, and that good evaluation surfaces structure rather than collapsing it.

  • Instead of one recommendation question, Koji runs AI-moderated conversational interviews using six structured question types (open-ended, scale, single- and multiple-choice, ranking, yes/no), so a low score is immediately followed by "what specifically, and why?"
  • Automatic thematic analysis turns open-text answers into ranked, actionable themes — the diagnostic layer NPS structurally cannot provide.
  • Distribution- and theme-level reporting at course, programme and institution level keeps the polarised-cohort case visible instead of netting it to a single figure.
  • For institutions that genuinely want a simple longitudinal sentiment line for governors, Koji can still capture a scale item — but it is contextualised by the qualitative evidence rather than presented as a standalone verdict.

The same conversational interview engine powers general user and customer research on the main Koji platform; the education product applies it to the specific methodological demands of course and programme evaluation.

If you have been asked to "add an NPS to the student survey," it is worth asking what decision the number is meant to inform — and whether a method built to diagnose teaching would serve that decision better. See how Koji for Education does it.

The takeaway

NPS is a fine headline gauge and a poor evaluation instrument. Its own marketing literature shows it is not specially predictive; its scoring rule discards the distribution that course improvement depends on; and "would you recommend" is simply the wrong question for assuring teaching quality. Track it at the institution level if you must, but do not let a corporate marketing metric quietly become how your university decides what good teaching looks like.