New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to docs
analysis-reporting8 min read

Should You Use Net Promoter Score for Courses? The "Would You Recommend" Question in Higher Education

Net Promoter Score is migrating from customer experience into student feedback. What Reichheld (2003) actually claimed, why Keiningham et al. (2007) failed to replicate its superiority, and whether a single recommend-question belongs in course evaluation.

Koji Education Team

Product

In brief: Net Promoter Score (NPS) — built from a single "how likely are you to recommend?" question — is increasingly proposed as a lightweight course-evaluation metric. The evidence does not support treating it as a superior or sufficient measure: the peer-reviewed replication by Keiningham et al. (2007) failed to confirm Reichheld's (2003) claim that NPS is the best predictor of growth, and the score discards information by collapsing an 11-point scale into three categories. NPS can be a useful, comparable engagement signal in course evaluation, but it should never replace multidimensional, diagnostic feedback.

Why this question keeps coming up

Quality teams are under pressure to shorten surveys and to produce a single, board-friendly number that trends over time. Net Promoter Score offers exactly that promise: one question, one number, instant benchmark. It is unsurprising that "Would you recommend this course/programme to a fellow student?" is showing up in module evaluations, programme dashboards, and student-experience reports across European higher education. The question is whether the metric earns the authority that is being placed on it.

What the research says

NPS originates with Reichheld (2003), "The One Number You Need to Grow" (Harvard Business Review, 81(12), 46-55). Reichheld reported that, across industries, responses to a single recommendation question — scored 0 to 10 — correlated with subsequent company growth at least as well as longer satisfaction batteries. The operational recipe is well known: classify 9-10 as promoters, 7-8 as passives, 0-6 as detractors, and compute NPS as the percentage of promoters minus the percentage of detractors, yielding a figure from -100 to +100.

The headline claim — that this single number is the best predictor of growth — did not survive independent scrutiny. Keiningham, Cooil, Andreassen and Aksoy (2007), "A Longitudinal Examination of Net Promoter and Firm Revenue Growth" (Journal of Marketing, 71(3), 39-51), used longitudinal data from 21 firms and more than 15,500 interviews in the Norwegian Customer Satisfaction Barometer to replicate Reichheld's analyses. Using the very industries Reichheld cited as exemplars, they failed to replicate the asserted superiority of NPS over established satisfaction measures such as the American Customer Satisfaction Index. The paper won the 2007 Marketing Science Institute / H. Paul Root Award, signalling that the discipline took the rebuttal seriously.

A further methodological critique runs through the literature: NPS may rest on a correlation-as-causation inference, and its category cut-points (the 6/7 and 8/9 boundaries) are not empirically derived, so two distributions with very different shapes can yield the same score. Independent replications of the customer-loyalty literature have likewise found that single-item recommendation does not consistently out-predict multi-item satisfaction. The honest summary is that NPS is a serviceable, comparable loyalty proxy, not the uniquely powerful instrument its marketing implies.

Why it matters for course evaluation in practice

Translating this to a teaching context, three implications follow.

First, a recommend-question measures advocacy, not learning or teaching quality. A student may enthusiastically recommend a module because it was enjoyable, well-organised, or lightly assessed, while learning relatively little — exactly the gap documented in the wider SET-versus-learning literature. Advocacy is a legitimate construct (it relates to programme reputation, recruitment, and word-of-mouth), but it is not a substitute for diagnostic teaching feedback.

Second, the three-box collapse throws away information. Reducing an 11-point scale to promoter/passive/detractor discards the granularity that committees need to detect small changes and to act. A module can hold a steady NPS while the reasons behind it shift entirely.

Third, NPS is genuinely useful as a comparable, low-burden tracking signal — provided it is reported with its category breakdown and, ideally, an accompanying open-text "why". The recommendation question travels well across modules and cohorts, is intuitive to non-specialist readers, and works as a high-level trend line on top of a richer instrument, not in place of one.

For accreditation and quality-cycle purposes, NPS alone is thin evidence. Reviewers under ESG-aligned frameworks expect to see that feedback was acted on; a single advocacy number does not show closing-the-loop. It is best treated as one indicator in a triangulated picture.

Limitations and honest caveats

Several caveats deserve emphasis for a critical reader. The Keiningham et al. (2007) replication was conducted in commercial markets; its direct external validity to higher education is an extrapolation, not a demonstration — but the structural critiques (information loss, arbitrary cut-points, correlation-causation) are domain-general and apply with full force to course evaluation.

The 0-10 scale also behaves differently across cultures: response-style differences mean that a Nordic cohort and a Southern-European cohort can produce different NPS values for an identically experienced course, a problem we examine under cross-cultural response styles. Comparing programme NPS across countries without measurement-invariance evidence is therefore hazardous. Finally, small class sizes make NPS extremely volatile: with twelve respondents, a single detractor can swing the score by double digits, so confidence intervals matter as much here as for any mean.

None of this makes NPS worthless. It makes NPS a bounded tool: good for trend-tracking and stakeholder communication, poor as a sole or high-stakes measure.

There is also a governance dimension that higher education should not overlook. Customer-experience teams adopted NPS partly because it is easy to set targets against and to tie to incentives. Importing that logic into teaching is risky: the moment a recommend-score becomes a target attached to staff appraisal, it invites the same gaming and grade-leniency pressures documented across the wider student-evaluation literature, while rewarding the affective qualities that drive advocacy over the harder-to-love features of demanding, high-learning courses. Used as a private, formative trend that prompts a conversation about why a cohort would or would not recommend a module, the recommend-question is healthy; used as a public league-table KPI, it reproduces every pathology the course-evaluation field has spent two decades cataloguing.

How Koji incorporates this

Koji's design philosophy is the opposite of one-number reductionism, which is precisely why it can use a recommendation question well.

  • NPS as one structured item, never the whole study. Koji supports scale and single_choice question types, so a recommend-question can sit alongside diagnostic items rather than standing alone. The metric becomes a comparable headline, not the entire evaluation.
  • The "why" behind every score. Where classic NPS leaves you guessing why a student is a detractor, Koji's AI-moderated conversational interview probes the rating in the student's own words and applies automatic thematic analysis across the cohort — restoring exactly the diagnostic information the three-box collapse discards.
  • Bias-aware, uncertainty-aware reporting. Koji's reporting is designed to surface category breakdowns and small-sample volatility rather than a bare score, mitigating the false-precision risk that the Keiningham critique highlights.
  • Triangulation and closing the loop. Because Koji captures structured ratings, open text, and cohort comparisons together, an institution can pair an NPS trend with the evidence of action that accreditors expect — turning advocacy tracking into part of a defensible quality cycle.

Koji frames NPS as a signal to be contextualised, not a verdict — designed to mitigate the over-interpretation the research warns against. Teams that also run customer and product research will recognise the pattern: the same AI-moderated interview engine powering koji.so lets product teams move past a bare NPS to the reasons behind it.

A worked example: reading an NPS for a 30-student module

Suppose a module receives 30 responses to "How likely are you to recommend this module to a fellow student?" Eighteen students answer 9-10 (promoters), six answer 7-8 (passives), and six answer 0-6 (detractors). The NPS is 60% minus 20%, or +40 — a figure a dashboard will render as a confident green. But look at what the single number hides. With only 30 respondents, the score is fragile: if two of those detractors had instead answered 7, the NPS would jump to roughly +47, and if three promoters had a worse week and answered 8, it would fall to +30. A swing of this size, driven by a handful of students, would in a customer-experience report be read as a "trend". In a small classroom it is sampling noise.

Now add the diagnostic layer the bare score omits. Imagine the six detractors all wrote, in open text, that feedback on coursework arrived too late to help with the next assignment, while the eighteen promoters praised the lecturer's enthusiasm. The +40 tells you nothing actionable; the reasons tell you to fix feedback turnaround. This is the core lesson of the Keiningham et al. (2007) critique applied to teaching: the recommend-question can track sentiment over time, but on its own it neither explains nor justifies a decision. A defensible report therefore pairs the NPS with its category counts, its response total, an indication of volatility, and the thematic summary of why students answered as they did — exactly the package a quality committee needs and a single number withholds.

Related Resources

References

  • Reichheld, F. F. (2003). The One Number You Need to Grow. Harvard Business Review, 81(12), 46-55. https://hbr.org/2003/12/the-one-number-you-need-to-grow
  • Keiningham, T. L., Cooil, B., Andreassen, T. W., & Aksoy, L. (2007). A Longitudinal Examination of Net Promoter and Firm Revenue Growth. Journal of Marketing, 71(3), 39-51. https://doi.org/10.1509/jmkg.71.3.039
  • Reichheld, F. F., & Markey, R. (2011). The Ultimate Question 2.0: How Net Promoter Companies Thrive in a Customer-Driven World. Harvard Business Review Press.

Related articles

analysis-reporting

Interpreting and Reporting Student Ratings Responsibly: What Linse (2017) Tells Evaluation Committees

The biggest threat to fair evaluation is not the survey — it is how committees read it. Linse (2017) and Boysen (2015) show administrators routinely over-interpret tiny mean differences. We translate the research into concrete rules for reporting student ratings so decisions are defensible.

best-practices

Does Closing the Feedback Loop Actually Matter? The Evidence on Acting on Student Evaluations

Universities are good at collecting student feedback and bad at acting on it visibly. The research — Watson (2003), Leckey & Neill (2001), Shah et al. (2017) — shows that failing to close the loop drives the scepticism and declining response rates that quietly destroy your evaluation data.

research-methods

What Do Student Evaluations Actually Measure? Marsh, the SEEQ, and the Case for Multidimensional Feedback

Herbert Marsh spent decades showing that well-built student evaluations are multidimensional, reliable, and stable — and that a single global score throws away most of what they can tell you. What his work establishes, where critics push back, and how to design feedback that is actually usable.

research-methods

Response Styles and Likert Scales: Why Cross-Cultural Evaluation Needs More Than Numbers

Acquiescence and extreme response styles vary systematically by culture (Harzing, 2006; Baumgartner & Steenkamp, 2001), which means raw Likert averages are not directly comparable across nationalities in Europe's multinational classrooms. This article explains the evidence and how to evaluate fairly across diverse cohorts.