Net Promoter Score is on more executive dashboards than any other customer metric. It is also, in the form most companies run it, close to meaningless — not because the idea is bad, but because the way it is calculated throws away most of the information and the way it is used ignores the rest.
This is not an argument against measuring customer sentiment. It is an argument for understanding what you are measuring, because a number you trust and should not is considerably more dangerous than no number at all.
The mechanics, precisely
One question: "How likely are you to recommend us to a friend or colleague?" Answered on a scale from 0 to 10.
Respondents are then bucketed:
- Promoters: 9 and 10
- Passives: 7 and 8
- Detractors: 0 through 6
And the score is: % promoters minus % detractors. Passives are discarded entirely. The result runs from −100 to +100.
So a company where every customer scores 8 — solidly satisfied, no complaints — has an NPS of zero. Exactly the same score as a company where half score 10 and half score 0. Our free NPS calculator will run the arithmetic, and the more instructive exercise is to feed it two wildly different distributions and watch them produce the same number.
Lie one: the buckets destroy the data
That example is not a corner case. It is the central defect.
A 6 and a 0 are both detractors. One is a customer who is broadly fine and had one bad week; the other is actively telling people not to buy from you. Treating them as identical is not a simplification, it is an error, and it is baked into the metric.
Meanwhile a 7 and an 8 — customers who like you — contribute nothing at all. And moving a customer from 8 to 9 changes the score by the same amount as moving one from 0 to 9, despite being an enormously smaller achievement.
The arithmetic consequence is that NPS is extremely sensitive to movement across two arbitrary thresholds and completely blind to everything else. A team that improves every customer's experience by one point may see no change at all, or a dramatic one, depending entirely on where their customers happened to be sitting relative to the cut lines.
The average rating, or the full distribution, contains all of this information. NPS deliberately discards it in exchange for a single memorable number, and the trade is much worse than it looks.
Lie two: the sample is not who you think
NPS is measured on the people who answered, and the people who answer are not a random sample of your customers.
Response rates for these surveys are typically low. Who responds? People with a strong feeling — delighted or furious — and people with a lot of free time. The vast, quiet middle, which is most of your customer base, does not answer, because they have nothing to say and something better to do.
This produces a bimodal sample from a population that is not bimodal, and it systematically overstates both promoters and detractors. Your NPS is a measurement of your extremes, presented as a measurement of your customers.
Worse, this bias is not stable. Anything that changes who bothers to respond — an incident, a price change, a survey redesign — changes the score without any underlying change in how customers feel. A great deal of NPS movement is sampling noise wearing a suit.
Lie three: nobody computes the confidence interval
NPS is reported to a whole number and discussed as though a change of three points means something.
It usually does not. With a few hundred responses, the confidence interval on an NPS figure is wide — comfortably wider than the quarter-on-quarter movements that get celebrated or investigated in review meetings. A shift from 34 to 37 is, in most real samples, indistinguishable from no change.
Two things follow. First, do not act on small movements. The urge to explain a three-point drop produces elaborate narratives about causes that do not exist. Second, do not compare small segments. "Enterprise NPS is 12 points higher than SMB" sounds like a finding, and with forty responses per segment it is noise.
If your NPS is computed from fewer than a few hundred responses, treat it as a rough directional indication and nothing more. That is a substantially less exciting claim than the dashboard implies, and it is the honest one.
Lie four: the question is not the one you care about
"Would you recommend us" is a proxy. What you actually want to know is whether they will stay, whether they will buy more, and whether they will in fact recommend you.
Stated intention to recommend is a weak predictor of actual recommendation. People say yes to be pleasant, and then never mention you to anyone. And in plenty of categories, nobody recommends anything to anyone — you may love your email verification vendor and never once bring it up at dinner.
The original claim, from the 2003 Harvard Business Review article that launched NPS, was that this one number predicts growth. That claim has been repeatedly challenged in the academic literature since, and the honest summary is that NPS correlates with growth about as well as other satisfaction metrics — which is to say, somewhat, in some industries, sometimes.
It is not magic. It was marketed as magic, extremely effectively, and the marketing outlived the evidence.
So why use it at all?
Because it has two genuine virtues, and they are not statistical.
It is comparable. Everyone runs the same question the same way, which means an external benchmark exists. That is worth something, even if the underlying number is crude.
It is universally understood. A board understands NPS. Explaining a distribution of ratings to a board is a harder sell than a single number that everyone has already agreed to care about, and organisational alignment has real value even when it is built on an imperfect measure.
These are political virtues, not analytical ones — and that is fine, as long as everyone knows which kind of virtue they are relying on. The failure is not using NPS. It is believing it.
The part that actually matters, which everyone skips
After the 0–10 question there is a second question: why?
That free-text answer is the entire value of the exercise. It is the only part that tells you what to do. The number tells you that something is wrong; the text tells you what.
And it is routinely ignored, because the number is easy to put on a slide and the text is not. Companies collect thousands of free-text responses, compute a score from the numbers beside them, and never read a word.
If you do one thing differently after reading this: read the comments. All of them, or a random sample of a hundred. Not summarised, not sentiment-scored, not turned into a word cloud. Read them, in the customer's own words. It is the cheapest customer research available to any business and almost nobody does it, because it takes an afternoon and produces no chart.
What to measure instead, or alongside
NPS asks about a hypothetical future behaviour. There are better questions available, and they are better precisely because they ask about things that actually happened.
- Actual retention. Did they renew? This is not a proxy for anything. It is the thing itself, and it is already in your billing data.
- Actual referrals. How many customers arrived through a referral? Also already in your data, and it measures the behaviour NPS merely asks about.
- CSAT, after a specific interaction. "How was that support conversation?" Narrow, immediate, actionable, and far less prone to the sampling problems above because the moment of asking is tied to the moment of experience.
- Customer Effort Score. "How easy was it to get this done?" A better predictor of loyalty than satisfaction in much of the research, and it points directly at something you can fix.
None of these will replace NPS on the board slide, and that is fine. Run NPS for the benchmark and the politics. Run these for the truth, and make sure everyone knows which is which.
How to run it so that it is not useless
If you are going to do it, do it in a way that produces information:
- Ask after a real interaction, not on a calendar schedule. A survey that arrives because it is the first of the quarter is answered by whoever happens to be free. Our review request generator drafts the ask, and the reputation score calculator puts the private feedback next to the public reviews, which is where it belongs.
- Ask everyone, not a self-selecting sample. The more you can raise the response rate, the less the bimodal bias distorts you.
- Keep the full distribution. Report the histogram alongside the score. It costs nothing and it is where the information is.
- Report the confidence interval. If it is wider than the change you are discussing, say so out loud in the meeting.
- Read the comments. Actually read them.
- Close the loop. Contact the detractors. Not to change their score — to find out what happened and fix it. This is the single highest-return activity in the whole programme and it is treated as optional.
That last one converts NPS from a measurement exercise into an operational one, which is the only version worth running.
Detractors are a list, not a number
Here is the reframing that makes the whole thing useful.
Review Management exists to run exactly this routing, and Social Proof is what turns the promoter end of it into something a prospect actually sees.
A detractor is not a percentage point. It is a named customer, with an email address, who has just told you they are unhappy and — critically — has just proven they are willing to talk to you, because they answered the survey.
That is not a data point. That is the warmest churn-prevention lead you will ever get, and most companies aggregate it into a score and throw the individual away.
The same applies at the other end. A promoter is a named customer who has just told you they are delighted, which makes them the ideal person to ask for a review, a referral, or a case study — and the moment they answered the survey is the moment of maximum willingness. Our piece on getting more Google reviews makes the case that the moment of delight is everything, and an NPS survey is a machine for identifying exactly that moment at scale.
Used this way, NPS stops being a metric and becomes a routing system: detractors to customer success, promoters to review requests, passives left alone. That is worth running even if you never look at the score.
Relational and transactional NPS are different animals
The literature distinguishes two versions, and conflating them is one of the more common ways an NPS programme goes wrong.
Relational NPS asks about the overall relationship, usually on a schedule — quarterly, annually. It is the version that appears on board slides. It is also the version most exposed to every problem in this article: the sample self-selects, the timing is arbitrary, and the respondent is being asked to summarise a year of experience into one digit, which nobody can actually do.
Transactional NPS asks immediately after a specific event — a support ticket, an onboarding, a delivery. It is far more useful, because the respondent is answering about something concrete that just happened, and because you know exactly what they are answering about.
The trap is averaging them together, or comparing them. A transactional score after a support interaction is measuring your support team. A relational score is measuring your product, your price, your competitors, and the mood of whoever opened the email. These are not the same quantity and one cannot be benchmarked against the other.
If you run only one, run the transactional version. It is closer to the experience, closer to something you can fix, and far less prone to the sampling bias that makes the relational number so unreliable.
Benchmarks, and why yours is probably not comparable
The great selling point of NPS is that you can compare yourself to others. In practice this is much shakier than it appears.
Scores vary enormously by industry — in some sectors a score of 30 is excellent and in others it is mediocre — and they vary by geography in a way that has nothing to do with satisfaction. Respondents in some cultures avoid the extremes of a scale as a matter of habit, which mechanically depresses the promoter count without any difference in how customers actually feel.
They also vary by how you asked. A survey sent by email to everyone produces a different score from an in-app prompt shown to active users, which produces a different score again from a question asked by a human at the end of a call. None of these differences are about the customer.
So a benchmark comparison is only meaningful when the methodology matches, and it almost never does, because the published benchmarks do not disclose their methodology in enough detail for you to check.
Use benchmarks the way you would use a weather forecast for a different city: mildly informative, not a basis for a decision. Your own score, measured consistently over time with a stable method, is far more useful than any comparison to a number whose provenance you cannot inspect.
Survey fatigue, and the cost of asking
Every survey has a price, and it is not the software licence.
Asking costs you a small amount of the customer's goodwill. Asking repeatedly costs more. Asking repeatedly and then visibly doing nothing about the answers costs a great deal — because you have now demonstrated that the survey is theatre, and the next one will be ignored by the people whose opinion you most needed.
This produces a slow, invisible degradation. Response rates fall, the remaining respondents skew further toward the extremes, and the score becomes progressively less meaningful, all while the dashboard continues to display a number to one decimal place.
Three rules that keep the cost down:
- Do not ask the same customer more than a couple of times a year. Suppress recent respondents.
- Do not ask if you are not going to read the answers. An unread survey is worse than no survey, because it is a promise you broke.
- Close the loop visibly. When you fix something a customer raised, tell them. This is the only thing that makes the next survey worth answering.
What good looks like
A functioning customer-sentiment programme, in practice, looks almost nothing like a dashboard.
It looks like a transactional survey fired after a real event, with a low-friction rating and one open question. It looks like every detractor being contacted by a human within a couple of days — not to argue, but to ask what happened. It looks like every promoter being asked, at that exact moment, for a review or a referral, because that is the moment of maximum willingness and it will not come again.
It looks like someone reading a hundred comments an afternoon a month, and it looks like a short list of the things customers actually said, circulated to the people who can change them.
And somewhere off to the side, once a quarter, it produces a number that goes on a slide. That number is the least valuable output of the entire system, and in most companies it is the only one that survives.
The honest summary
NPS is a crude, statistically fragile, widely-misunderstood metric that discards most of its own data and measures a hypothetical rather than a behaviour.
It is also comparable across companies, universally understood, and — if you read the comments and act on the individuals rather than the average — capable of surfacing exactly the two lists you most want: the customers about to leave, and the customers ready to advocate.
Both statements are true. The mistake is to hold only one of them. Run it, benchmark with it, and never, ever make a decision because the number moved three points — because in almost every real sample, it did not.
The scale itself is doing strange things
An underappreciated problem: the 0–10 scale is not used the way its designers assumed, and the way people use it varies systematically in ways that have nothing to do with satisfaction.
Many respondents treat the scale as if it were a school grade. In that frame, 7 out of 10 is a perfectly respectable mark — a pass, a solid B — and someone awarding it believes they have given you a compliment. NPS files them as a detractor. The customer thinks they said “good”; the metric records “actively harmful”. Nobody involved is aware of the disagreement.
There is also a well-documented tendency for respondents in some cultures to avoid the extremes of any scale, which mechanically suppresses the 9s and 10s that NPS depends on entirely. A company with customers in those markets will post a lower score than an identical company elsewhere, and no amount of improvement will close a gap that was never about the product.
And there is the number-preference effect: people gravitate to round numbers and to the ends of scales, so 0, 5 and 10 are over-represented relative to 1, 4 and 9. Since NPS cares intensely about the 9/10 boundary and the 6/7 boundary, this quantisation lands directly on the thresholds that determine the score.
None of this is fixable within NPS, because it is a property of the instrument rather than of the analysis. It is simply worth knowing that the number is noisier than it looks before you build an incentive scheme on top of it — which, if you do, will produce exactly the behaviour you would expect: employees asking customers to give them a 9 or a 10, which corrupts the measurement and teaches your customers that the survey is a favour rather than a question.
If you only remember one thing
The score is the least interesting thing your survey produces. Beneath it sit two lists: the people about to leave, and the people ready to advocate for you. Both are named, both are reachable, and both have just proven they will engage with you, because they answered.
Aggregating those two lists into a single integer and putting it on a slide destroys almost everything of value in them. It is, in a fairly precise sense, taking the most actionable customer intelligence you will ever be handed and compressing it until it fits in a box on a dashboard, at which point it can no longer be acted on by anybody.
Keep the number if it helps you talk to your board. But work the lists. That is where the retention is, and that is where the referrals are, and no amount of movement in the headline figure will substitute for a conversation with the customer who told you exactly what was wrong and then waited to see whether anyone would call.
A score is only useful if it changes what you do next. Govarova Review Management turns promoters into public reviews, Sequences runs the survey and the follow-up, and you can start free.