Business Growth

CX Measurement That Matters: NPS vs CSAT vs CES for Service Firms

Fred Reichheld's 2003 Harvard Business Review article promised 'the one number you need to grow,' and two decades later most service firms still run their client experience on a single survey question. The problem is that the strongest replication study in the literature could not reproduce the claim, and competing metrics, customer satisfaction (CSAT) and customer effort score (CES), each won their own evidence battles in different contexts. For a 5-7 figure service business with thirty or fifty active clients, the question is not which acronym wins, but which combination of signals actually predicts renewal, expansion, and referral with a sample size that small. This article walks through the evidence on all three metrics, the critiques, and a layered measurement stack built for small firms.

Joshua Agonya Pi'Rwot

By Joshua Agonya Pi'Rwot

Founder, Business Growth Accelerator

Executive summary

NPS promised one number to rule them all; the replication evidence disagrees. This guide weighs NPS, CSAT, and CES research and assembles a lean measurement stack for service firms that ties scores to revenue decisions.

Section 1

The five challenges at a glance

Service firms do not fail at CX measurement because they lack survey tools; they fail because they import enterprise measurement habits into a small-sample business. The five recurring failure modes below come straight from the academic record and from how 5-7 figure firms actually operate. Each one has a documented root cause and an evidence trail. The pattern worth noticing: most failures involve treating a measurement instrument as a strategy. A score is a thermometer, not a treatment plan. Reichheld himself positioned NPS as a management system rather than a number, yet the number is what most operators adopted. Meanwhile the replication literature, the effort research, and the complaint-handling studies each point to a different part of the experience that the dominant metric misses. The table summarizes who gets hurt and what the evidence says, and the following three sections analyze the most expensive failures in depth. Treat it as a diagnostic checklist: most firms exhibit at least three of the five simultaneously, and fixing the closed loop typically pays back fastest, because it converts data the firm already collects into retention actions with named owners rather than adding another instrument to a system nobody acts on.

Section 2

Challenge 1: The NPS evidence and the replication that failed

Reichheld's research, published as 'The One Number You Need to Grow' (HBR, 2003), linked survey responses to purchasing and referral behavior across industries and concluded that the likelihood-to-recommend question, scored 0-10, with promoters (9-10) minus detractors (0-6), correlated with relative growth rates in most industries studied. The claim spread because it was simple, comparable, and tied to a compelling growth story. The strongest counter-evidence arrived in 2007. Keiningham, Cooil, Andreassen, and Aksoy used longitudinal data from 21 firms and more than 15,500 interviews from the Norwegian Customer Satisfaction Barometer to replicate the original analyses (Journal of Marketing, 2007). They found NPS was not superior to other loyalty metrics, including the American Customer Satisfaction Index, in predicting revenue growth, a finding strong enough to win the Marketing Science Institute's H. Paul Root Award. The practical reading for founders is not that NPS is useless. The recommend question still captures relationship strength, and detractor comments are reliably diagnostic. The reading is that NPS is one indicator among several, its industry benchmarks are noisy, and decisions about pricing, staffing, or service redesign should never rest on the score alone. Advanced operators treat NPS movement as a prompt for investigation, not as a verdict.

Section 3

Challenge 2: What CSAT and CES actually predict

CSAT, satisfaction with a specific interaction or milestone, is the oldest of the three metrics and the most context-bound. Its weakness is well documented: satisfied customers defect anyway, which is precisely the gap Reichheld used to justify NPS. Its strength is diagnostic precision, because a low CSAT on a deliverable tells you exactly where the experience broke. Customer Effort Score comes from a different evidence base. Dixon, Freeman, and Toman's CEB study of more than 75,000 customers interacting with contact centers and self-service channels found that exceeding expectations barely moved loyalty, while reducing customer effort, repeat contacts, channel switching, re-explaining problems, strongly predicted it ('Stop Trying to Delight Your Customers,' HBR, 2010). Their conclusion: CES outperformed both CSAT and NPS in predicting loyalty in service contexts. The critique runs the other way too: CES was validated in transactional support settings, not in long-cycle professional relationships, so it cannot carry relationship-level judgment on its own. The synthesis the evidence supports is a division of labor. CES belongs at high-friction touchpoints, onboarding, handoffs, problem resolution, where effort is the loyalty killer. CSAT belongs at milestones where quality judgment matters. The recommend question belongs at the relationship level, asked sparingly. None of them substitutes for behavioral truth: renewal, expansion, and whether referrals actually arrive.

Section 4

Challenge 3: Small samples and the missing closed loop

Enterprise CX programs survive statistical noise because they survey thousands. A 40-client firm cannot. With 25 survey responses, one annoyed client moves NPS by eight points, and quarter-over-quarter score tracking becomes the study of randomness. The honest small-firm posture is to treat scores as census data about individuals, not statistics about populations: every detractor response is a named account requiring action, not a data point in a trend line. That reframe connects measurement to the second documented gap, the closed loop. Homburg and Fürst's dyadic study of complaint handling (Journal of Marketing, 2005) matched company-side complaint management practices with customer-side perceptions, and found that systematic, guideline-driven complaint handling (the mechanistic approach) had a stronger total effect on justice perceptions, satisfaction, and loyalty than relying on employee culture alone. In other words, the value of measurement is realized in the documented response process, not the dashboard. Small firms hold an underused advantage here: the principal can personally call every detractor within 48 hours, something no enterprise can do. Firms that operationalize that, score triggers a call, the call triggers a fix, the fix gets logged and verified, convert measurement into retention. Firms that simply chart scores convert measurement into overhead.

Section 5

Innovative solutions

The most useful innovations in small-firm CX measurement abandon the enterprise survey model rather than miniaturize it. First, interview-based relationship reviews: a 20-minute structured conversation each quarter with every key account, scored on a consistent rubric, generates richer signal than any emailed survey and doubles as an expansion touchpoint. Second, embedded micro-CES: a single one-click effort question fired immediately after onboarding, after each major handoff, and after any problem resolution, moments the effort research identifies as loyalty-critical (Dixon et al., 2010). Third, verbatim mining: with small samples, the open-text comment is worth more than the number, and modern text analysis makes theming fifty comments trivial; the score locates the account, the verbatim locates the cause. Fourth, behavioral triangulation: build a simple client health view that places survey signals next to ground-truth behavior, invoice payment speed, meeting attendance, response latency, scope expansion, referral activity. The Keiningham critique implies survey attitudes are imperfect proxies for behavior, so let observed behavior arbitrate. Finally, benchmark against your own baseline rather than published industry NPS tables, which mix incomparable survey methodologies. The relevant question for a service firm is never 'are we above the agency benchmark?' but 'which named accounts moved, in which direction, and why?'

Section 6

Solution framework

The small-firm measurement stack has four layers, each answering a different question. Layer one, transactional effort (CES): one-question pulse after onboarding, handoffs, and issue resolution, answering 'are we hard to work with right now?' Layer two, milestone satisfaction (CSAT): a short quality check after major deliverables, answering 'did this meet the standard?' Layer three, relationship strength: the recommend question plus a why, asked twice a year, ideally inside a structured review conversation, answering 'is this relationship an asset or a risk?' Layer four, behavioral ground truth: net revenue retention, logo churn, expansion rate, and realized referral count, reviewed monthly, the layer that audits the other three, per the replication evidence that attitudes do not guarantee outcomes (Keiningham et al., 2007). Two operating rules hold the stack together. Rule one: every negative signal at any layer triggers a named-owner response within 48 hours, logged and verified, reflecting the complaint-handling evidence that systematic process beats good intentions (Homburg and Fürst, 2005). Rule two: no compensation is tied to any score, removing the gaming incentive that corrupts the data. Total instrument burden on clients: under four questions per quarter. Total build cost: a form tool, a spreadsheet or lightweight CRM view, and discipline.

Section 7

Evidence-based action plan

Days 1-30: establish ground truth before adding surveys. Calculate trailing 12-month logo churn, net revenue retention, and referral count by account. Tag every current account red, yellow, or green based on observable behavior. This baseline is what every future score gets validated against. Days 31-60: deploy the minimal instrument set. Add the one-click CES question to onboarding completion and ticket resolution. Add a two-question milestone CSAT to your delivery workflow. Draft the 48-hour response protocol, who calls, what gets logged, how fixes get verified, before the first response arrives, because the complaint-handling evidence says the process is the product (Homburg and Fürst, 2005). Days 61-90: run the first relationship review cycle. Book 20-minute structured conversations with your top accounts; ask the recommend question live and follow the why. Theme the verbatims. Then hold a measurement review: which signals predicted the behavior you already observed, and which were noise? Kill anything that produced no decision. Quarterly thereafter, audit the stack against renewals and referrals, the Keiningham finding is a standing warning that scores can flatter while revenue leaks. The output that matters is not a trend line; it is a list of named accounts with actions attached. For adjacent evidence in this pillar, see [The Moment-of-Truth Map: Peak-End Experience Design for Service Firms](/blog/growth-moment-of-truth-journey-map-peak-end) and [Client Offboarding and Win-Back: The Economics of Graceful Exits](/blog/growth-client-offboarding-win-back-economics).

FAQ

Direct answers for operators.

Is NPS still worth using after the Keiningham critique?

Yes, with reduced expectations. The 2007 Journal of Marketing replication showed NPS is not a superior predictor of revenue growth, but the recommend question remains a useful relationship-strength signal and detractor verbatims are reliably diagnostic. Use it twice a year at the relationship level, never as your only metric, and always validated against renewals and realized referrals.

Which metric is best for a service firm with under 50 clients?

None alone. At that scale, scores are census data, not statistics, every response is a named account needing action. The evidence supports layering: CES at high-friction touchpoints (Dixon et al., 2010), CSAT at deliverable milestones, the recommend question in live quarterly reviews, and behavioral data, churn, expansion, referrals, as the arbiter of all three.

Should we tie team bonuses to NPS or CSAT?

The evidence argues against it. Incentivized teams beg for top scores and filter who gets surveyed, corrupting the data the firm needs for decisions, a problem widely documented in NPS practitioner critiques. Pay on behavioral outcomes like retention and expansion instead, and keep survey data clean for diagnosis rather than performance management.

How does customer effort score apply outside call centers?

The CEB research validated CES in transactional support, but the underlying finding, reducing effort drives loyalty more than exceeding expectations, transfers to service delivery. For agencies and consultancies, effort shows up as repeated explanations, chased updates, and clunky handoffs. A one-click effort question after onboarding and issue resolution catches these loyalty killers early.

Joshua Agonya Pi'Rwot

Written by

Joshua Agonya Pi'Rwot

Founder, Business Growth Accelerator · Country Director, AVODA Group Uganda · EMBA

Joshua helps service-business operators turn scattered marketing into a clear path from first attention to booked call. He is Founder of Business Growth Accelerator and Country Director of AVODA Group Uganda.