Business Growth

The 5% Club: Inside AI Deployments That Actually Move the P&L

The arithmetic of enterprise AI is now public and uncomfortable. MIT's NANDA initiative found that roughly 95% of generative AI pilots deliver little or no measurable P&L impact, while about 5% achieve rapid value acceleration (MIT NANDA, 2025). McKinsey's global survey tells the same story from a different angle: 88% of organizations now use AI in at least one function, but only 39% report any enterprise-level EBIT impact, and a mere 6% qualify as high performers (McKinsey, 2025). Yet inside the 5% club, the numbers are striking, Klarna attributed a $40 million annual profit improvement to one assistant (Klarna, 2024). This article extracts the shared patterns from documented winners so growth-stage service firms can copy the playbook, not the hype.

Joshua Agonya Pi'Rwot

By Joshua Agonya Pi'Rwot

Founder, Business Growth Accelerator

Executive summary

MIT found 95% of generative AI pilots deliver no P&L impact, yet a small club books eight-figure returns. This evidence review dissects Klarna, Lumen, and best-in-class AP teams to isolate the patterns that separate the 5%.

Section 1

The five challenges at a glance

The gap between AI adoption and AI profit is not one problem but five compounding ones. Adoption is nearly universal, 88% of organizations report regular AI use in at least one function (McKinsey, 2025), yet most deployments never touch the income statement. The evidence base now lets us name the failure modes precisely. MIT's research, built on 150 executive interviews, a 350-employee survey, and analysis of 300 public deployments, attributes most failures not to model quality but to a 'learning gap': tools that never adapt to real workflows (MIT NANDA, 2025). RAND's interviews with 65 data scientists found misalignment between leaders and builders as the leading root cause, with AI projects failing at twice the rate of ordinary IT projects (RAND, 2024). The table below maps the five challenges, their root causes, and who pays the steepest price, typically founder-led service firms that lack slack budget to absorb a dead pilot.

Section 2

Challenge one: the adoption-impact gap

The defining statistic of this cycle is the spread between use and value. McKinsey's 2025 survey of nearly 2,000 respondents across 105 countries found 88% of organizations using AI in at least one function, up from 78% a year earlier, but only 39% reporting EBIT impact at the enterprise level, and most of those attributing less than 5% of EBIT to AI (McKinsey, 2025). Stanford's AI Index corroborates: most companies reporting financial benefits from AI estimate them at low levels (Stanford HAI, 2025). The high performers are rare and behave differently. McKinsey identifies roughly 6% of organizations as AI high performers, and they are distinguished less by spending than by ambition and design: they set growth and innovation objectives rather than pure efficiency targets, and most are redesigning workflows rather than bolting AI onto existing ones (McKinsey, 2025). For a 5-7 figure service business, the implication is blunt. Buying licenses is table stakes that competitors match within a quarter. P&L impact comes from choosing a workflow with a known cost or revenue baseline, instrumenting it before deployment, and holding the deployment to that baseline. If a use case cannot name the line item it will move, the evidence says it will join the 61% that produce nothing measurable.

Section 3

Challenge two: why 95% of pilots stall

MIT's GenAI Divide research is the most-cited failure study of the era, and its mechanism matters more than its headline. The roughly 95% failure rate is not driven by model quality or regulation, the researchers argue, but by a learning gap: generic tools like ChatGPT excel for individuals because they are flexible, but stall in organizational use because they do not learn from or adapt to specific workflows (MIT NANDA, 2025). RAND's complementary study, structured interviews with 65 experienced data scientists and engineers, found that the most common root cause is leaders misunderstanding or miscommunicating the problem AI should solve, and noted AI projects fail at twice the rate of non-AI IT projects (RAND, 2024). Two practical findings follow for smaller firms. First, the buy-versus-build asymmetry is large: purchasing from specialized vendors and building partnerships succeeded about 67% of the time in MIT's data, while internal builds succeeded roughly one-third as often (MIT NANDA, 2025). Second, adoption driven by line managers, the people who own the workflow, outperformed programs run from central innovation teams (MIT NANDA, 2025). A boutique consultancy or agency should read this as permission: you do not need a data science team. You need a process owner, a specialized tool, and a contract that lets the tool learn your workflow.

Section 4

Challenge three: the ROI is hiding in the back office

Where firms spend and where returns appear are nearly inverted. MIT found that more than half of generative AI budgets flow to sales and marketing tools, while the largest measured ROI sits in back-office automation, reducing outsourced business processes, cutting external agency spend, and streamlining operations (MIT NANDA, 2025). The finance function offers the cleanest benchmark data in the economy. Ardent Partners' ePayables research finds best-in-class accounts payable teams process an invoice for roughly $2.88 versus $12.88 for all others, in 3.1 days versus 17.4, with a 9% exception rate versus 22% (Ardent Partners, 2024). That is a 78% unit-cost gap on a process every service business runs, with automation and AI as the primary drivers separating the cohorts. The pattern recurs in service operations: Klarna's assistant handled two-thirds of customer-service chats in its first month, 2.3 million conversations, resolving issues in under two minutes versus eleven, with a 25% drop in repeat inquiries, work equivalent to 700 full-time agents and an estimated $40 million profit improvement (Klarna, 2024). The lesson for the 5% club aspirant: back-office and service-operations processes have existing unit-cost baselines, which makes ROI provable. Front-office creativity rarely does.

Section 5

Innovative solutions

The documented winners offer a small but consistent set of moves worth copying. Klarna deployed a single OpenAI-powered assistant against one workflow, customer service, with pre-existing metrics for handle time, resolution, and repeat contacts, then published results against those baselines: equivalent of 700 agents, available in 23 markets and more than 35 languages, with customer satisfaction on par with humans (Klarna, 2024; OpenAI, 2024). Lumen Technologies aimed Copilot at one bottleneck, seller research that took four hours per customer, and cut it to about fifteen minutes, returning roughly four hours per seller per week, which the company values at $50 million annually (Microsoft, 2024). Neither firm ran a portfolio of twenty pilots; both picked one measurable workflow and scaled what worked. McKinsey's high-performer data adds the organizational layer: winners redesign workflows, put senior leaders in visible ownership roles, and track well-defined KPIs for AI initiatives (McKinsey, 2025). MIT adds the procurement layer: specialized vendors over internal builds, and tools selected for their ability to integrate deeply and improve with use (MIT NANDA, 2025). For a service firm, the synthesis is a 'one workflow, one owner, one number' doctrine, a deliberately boring strategy that the evidence says beats innovation theater roughly twenty to one.

Section 6

Solution framework

Translate the case evidence into a four-gate framework any growth-stage firm can run. Gate one: baseline. Choose a workflow where you already know the unit economics, cost per invoice, hours per proposal, cost per resolved ticket. If no baseline exists, build one first; the 5% club is defined by before-and-after numbers, and Ardent Partners' cost-per-invoice benchmarks show how stark those numbers can be (Ardent Partners, 2024). Gate two: ownership. Assign the line manager who runs the workflow, not an innovation committee, MIT found line-manager-driven adoption materially outperforms central programs (MIT NANDA, 2025). Gate three: procurement. Default to specialized vendors with deep integration and learning capability; reserve internal builds for genuine differentiators, accepting that they succeed about one-third as often (MIT NANDA, 2025). Gate four: the kill rule. Set a 90-day checkpoint against the baseline metric; pilots that cannot show movement get stopped, redirected, or renegotiated. This addresses RAND's core finding that failures stem from misalignment and from organizations chasing technology rather than problems (RAND, 2024). The framework's power is its symmetry: the same number that justified the pilot decides its fate. That single discipline, refusing to let success criteria drift after launch, is the cheapest membership fee the 5% club charges.

Section 7

Evidence-based action plan

Week one: inventory your workflows and rank them by baseline quality, not excitement. Finance operations, scheduling, intake, and customer service usually surface first because unit costs are already visible, and the benchmark gaps are enormous, with best-in-class invoice processing at roughly $2.88 versus $12.88 (Ardent Partners, 2024). Weeks two to four: select one workflow, document its current cost and cycle time, and shortlist two or three specialized vendors, weighting integration depth and the tool's ability to learn your process (MIT NANDA, 2025). Month two: deploy with the line manager as owner, define one financial KPI and one quality KPI, and instrument both. Month three: review against baseline; scale, fix, or kill. Throughout, resist the budget gravity that pulls spend toward sales and marketing tools, MIT's data shows over half of genAI budgets land there while measured ROI concentrates in back-office automation (MIT NANDA, 2025). Expect honest timelines: only 39% of adopters report EBIT impact, and high performers earned theirs through workflow redesign, not tool acquisition (McKinsey, 2025). For a service firm doing this with discipline, the realistic prize is not a press release, it is a two-to-four-point margin improvement on one process, repeated across processes until the P&L notices. For adjacent evidence in this pillar, see [AI and the Small-Firm Productivity Premium: Who Gains Most](/blog/growth-ai-small-firm-productivity-premium) and [Data Readiness Is the Growth Bottleneck Your AI Plan Ignores](/blog/growth-data-readiness-ai-bottleneck).

FAQ

Direct answers for operators.

What does the research actually say about AI pilot failure rates?

MIT's NANDA initiative found roughly 95% of generative AI pilots deliver little or no measurable P&L impact, based on 150 interviews, a 350-employee survey, and 300 public deployments (MIT NANDA, 2025). RAND separately found more than 80% of AI projects fail, twice the rate of non-AI IT projects, with leader-builder misalignment as the top root cause (RAND, 2024).

Which companies have documented financial results from AI?

Klarna reported its AI assistant performed work equivalent to 700 full-time agents and estimated a $40 million profit improvement in 2024, with resolution times falling from eleven minutes to under two (Klarna, 2024). Lumen Technologies valued Copilot-driven seller time savings, about four hours per week per seller, at $50 million annually (Microsoft, 2024).

Should a small service firm build or buy AI tools?

Buy, almost always. MIT found purchasing from specialized vendors and building partnerships succeeded about 67% of the time, while internal builds succeeded roughly one-third as often (MIT NANDA, 2025). For 5-7 figure service firms without ML engineering depth, the evidence favors specialized tools that integrate deeply with existing workflows and improve with use, owned by the line manager who runs the process.

Where should we look first for measurable AI ROI?

Back-office and service operations. MIT found the largest ROI in back-office automation even though over half of budgets go to sales and marketing (MIT NANDA, 2025). Finance operations offer the clearest benchmarks: best-in-class AP teams process invoices for roughly $2.88 versus $12.88 for everyone else, a 78% unit-cost gap driven largely by automation (Ardent Partners, 2024).

Joshua Agonya Pi'Rwot

Written by

Joshua Agonya Pi'Rwot

Founder, Business Growth Accelerator · Country Director, AVODA Group Uganda · EMBA

Joshua helps service-business operators turn scattered marketing into a clear path from first attention to booked call. He is Founder of Business Growth Accelerator and Country Director of AVODA Group Uganda.