Business Growth

Sequencing AI Investments: Which Functions Pay Back First

AI budgets follow attention, and attention follows revenue functions: MIT found more than half of generative AI spending flows to sales and marketing tools. The same research found the biggest measured ROI somewhere else entirely, back-office automation that eliminates outsourcing costs and streamlines operations (MIT NANDA, 2025). That inversion is the most expensive mistake a growth-stage service firm can make in 2026, because capital is finite and sequencing determines compounding. This article ranks AI investment targets by the strength of the payback evidence, finance operations, customer service, sales support, knowledge work, and the expert-domain caution zone, drawing on benchmark data from Ardent Partners, deployment results from Klarna and Lumen, and the controlled studies that reveal where gains concentrate.

Joshua Agonya Pi'Rwot

By Joshua Agonya Pi'Rwot

Founder, Business Growth Accelerator

Executive summary

Over half of generative AI budgets flow to sales and marketing, yet MIT research finds the strongest returns in back-office automation. An evidence-ranked sequencing model for service businesses deciding where AI goes first.

Section 1

The five challenges at a glance

Sequencing fails for predictable reasons. Budgets chase visibility: sales and marketing AI demos well, so it absorbs over half of genAI spend while back-office automation quietly generates the measurable returns (MIT NANDA, 2025). Firms also sequence by enthusiasm, whichever partner or department head lobbies hardest goes first, reproducing the misalignment RAND identifies as the top root cause of AI project failure (RAND, 2024). And the skill-gradient evidence is routinely ignored: gains concentrate among less-experienced workers doing pattern-rich work (Brynjolfsson et al., 2025), while expert-domain automation shows weak or even negative measured returns (METR, 2025; Dahl et al., 2024). The result is a portfolio of pilots ordered by politics rather than payback, which is one reason only 39% of adopting organizations report EBIT impact (McKinsey, 2025). The table below names the five sequencing challenges and the evidence that resolves each one in favor of a payback-ranked order.

Section 2

Tier one: finance and back-office operations pay back first

The strongest payback evidence in the entire AI literature sits in the least glamorous function. MIT's deployment analysis found the biggest ROI in back-office automation, eliminating business process outsourcing, cutting external agency costs, and streamlining operations, even as budgets flowed elsewhere (MIT NANDA, 2025). The benchmark data explains why. Ardent Partners' ePayables research shows best-in-class accounts payable operations process an invoice for roughly $2.88 versus $12.88 for all others, a 78% unit-cost gap, in 3.1 days versus 17.4, with exception rates of 9% versus 22% (Ardent Partners, 2024). Automation maturity is the primary separator between the cohorts. Three properties make this tier sequence-first. The baseline is already quantified, so ROI is provable within a quarter. The work is rule-rich and repetitive, the profile where AI accuracy is highest and the human-review burden lowest. And the savings are annuities: every invoice, every month, forever. For a service firm, tier one extends beyond AP to invoicing, collections follow-up, expense categorization, scheduling, and intake processing. None of it will headline your marketing. All of it shows up in the only document that matters for the pillar this article serves: the P&L, typically within one to two quarters.

Section 3

Tier two: customer service and client communications

Customer service holds the single most complete public case study of AI payback. Klarna's assistant handled two-thirds of customer-service chats, 2.3 million conversations, within its first month, performing work equivalent to 700 full-time agents, cutting resolution time from eleven minutes to under two, reducing repeat inquiries by 25%, and matching human agents on customer satisfaction; the company estimated a $40 million profit improvement in 2024 (Klarna, 2024; OpenAI, 2024). The controlled research explains the mechanism and generalizes it: in Brynjolfsson, Li, and Raymond's study of over 5,000 support agents, AI assistance raised resolved issues per hour by 14% on average and 34% for novices, by diffusing the conversational patterns of top performers to everyone (Brynjolfsson et al., 2025). Customer service ranks second rather than first for one reason: brand risk. Errors in tier one are internal; errors here reach clients, and the liability evidence, Air Canada was held responsible for its chatbot's incorrect fare advice (BC CRT, 2024), argues for staged deployment. Start with AI-assisted human responses, the Brynjolfsson configuration where the human remains the sender. Graduate to direct automation only on inquiry types with high volume and low ambiguity, with escalation paths and human review of edge cases. Sequenced that way, the payback is fast and the downside is bounded.

Section 4

Tiers three and four: sales support, then knowledge work, and a caution zone

Tier three is sales support, not sales replacement. Lumen Technologies found seller research that took four hours per customer could be done in about fifteen minutes with Copilot, returning roughly four hours per seller per week, which the company values at $50 million annually across its sales force (Microsoft, 2024). The pattern generalizes to any service firm's business development: account research, meeting preparation, proposal first drafts, follow-up sequences. Note what the evidence supports, augmenting preparation time, not automating relationships. Tier four is junior-heavy knowledge work: research synthesis, report drafting, documentation. The experimental evidence is solid, 40% faster task completion with 18% higher quality, gains concentrated among less-experienced workers (Noy & Zhang, 2023), but payback is diffuse, arriving as capacity rather than a named line item, which is why it sequences after the tiers with hard baselines. Last comes the caution zone: expert-domain automation. METR's trial found experienced developers 19% slower with AI on deep-context work (METR, 2025), and Stanford's audit of general-purpose LLMs found hallucination rates of 58% to 88% on specific, verifiable legal questions (Dahl et al., 2024). High-stakes professional judgment, legal positions, tax advice, compliance calls, should be the destination of your AI sequence, approached with verification infrastructure, never the starting point.

Section 5

Innovative solutions

The sequencing leaders share three practices worth stealing. First, they price their workflows before pricing tools. A service firm that calculates its own cost per invoice, cost per resolved client inquiry, and hours per proposal can rank AI candidates by gap-to-benchmark, the Ardent Partners method applied internally (Ardent Partners, 2024), instead of by vendor narrative. Second, they sequence for compounding evidence: tier-one wins generate the baselines, the data hygiene, and the organizational trust that de-risk tier two, and so on down the stack. McKinsey's high performers exhibit exactly this pattern, well-defined KPIs, workflow redesign, and scaling discipline rather than scattered pilots (McKinsey, 2025). Third, they buy specialized and integrate deeply: MIT's finding that vendor partnerships succeed about 67% of the time versus roughly one-third for internal builds (MIT NANDA, 2025) matters doubly for small firms, where tier-one functions are served by mature, affordable tools embedded in existing accounting and CRM platforms. A fourth, contrarian practice is emerging among service firms specifically: treating the savings from tiers one and two as the funding mechanism for tiers three and four. Instead of a single annual AI budget allocated by committee, each tier's documented payback finances the next deployment, turning the sequence itself into a self-funding growth loop that never requires a leap of faith larger than one quarter's evidence.

Section 6

Solution framework

Adopt a payback-ranked sequence with explicit gates. Quarter one, tier one: deploy AI on finance and admin operations. Baseline unit costs first, target the benchmark gap (Ardent Partners, 2024), and hold a 90-day review against the baseline. Gate: documented savings or kill. Quarter two, tier two: move to customer service in assisted mode, where AI drafts and humans send, the configuration with the strongest controlled evidence (Brynjolfsson et al., 2025). Gate: response time and satisfaction metrics improve without error incidents. Quarter three, tier three: sales support, replicating the Lumen pattern of compressing research and preparation time (Microsoft, 2024), measured in recovered selling hours and pipeline activity. Quarter four, tier four: junior knowledge-work augmentation with senior review gates, targeting the experimentally documented speed and quality gains (Noy & Zhang, 2023). Throughout, the caution-zone rule stands: no autonomous AI in high-stakes professional judgment, where hallucination evidence remains damning (Dahl et al., 2024), until earlier tiers have built your verification muscle. Two design principles hold the framework together. Every tier inherits infrastructure from the last, baselines, clean data, review workflows, vendor management. And every gate is financial, not sentimental: the sequence only advances on measured payback, which is precisely the discipline separating the 39% who report EBIT impact from the 61% who do not (McKinsey, 2025).

Section 7

Evidence-based action plan

Week one: compute three numbers, your cost per invoice processed, cost per client inquiry resolved, and hours per proposal produced. These are the sequencing keys; firms that skip this step default to sequencing by enthusiasm. Weeks two to four: compare against benchmarks, the $2.88 versus $12.88 invoice spread and the 3.1-day versus 17.4-day cycle gap (Ardent Partners, 2024), and select the tier-one process with the widest gap. Shortlist specialized vendors; the build-versus-buy evidence is lopsided (MIT NANDA, 2025). Months two to three: deploy, measure, and document the win in currency, not percentages, a real number convinces a skeptical team faster than any deck. Month four onward: advance one tier per quarter through the gated sequence, customer service in assisted mode, then sales support, then knowledge work, letting each tier's savings fund the next and each tier's infrastructure de-risk the next. At every stage, resist the inversion that traps most firms: the pull of half-your-budget toward sales and marketing tools whose measured ROI trails the back office (MIT NANDA, 2025). Twelve months of this discipline produces what the 5% club actually looks like at small-firm scale: four functions running measurably cheaper or faster, a P&L that shows it, and an organization that has learned to demand evidence from every tool that knocks. For adjacent evidence in this pillar, see [The Human-in-the-Loop Dividend: Why Oversight Pays](/blog/growth-human-in-the-loop-dividend) and [The Shadow AI Problem: Converting Unsanctioned Use Into Governed Advantage](/blog/growth-shadow-ai-problem-governed-advantage).

FAQ

Direct answers for operators.

Which business function shows the fastest AI payback?

Finance and back-office operations, on current evidence. MIT found the biggest genAI ROI in back-office automation despite most budgets flowing to sales and marketing (MIT NANDA, 2025), and Ardent Partners benchmarks show best-in-class AP teams processing invoices at roughly $2.88 versus $12.88 for everyone else, a 78% unit-cost gap with quantified baselines that make ROI provable within a quarter (Ardent Partners, 2024).

Why not start with sales and marketing AI?

Because the spend-to-return evidence is inverted there. MIT found more than half of generative AI budgets devoted to sales and marketing tools while measured ROI concentrated in back-office automation (MIT NANDA, 2025). Sales support, research, preparation, drafting, does pay back, as Lumen's $50 million in recovered seller time shows (Microsoft, 2024), but it sequences best after cheaper, more provable wins.

When should AI touch customer-facing work?

Second in the sequence, and in assisted mode first. The controlled evidence for AI-assisted agents is strong, 14% average productivity gains, 34% for novices (Brynjolfsson et al., 2025), and Klarna's direct automation produced an estimated $40 million profit improvement (Klarna, 2024). But the Air Canada chatbot ruling shows firms bear full liability for AI errors (BC CRT, 2024), so graduate to automation only on high-volume, low-ambiguity inquiries.

What should come last in an AI investment sequence?

High-stakes expert judgment. General-purpose LLMs hallucinated on 58% to 88% of verifiable legal questions in Stanford's audit (Dahl et al., 2024), and METR found experienced practitioners 19% slower with AI on deep-context work (METR, 2025). Automate the surrounding workflow, research retrieval, drafting, formatting, but keep professional conclusions human until your firm has built verification infrastructure through earlier tiers.

Joshua Agonya Pi'Rwot

Written by

Joshua Agonya Pi'Rwot

Founder, Business Growth Accelerator · Country Director, AVODA Group Uganda · EMBA

Joshua helps service-business operators turn scattered marketing into a clear path from first attention to booked call. He is Founder of Business Growth Accelerator and Country Director of AVODA Group Uganda.