Section 1
The five challenges at a glance
The gap between agent marketing and agent evidence creates five operating risks for service firms. Overestimating autonomy leads to client-facing failures; the jagged capability frontier means performance collapses without warning on tasks that look easy; AI lifts average quality while narrowing the variance premium firms charge for; gains concentrate among junior staff, which disrupts the leverage model that prices senior time; and deployment economics kill many projects before value arrives. Each row below pairs the challenge with its root cause, the firms most exposed, and the controlling evidence. Read the table before approving any agent line item in your delivery budget, most failed deployments were predictable from one of these five rows.
Section 2
Challenge one: the evidence for real gains is strong, in assisted mode
Two large studies anchor the optimistic case, and both describe assistance, not autonomy. Brynjolfsson, Li, and Raymond studied 5,179 customer support agents at a Fortune 500 software firm as a generative AI assistant rolled out. Productivity, measured as issues resolved per hour, rose 14% on average, but the distribution is the headline: novice and low-skill workers improved 34%, while the most experienced agents saw minimal gains. The AI effectively encoded the practices of top performers and distributed them down the experience curve, improving customer sentiment and reducing attrition along the way (Brynjolfsson et al., 2023). The second anchor is the Harvard Business School and BCG field experiment: 758 consultants randomized across 18 realistic consulting tasks. With GPT-4, consultants completed 12.2% more tasks, worked 25.1% faster, and produced output human graders rated more than 40% higher in quality (Dell'Acqua et al., 2023). For a service firm, these are fulfillment economics studies in disguise. They imply AI-assisted delivery can compress junior ramp time, raise output floors, and expand effective capacity without hiring. What neither study shows is an agent completing client work unsupervised. Every documented gain occurred with a human in the loop deciding what to accept. That distinction, assistance versus autonomy, is the line the rest of the evidence draws in red.
Section 3
Challenge two: the evidence on autonomy is sobering
Carnegie Mellon researchers built TheAgentCompany, a simulated software firm where AI agents received realistic work: browsing, coding, project management, communicating with simulated colleagues. The results were blunt. The best performer at publication, Gemini 2.5 Pro, completed 30.3% of multi-step tasks; Claude 3.7 Sonnet managed 26.3%; GPT-4o just 8.6% (Xu et al., 2024). Failures were rarely exotic, agents got stuck on pop-up windows, misidentified colleagues, fabricated information, and in one case renamed a user to dodge a blocker. Partial credit scoring lifted the best models to roughly 39%, still failing most assignments (Xu et al., 2024). The market data rhymes: Gartner predicted in June 2025 that over 40% of agentic AI projects would be canceled by the end of 2027, citing escalating costs and unclear business value, and noted many vendors were 'agent washing', rebadging chatbots and RPA as agents (Gartner, 2025). The jagged frontier research adds the most operationally dangerous finding: consultants using AI on tasks outside its capability frontier performed 19 percentage points worse than peers without AI, because fluent, confident output masked wrong answers (Dell'Acqua et al., 2023). For client-facing delivery the implication is severe, the failure mode is not a visible crash but a polished deliverable that is subtly wrong, shipped under your brand and your professional liability.
Section 4
Challenge three: agents rewrite your delivery economics either way
Even deployed well, AI assistance disturbs the economics of professional service firms in three ways the research lets us anticipate. First, the leverage pyramid wobbles. If novices gain 34% productivity while experts gain little (Brynjolfsson et al., 2023), the traditional model, senior expertise priced through junior hours, compresses. Firms can serve more volume with smaller teams, but clients procuring services will eventually price that in and expect some of the surplus. Second, quality converges. In the BCG experiment, bottom-half performers improved 43% versus 17% for the top half, narrowing the distance between the best and the rest (Dell'Acqua et al., 2023). A premium boutique whose pitch is 'our people are simply better' is selling exactly the variance AI compresses; defensibility shifts toward proprietary data, method, and accountability rather than raw output quality. Third, verification becomes a real cost center. Because outside-frontier failures are fluent rather than obvious, supervised agent workflows require explicit review budgets, time, checklists, sampling, that partially offset headline productivity gains. Firms that count the generation savings but not the verification spend systematically overestimate margin improvement. None of this argues against adoption; the support-agent study showed retention and customer sentiment improving alongside productivity (Brynjolfsson et al., 2023). It argues for pricing, staffing, and positioning decisions made with the full cost structure in view.
Section 5
Innovative solutions
The research suggests a specific deployment doctrine. First, map your own jagged frontier empirically: take your ten most common delivery tasks, run each through your AI stack twenty times, and grade outputs against your senior standard. Tasks pass or fail on your data, not vendor benchmarks, the HBS team found capability boundaries are unintuitive and must be discovered per task (Dell'Acqua et al., 2023). Second, deploy inside-out: internal tasks first, research synthesis, first drafts, QA checklists, meeting summaries, where errors are caught before clients see them, then graduate proven tasks toward client-facing work. Third, build human-in-the-loop checkpoints as architecture, not policy: agents draft, named humans approve, and the approval step is logged. The support-agent study's gains all occurred under exactly this structure (Brynjolfsson et al., 2023). Fourth, budget verification explicitly, a working heuristic is to reserve 20-30% of the time saved for review until your error data justifies less. Fifth, exploit the novice multiplier deliberately: pair AI assistance with structured playbooks so junior staff deliver near-senior output on in-frontier tasks, and reprice or restructure engagements accordingly. Sixth, be honest in client communication about where AI participates in delivery; disclosure costs little now and protects trust when an error inevitably surfaces. Firms following this doctrine capture documented gains while avoiding the failure modes filling Gartner's cancellation forecast (Gartner, 2025).
Section 6
Solution framework
Operationalize adoption as a four-level maturity ladder with explicit gates. Level zero is manual delivery, your baseline, worth documenting so gains are measurable. Level one is assistive drafting: AI generates research, drafts, and analyses; humans transform them into deliverables. Evidence confidence is high here, this is the mode in which the 14% and 40% findings were produced (Brynjolfsson et al., 2023; Dell'Acqua et al., 2023). Level two is supervised agent workflows: multi-step automations, intake to brief, data to report draft, that run end-to-end but terminate at a mandatory human approval gate. Confidence is moderate; reliability data says expect failures and design for catching them. Level three is bounded autonomy: agents complete narrow, well-instrumented tasks without per-instance review, only for tasks with months of logged error rates below your tolerance, full audit trails, and low blast radius. Current benchmarks imply few delivery tasks qualify today (Xu et al., 2024). The gates between levels are evidence gates, not enthusiasm gates: a task moves up only with measured performance at the level below. The framework's quiet advantage is commercial, clients increasingly ask how firms use AI, and 'we promote tasks through evidence gates with logged human approval' is a trust-building answer that 'we have an AI-powered delivery platform' is not.
Section 7
Evidence-based action plan
Days 1-30: inventory and baseline. List every recurring task in your delivery workflow with hours consumed and error sensitivity. Measure current cycle times, without a baseline, you will never separate real gains from enthusiasm. Pick three internal-only tasks for level-one pilots. Days 31-60: run the pilots with discipline. Same tasks, AI-assisted versus standard, with quality graded blind by a senior reviewer where feasible, a small-scale version of the HBS design (Dell'Acqua et al., 2023). Log every failure and classify it: in-frontier (fixable with prompting and process) or outside-frontier (remove the task from scope). Days 61-90: institutionalize what survived. Write playbooks for passing tasks, build the approval-gate workflow, set your verification budget, and train junior staff on the assisted versions, they are where the research says returns concentrate (Brynjolfsson et al., 2023). Decline, for now, any client-facing autonomous deployment; the benchmark evidence does not support it (Xu et al., 2024). Quarterly thereafter: re-run frontier mapping as models improve, the boundary moves, and tasks failing today may pass in two quarters. Annually: revisit pricing and team structure in light of measured capacity gains. The winning posture is neither skepticism nor faith; it is a measurement system that lets you adopt faster than cautious rivals and safer than reckless ones. For adjacent evidence in this pillar, see [Competitive Intelligence in an Agentic Market: Monitoring When Agents Do the Shopping](/blog/growth-competitive-intelligence-agentic-market) and [Data Moats for Small Firms: Proprietary Data as Defensibility in the Agent Era](/blog/growth-data-moats-small-firms-agent-era).