Business Growth

When Not to Automate: Evidence-Based Criteria for Keeping Humans in the Workflow

The automation literature has quietly assembled a counter-canon: rigorous studies where adding AI made things worse. A field experiment found that disclosing a chatbot's identity before a sales call cut purchases by more than 79.7% (Luo et al., 2019). A randomized trial found experienced developers were 19% slower with AI assistance, while believing they were faster (METR, 2025). Klarna, the poster child for AI-replaces-humans, reversed course and began rehiring support staff after quality fell (Bloomberg, 2025). None of this argues against automation; it argues for criteria. This article distills the evidence into a decision framework for growth-stage service businesses: which workflows to automate fully, which to augment, and which to deliberately keep human.

Joshua Agonya Pi'Rwot

By Joshua Agonya Pi'Rwot

Founder, Business Growth Accelerator

Executive summary

Disclosed chatbots cut purchases by 79.7% and AI tools made expert developers 19% slower. The research-backed criteria for deciding when automation destroys value, and where humans still win for service businesses.

Section 1

The five challenges at a glance

The decision to automate fails most often not because the AI cannot do the task, but because the task was misclassified. Five recurring misclassifications show up in the evidence. Trust-sensitive interactions: customers penalize known machines in persuasion contexts, independent of performance (Luo et al., 2019). Expert work: AI assistance reliably lifts novices, support agents improved 34% for new workers versus 14% on average (Brynjolfsson et al., 2023), but can slow seasoned experts on complex, context-heavy work (METR, 2025). Edge-case-heavy workflows: automation handles the median case and fails the tail, which is where service reputations live (Klarna case, 2025). Fragile-trust environments: people abandon algorithms after seeing a single error, even when the algorithm is superior on average (Dietvorst et al., 2015), so visible failures poison whole programs. And cost-fragile automation: assumed savings can invert, Gartner projects GenAI cost per service resolution will exceed many offshore human agents by 2030 (Gartner, 2026). The table summarizes the five, with root causes and evidence.

Section 2

Challenge one: the disclosure penalty, trust is the product

The strongest single result in the not-to-automate literature is Luo and colleagues' field experiment with over 6,200 customers receiving outbound sales calls. Undisclosed chatbots performed remarkably well, as effective as proficient human workers and four times more effective than inexperienced ones. But when the chatbot's identity was disclosed before the conversation, purchase rates fell by more than 79.7%, because customers perceived the disclosed bot as less knowledgeable and less empathetic, before it had said anything substantive (Luo et al., Marketing Science, 2019). The implication is not to hide bots; concealment is increasingly untenable legally and reputationally, and regulation is moving toward mandatory disclosure. The implication is that in persuasion and trust contexts, the human is part of the product. Buyers of professional services are purchasing accountability, judgment, and a relationship, someone who can be embarrassed if the work fails. The behavioral economics compounds this: Dietvorst and colleagues showed people abandon algorithms after seeing them err, even when the algorithm demonstrably outperforms humans (Dietvorst, Simmons, and Massey, 2015). Trust in machines is asymmetric, slower to build, faster to destroy, so a client who catches one AI-generated error in a deliverable may discount everything thereafter. Practical criteria: keep humans on the interactions where the counterparty is deciding whether to trust you, sales conversations, pricing discussions, bad-news delivery, renewal saves, and apply AI behind those humans, drafting and preparing, where its speed compounds without its trust penalty (Luo et al., 2019).

Section 3

Challenge two: the expertise inversion

AI's productivity effects are not uniform across skill levels, they often run opposite to intuition. The field evidence shows AI assistance functions as expertise transfer: in Brynjolfsson, Li, and Raymond's study of 5,179 support agents, AI lifted productivity 14% on average but 34% for novice workers, with minimal impact on the most experienced, because the system encoded what top performers already knew (Brynjolfsson et al., 2023). Noy and Zhang's randomized writing experiment found ChatGPT compressed the performance gap between stronger and weaker writers (Noy and Zhang, Science, 2023). Then METR's randomized trial completed the picture at the high end: 16 experienced open-source developers working on their own large, mature codebases were 19% slower when allowed to use AI tools, despite predicting 24% speedup beforehand and believing they had achieved 20% speedup afterward (METR, 2025). The mechanism is overhead: on complex tasks rich in tacit context, reviewing, correcting, and integrating AI suggestions costs more than drafting from expertise, and the perception gap means your seniors cannot self-report this honestly, because the tool feels fast even when it is slow. Decision criteria follow directly. Automate aggressively where your people are novices relative to the task: research synthesis, first drafts in unfamiliar domains, boilerplate. Be skeptical where your people are genuine experts working in deep context: senior strategy, complex client situations, your firm's signature methodology. And measure rather than poll, task-level timing data, not satisfaction surveys, because the METR result proves felt productivity and real productivity diverge (METR, 2025).

Section 4

Challenge three: edge cases, error costs, and the Klarna lesson

Klarna provides the clearest corporate natural experiment. The company replaced the workload of roughly 700 customer service staff with an AI assistant that handled about two-thirds of queries, and its CEO became the most quotable evangelist for AI-driven headcount reduction (CNBC, 2025). By May 2025, Klarna was rehiring humans for support. CEO Sebastian Siemiatkowski's diagnosis: 'As cost unfortunately seems to have been a too predominant evaluation factor when organizing this, what you end up having is lower quality' (Bloomberg, 2025). The failure pattern is generalizable: AI resolved the median ticket adequately, but quality collapsed on the tail, emotionally charged cases, multi-step problems, exceptions requiring discretion, and in services, reputation is priced on the tail, not the median. A customer remembers the one catastrophic interaction, not the forty adequate ones. The error-cost arithmetic should drive the decision: automation suits tasks where errors are cheap, detectable, and reversible (internal drafts, classification, scheduling) and fails where errors are expensive, hidden, or irreversible (compliance findings, financial advice, escalations, anything a client acts on without checking). Volume matters too: low-volume complex workflows never repay their automation overhead, while Gartner's cost-per-resolution forecast warns that even high-volume automation can lose its cost case as subsidies end and cases grow complex (Gartner, 2026). Gartner's broader prediction that over 40% of agentic AI projects will be canceled by 2027 for escalating costs, unclear value, or inadequate risk controls is, in effect, the Klarna pattern industrialized (Gartner, 2025).

Section 5

Innovative solutions

The firms getting this right are not choosing between humans and AI; they are engineering the seam. The leading pattern is AI-drafts-human-sends: AI produces the first version of everything, responses, proposals, analyses, and a named human owns verification and delivery on anything client-facing. This captures the documented drafting gains (40% faster, 18% higher quality in randomized testing, Noy and Zhang, 2023) while keeping the trust-bearing surface human, sidestepping the disclosure penalty entirely (Luo et al., 2019). Second, tiered escalation by design: automate tier one with explicit, generous triggers to human tier two, emotional language, repeat contact, high account value, low model confidence, encoding the Klarna lesson that the tail must route to people (Bloomberg, 2025). Klarna's own post-reversal model blends AI for routine queries with humans for empathy and discretion. Third, confidence-gated automation: expose model uncertainty and route low-confidence outputs to review queues rather than to clients. Fourth, error-budget governance: every automated workflow gets a defined acceptable error rate and a kill threshold, with blameless review protocols that prevent the single-error trust collapse Dietvorst documented from triggering wholesale abandonment of genuinely useful systems (Dietvorst et al., 2015). Fifth, periodic re-evaluation both directions: tasks move across the automate/augment/human boundary as models improve and as costs shift, Gartner's rising cost-per-resolution curve means some of today's automation will deserve de-automation by 2030 (Gartner, 2026). The seam, not the software, is the design object.

Section 6

Solution framework

Run every workflow through a five-question automation gate before deploying AI. One, is the interaction trust-bearing? If the counterparty is deciding whether to believe, buy, or forgive, keep a human in front; the disclosure evidence prices the penalty at most of your conversion (Luo et al., 2019). Two, who performs this task today, novices or experts? Novice-heavy tasks are prime automation targets (34% gains, Brynjolfsson et al., 2023); expert tasks in deep context risk the METR inversion and demand timing data before rollout (METR, 2025). Three, what does an error cost, and who catches it? Cheap, detectable, reversible errors permit full automation; expensive, hidden, or irreversible errors require human verification or human execution. Four, how fat is the tail? Estimate the share of cases that are exceptions; above roughly 20% edge cases, automate the median with engineered escalation, and below high volume thresholds, do not automate at all, the overhead never repays. Five, does the cost case survive subsidy decay? Model unit costs at realistic token consumption and assume vendor prices normalize upward, per Gartner's trajectory (Gartner, 2026). Scoring yields three lanes: Automate (routine, high-volume, error-tolerant, trust-neutral), Augment (AI drafts, human verifies and delivers, the default for client-facing knowledge work), and Human (trust-bearing, expert-context, fat-tailed, or error-catastrophic). Document the lane assignment per workflow, revisit quarterly, and treat lane changes as deliberate decisions with evidence, not as silent drift driven by tool enthusiasm or vendor pressure.

Section 7

Evidence-based action plan

Days 1-30: inventory your top fifteen workflows and run each through the five-question gate. Expect roughly a third in each lane for a typical service firm. For anything currently automated that sits in the Human or Augment lane, client-facing chat with no escalation triggers, auto-sent deliverables, install human verification immediately; you are carrying Klarna risk (Bloomberg, 2025). Days 31-60: engineer the seams. For Augment-lane workflows, implement AI-drafts-human-sends with named verification owners. For Automate-lane workflows, define escalation triggers, error budgets, and kill thresholds. Where senior staff use AI on expert work, run a two-week timing comparison, measured, not self-reported, because the METR perception gap guarantees self-reports mislead (METR, 2025). Days 61-90: measure the lanes. Track conversion and satisfaction on trust-bearing interactions, error rates and escalation volumes on automated tiers, and unit costs against the human alternative per Gartner's framing (Gartner, 2026). Publish the lane map internally so every hire knows which work is AI-first and which is deliberately human. Success criteria at day 90: zero unverified AI output reaching clients, every automated workflow carrying escalation triggers and an error budget, timing data on expert-AI combinations, and a documented automation map reviewed quarterly. The strategic payoff runs deeper than risk avoidance: as competitors automate indiscriminately, the firm that keeps humans precisely where humans win, trust, judgment, the tail, turns 'you will talk to a person who owns your outcome' into a priced, defensible differentiator. For adjacent evidence in this pillar, see [From AI Pilots to P&L: Why Most Initiatives Never Reach the Income Statement, and the Operator Playbook for the Ones That Do](/blog/growth-from-ai-pilots-to-pl-operator-playbook) and [The AI ROI Measurement Problem: Why EBIT Impact Stays Invisible and the Metrics Framework That Fixes It](/blog/growth-ai-roi-measurement-problem-ebit-metrics).

FAQ

Direct answers for operators.

What is the strongest evidence against automating customer-facing sales?

A Marketing Science field experiment with over 6,200 customers: chatbots performed as well as proficient human sales agents, but when the bot's identity was disclosed before the conversation, purchase rates fell by more than 79.7%, because customers pre-judged the machine as less knowledgeable and empathetic (Luo et al., 2019). Since hiding bots is not a viable strategy, the practical conclusion is to keep humans on trust-bearing conversations and use AI behind them for preparation and drafting.

Does AI really slow down experienced professionals?

It can. METR's 2025 randomized trial found experienced open-source developers working on their own mature codebases were 19% slower with AI tools, despite predicting a 24% speedup and believing afterward they had been 20% faster. The overhead of reviewing and correcting AI output exceeded its drafting value on complex, context-heavy work. The lesson: measure expert-AI combinations with timing data rather than self-reports, because perceived and actual productivity diverge.

What did Klarna's AI reversal actually prove?

That median-case automation fails tail-case businesses. Klarna's AI assistant handled roughly two-thirds of support queries, the routine ones, but quality collapsed on emotional, multi-step, and exceptional cases, and customer satisfaction fell. CEO Sebastian Siemiatkowski admitted cost had been 'a too predominant evaluation factor' producing 'lower quality' (Bloomberg, 2025). Klarna now runs a hybrid: AI for routine queries, rehired humans for empathy, discretion, and escalation, which is the design the evidence supported all along.

Which tasks should a service business automate first, and last?

First: high-volume, routine, error-tolerant, trust-neutral work, classification, scheduling, research synthesis, internal first drafts, especially tasks your junior people do, where evidence shows the largest gains (34% for novices, Brynjolfsson et al., 2023). Last or never: trust-bearing client conversations, expert judgment in deep context, fat-tailed exception handling, and anything where errors are expensive, hidden, or irreversible. The default for client-facing knowledge work is the middle lane: AI drafts, a named human verifies and delivers.

Joshua Agonya Pi'Rwot

Written by

Joshua Agonya Pi'Rwot

Founder, Business Growth Accelerator · Country Director, AVODA Group Uganda · EMBA

Joshua helps service-business operators turn scattered marketing into a clear path from first attention to booked call. He is Founder of Business Growth Accelerator and Country Director of AVODA Group Uganda.