Business Growth

The Agent-Washing Problem: How to Evaluate AI Agent Vendors Without Getting Burned

In June 2025, Gartner published one of the most quietly damning statistics in enterprise software: of the thousands of vendors claiming to sell agentic AI, the firm estimated only about 130 were real (Gartner, 2025). The same release predicted that more than 40% of agentic AI projects will be canceled by the end of 2027, undone by escalating costs, unclear business value and inadequate risk controls. Gartner named the underlying behavior 'agent washing', rebranding chatbots, RPA and AI assistants as agents without substantial agentic capability. For a 5-7 figure service business, where a failed platform bet can consume a year's improvement budget, the finding is a reason not to abstain, but to buy like an auditor. This article turns Gartner's evidence into a working due-diligence system.

Joshua Agonya Pi'Rwot

By Joshua Agonya Pi'Rwot

Founder, Business Growth Accelerator

Executive summary

Gartner estimates only about 130 of thousands of vendors selling 'AI agents' offer real agentic capability, and predicts over 40% of projects will be canceled by 2027. A due-diligence framework for buying safely.

Section 1

The five challenges at a glance

The agent-washing problem is really five distinct failure modes that share a marketing department. The table below separates them, because the defenses differ. The first is misrepresentation at the product level: scripted chatbots and rule-based RPA relabeled as autonomous agents, the literal washing Gartner identified (Gartner, 2025). The second is project mortality even with honest vendors: the predicted 40%-plus cancellation rate stems from cost escalation, unclear value and weak risk controls, problems that originate on the buyer's side as often as the seller's (Gartner, 2025). The third is hype-driven procurement: Gartner's January 2025 poll found organizations split between conservative investment and waiting, with only 19% making significant commitments, a market structurally primed for fear-of-missing-out purchases (Gartner, 2025). The fourth is demo opacity: agentic behavior is precisely the kind of capability a rehearsed demonstration can fake, because autonomy only reveals itself against unscripted inputs. The fifth is the backlash trap: burned buyers swing to blanket refusal, even as Gartner separately predicts 40% of enterprise applications will carry task-specific agents by the end of 2026 and a third of enterprise software will include agentic AI by 2028 (Gartner, 2025). Mature buyers hold both truths: most vendor claims are inflated, and the underlying capability trend is real. The discipline is distinguishing the two on evidence rather than narrative.

Section 2

What Gartner actually found, and what it did not say

Precision protects budgets, so state the findings exactly. Gartner's June 2025 release predicted that more than 40% of agentic AI projects will be canceled by the end of 2027, attributing the mortality to escalating costs, unclear business value and inadequate risk controls (Gartner, 2025). It defined agent washing as vendors rebranding existing products, AI assistants, robotic process automation, chatbots, without substantial agentic capabilities, and estimated only about 130 of the thousands of self-described agentic vendors were real (Gartner, 2025). Senior Director Analyst Anushree Verma characterized the current project landscape as 'early stage experiments or proof of concepts that are mostly driven by hype and are often misapplied' (Gartner, 2025). The supporting poll, taken in January 2025 across 3,412 webinar attendees, found 19% reporting significant agentic investment, 42% conservative investment, 8% none, and 31% waiting or unsure (Gartner, 2025). Now the part the doom headlines dropped: the same Gartner release predicted that by 2028, at least 15% of day-to-day work decisions will be made autonomously by agentic AI and 33% of enterprise software applications will include it (Gartner, 2025); two months later the firm predicted 40% of enterprise applications would feature task-specific agents by the end of 2026, up from under 5% in 2025 (Gartner, 2025). Gartner did not say agents are fake. It said most vendors selling them are exaggerating, a market-timing observation, not a technology verdict.

Section 3

The anatomy of a real agent versus a wrapped chatbot

Gartner's definitional line gives buyers a usable test: agentic AI delivers autonomous or semi-autonomous goal pursuit, planning, deciding and acting across steps, whereas washed products execute predefined flows with a language-model interface on top (Gartner, 2025). Translate that into five inspectable properties. Autonomy: can the system pursue a goal it was given, choosing its own intermediate steps, or does it follow an authored decision tree? Ask the vendor to show the plan the system generated, not the output it produced. Tool use: does it invoke external systems, query a database, call an API, update a record, and handle the failures those calls produce? A system that cannot act on your stack is an advisor, not an agent. Multi-step planning: give it a task requiring sequencing it has not seen; scripted products break visibly here. Statefulness: does it carry context across steps and sessions, or does each interaction start cold? Environmental feedback: when an action fails or returns surprises, does it adapt or halt? The procurement implication is a single rule, never evaluate from the vendor's demo environment. Insist on a time-boxed pilot against your data, your edge cases and your failure modes, with the vendor's engineers observable as they configure it. How much human authoring happens during setup tells you what the autonomy claims are worth. Real vendors accept these terms routinely; washed vendors negotiate them away.

Section 4

The cost of getting it wrong, in both directions

The visible failure is the wasted spend, and Gartner's cancellation drivers explain its anatomy: costs escalate because washed products require unbudgeted human scaffolding to simulate the promised autonomy; value stays unclear because the project was scoped to a demo narrative rather than a measured workflow; risk controls lag because nobody planned for a system that acts (Gartner, 2025). For a service business running 20-40% margins, a six-figure failed deployment is not a line item, it is the year's entire improvement capacity, plus the opportunity cost of the team's attention. But the inverse error compounds longer. The firm that responds to one burn by freezing all agentic investment exits the learning curve precisely as the capability becomes table stakes: 40% of enterprise applications carrying task-specific agents by end-2026 means your competitors' software will ship with working agents inside it regardless of anyone's procurement courage (Gartner, 2025). Verma's prescription points at the discriminating variable: value comes from enterprise productivity rather than individual task augmentation, deploying agents where decisions are needed, automation for routine workflows, and assistants for simple retrieval (Gartner, 2025). That hierarchy is a budgeting tool. Most washed products are assistants priced as agents; paying agent prices for retrieval is the quiet, recurring version of the burn. Matching the tool class to the job, and paying accordingly, converts Gartner's taxonomy into protection.

Section 5

Innovative solutions

Sophisticated buyers are converging on procurement patterns that price vendor risk explicitly. The first is the proof-of-value sprint: a paid two-to-four-week pilot on the buyer's real workflow with success metrics fixed in writing beforehand, conversion, cycle time, error rate, and a pre-agreed kill decision if they are missed. Paying for the pilot is deliberate; it buys the right to dictate scope and keeps the vendor's best engineers on it. The second is outcome-linked contracting: tying a meaningful fee share to the measured business result, which washed vendors systematically refuse because their unit economics depend on hidden human labor. The third is the autonomy audit: a standing requirement that vendors demonstrate, live and unscripted, the five properties, autonomy, tool use, planning, state, feedback handling, against scenarios the buyer authors. The fourth is reference diligence with teeth: speaking to two customers at comparable scale who have run the product for six months or more, asking specifically what human effort the system still requires; the gap between demo autonomy and operational autonomy lives in that answer. The fifth is the portfolio cadence: one pilot per quarter, maximum, each scoped to a single workflow with a named owner, a rhythm consistent with the conservative-majority posture Gartner's polling shows most organizations have sensibly adopted (Gartner, 2025). Together these flip the information asymmetry: the vendor now proves capability under your conditions, at their risk, before your capital commits.

Section 6

Solution framework

Compress the diligence into five questions, asked in order, each with a disqualifying answer. One: what does this system decide on its own? If the honest answer enumerates branches a human authored, you are buying automation, possibly worth buying, never worth agent pricing (Gartner, 2025). Two: show me a failure. Ask the vendor to demonstrate the system encountering an unplanned obstacle; real agents degrade visibly and recover or escalate; washed products either cannot run the scenario or fail silently. Three: what does setup actually involve? If configuration means consultants encoding your rules for weeks, the autonomy lives in the consultants. Four: will you contract on outcomes? Refusal is information; hedged acceptance with measurable triggers is the credible middle. Five: who else at my scale runs this in production, and what humans does it still need? Then score the deployment itself against Gartner's three documented killers, cost trajectory, value clarity, risk controls (Gartner, 2025), before signing: a budget that includes integration and supervision labor, a metric the CFO accepts as value, and a written answer to what the agent is permitted to do unsupervised. Firms that institutionalize the five questions report a useful side effect: vendor conversations get shorter. The washed sellers identify themselves early, usually at question two, which returns the scarcest resource a growth-stage founder spends on procurement, attention.

Section 7

Evidence-based action plan

Days 1-30: inventory and educate. List every tool in your stack marketed as agentic, including features inside software you already own, recall that task-specific agents are arriving embedded, predicted in 40% of enterprise apps by end-2026 (Gartner, 2025). Classify each against the assistant-automation-agent hierarchy (Gartner, 2025). Brief whoever buys software on the five-question screen; agent washing succeeds through unbriefed buyers. Days 31-60: structure the pipeline. Select the single workflow where an agent would create measurable value, decision-heavy, repetitive, instrumented. Draft your standard pilot agreement: paid sprint, written metrics, kill date, unscripted demonstration rights. Shortlist three vendors and run the five questions; expect the shortlist to shrink. Days 61-90: run one pilot. Execute against real work with a named internal owner, log the human effort the system actually consumes, and hold the kill discipline, Gartner's 40% cancellation prediction describes projects that lacked one (Gartner, 2025). Decide on the metrics, not the relationship. Quarterly thereafter: revisit the inventory, because embedded agents will keep appearing in tools you already pay for, and each deserves the same classification. The closing posture is Verma's, worth adopting verbatim as policy: pursue enterprise productivity, not task theater, agents where decisions are needed, automation for routine workflows, assistants for retrieval (Gartner, 2025). A firm that buys on that taxonomy is nearly impossible to agent-wash, and positioned to move fast when the roughly 130 real vendors become 500. For adjacent evidence in this pillar, see [Human Relationships in an Agentic Market: The Evidence for Pricing the Human Layer](/blog/growth-human-relationships-agentic-market) and [Agentic Commerce Protocols: What ACP, MCP, and the New Payment Rails Mean for Operators](/blog/growth-agentic-commerce-protocols-acp-mcp-payment-rails).

FAQ

Direct answers for operators.

What is agent washing?

Agent washing is vendors rebranding existing products, chatbots, robotic process automation, AI assistants, as 'AI agents' without substantial agentic capability. Gartner identified the practice in June 2025 and estimated that only about 130 of the thousands of vendors claiming agentic AI were real. The tell is autonomy: washed products execute predefined flows, while real agents plan and act across steps toward goals.

Why does Gartner predict so many agentic AI projects will fail?

Gartner predicts more than 40% of agentic AI projects will be canceled by the end of 2027, citing three drivers: escalating costs, unclear business value and inadequate risk controls. Most current projects are hype-driven experiments or proofs of concept, often misapplied. Notably, the same research expects adoption to keep rising, failure of projects and growth of the category are happening simultaneously.

How do I test whether a vendor's AI agent is real?

Test five properties in your own environment, never the vendor's demo: autonomy (does it choose steps or follow authored branches), tool use (can it act on your systems and handle failures), multi-step planning on unseen tasks, statefulness across sessions, and adaptation when actions fail. Insist on a paid, time-boxed pilot with written success metrics and a kill date. Washed vendors negotiate these terms away.

Should a small service business avoid agentic AI until the market matures?

Avoidance carries its own cost: Gartner predicts 40% of enterprise applications will feature task-specific agents by the end of 2026, so the capability arrives embedded in software you already use. The evidence supports cadence, not abstinence, one disciplined pilot per quarter, scoped to a measurable workflow, screened with the five-question framework, and killed without sentiment if metrics miss.

Joshua Agonya Pi'Rwot

Written by

Joshua Agonya Pi'Rwot

Founder, Business Growth Accelerator · Country Director, AVODA Group Uganda · EMBA

Joshua helps service-business operators turn scattered marketing into a clear path from first attention to booked call. He is Founder of Business Growth Accelerator and Country Director of AVODA Group Uganda.