Section 1
The five challenges at a glance
The agent-washing problem is really five distinct failure modes that share a marketing department. The table below separates them, because the defenses differ. The first is misrepresentation at the product level: scripted chatbots and rule-based RPA relabeled as autonomous agents, the literal washing Gartner identified (Gartner, 2025). The second is project mortality even with honest vendors: the predicted 40%-plus cancellation rate stems from cost escalation, unclear value and weak risk controls, problems that originate on the buyer's side as often as the seller's (Gartner, 2025). The third is hype-driven procurement: Gartner's January 2025 poll found organizations split between conservative investment and waiting, with only 19% making significant commitments, a market structurally primed for fear-of-missing-out purchases (Gartner, 2025). The fourth is demo opacity: agentic behavior is precisely the kind of capability a rehearsed demonstration can fake, because autonomy only reveals itself against unscripted inputs. The fifth is the backlash trap: burned buyers swing to blanket refusal, even as Gartner separately predicts 40% of enterprise applications will carry task-specific agents by the end of 2026 and a third of enterprise software will include agentic AI by 2028 (Gartner, 2025). Mature buyers hold both truths: most vendor claims are inflated, and the underlying capability trend is real. The discipline is distinguishing the two on evidence rather than narrative.
Section 2
What Gartner actually found, and what it did not say
Precision protects budgets, so state the findings exactly. Gartner's June 2025 release predicted that more than 40% of agentic AI projects will be canceled by the end of 2027, attributing the mortality to escalating costs, unclear business value and inadequate risk controls (Gartner, 2025). It defined agent washing as vendors rebranding existing products, AI assistants, robotic process automation, chatbots, without substantial agentic capabilities, and estimated only about 130 of the thousands of self-described agentic vendors were real (Gartner, 2025). Senior Director Analyst Anushree Verma characterized the current project landscape as 'early stage experiments or proof of concepts that are mostly driven by hype and are often misapplied' (Gartner, 2025). The supporting poll, taken in January 2025 across 3,412 webinar attendees, found 19% reporting significant agentic investment, 42% conservative investment, 8% none, and 31% waiting or unsure (Gartner, 2025). Now the part the doom headlines dropped: the same Gartner release predicted that by 2028, at least 15% of day-to-day work decisions will be made autonomously by agentic AI and 33% of enterprise software applications will include it (Gartner, 2025); two months later the firm predicted 40% of enterprise applications would feature task-specific agents by the end of 2026, up from under 5% in 2025 (Gartner, 2025). Gartner did not say agents are fake. It said most vendors selling them are exaggerating, a market-timing observation, not a technology verdict.
Section 3
The anatomy of a real agent versus a wrapped chatbot
Gartner's definitional line gives buyers a usable test: agentic AI delivers autonomous or semi-autonomous goal pursuit, planning, deciding and acting across steps, whereas washed products execute predefined flows with a language-model interface on top (Gartner, 2025). Translate that into five inspectable properties. Autonomy: can the system pursue a goal it was given, choosing its own intermediate steps, or does it follow an authored decision tree? Ask the vendor to show the plan the system generated, not the output it produced. Tool use: does it invoke external systems, query a database, call an API, update a record, and handle the failures those calls produce? A system that cannot act on your stack is an advisor, not an agent. Multi-step planning: give it a task requiring sequencing it has not seen; scripted products break visibly here. Statefulness: does it carry context across steps and sessions, or does each interaction start cold? Environmental feedback: when an action fails or returns surprises, does it adapt or halt? The procurement implication is a single rule, never evaluate from the vendor's demo environment. Insist on a time-boxed pilot against your data, your edge cases and your failure modes, with the vendor's engineers observable as they configure it. How much human authoring happens during setup tells you what the autonomy claims are worth. Real vendors accept these terms routinely; washed vendors negotiate them away.
Section 4
The cost of getting it wrong, in both directions
The visible failure is the wasted spend, and Gartner's cancellation drivers explain its anatomy: costs escalate because washed products require unbudgeted human scaffolding to simulate the promised autonomy; value stays unclear because the project was scoped to a demo narrative rather than a measured workflow; risk controls lag because nobody planned for a system that acts (Gartner, 2025). For a service business running 20-40% margins, a six-figure failed deployment is not a line item, it is the year's entire improvement capacity, plus the opportunity cost of the team's attention. But the inverse error compounds longer. The firm that responds to one burn by freezing all agentic investment exits the learning curve precisely as the capability becomes table stakes: 40% of enterprise applications carrying task-specific agents by end-2026 means your competitors' software will ship with working agents inside it regardless of anyone's procurement courage (Gartner, 2025). Verma's prescription points at the discriminating variable: value comes from enterprise productivity rather than individual task augmentation, deploying agents where decisions are needed, automation for routine workflows, and assistants for simple retrieval (Gartner, 2025). That hierarchy is a budgeting tool. Most washed products are assistants priced as agents; paying agent prices for retrieval is the quiet, recurring version of the burn. Matching the tool class to the job, and paying accordingly, converts Gartner's taxonomy into protection.
Section 5
Innovative solutions
Sophisticated buyers are converging on procurement patterns that price vendor risk explicitly. The first is the proof-of-value sprint: a paid two-to-four-week pilot on the buyer's real workflow with success metrics fixed in writing beforehand, conversion, cycle time, error rate, and a pre-agreed kill decision if they are missed. Paying for the pilot is deliberate; it buys the right to dictate scope and keeps the vendor's best engineers on it. The second is outcome-linked contracting: tying a meaningful fee share to the measured business result, which washed vendors systematically refuse because their unit economics depend on hidden human labor. The third is the autonomy audit: a standing requirement that vendors demonstrate, live and unscripted, the five properties, autonomy, tool use, planning, state, feedback handling, against scenarios the buyer authors. The fourth is reference diligence with teeth: speaking to two customers at comparable scale who have run the product for six months or more, asking specifically what human effort the system still requires; the gap between demo autonomy and operational autonomy lives in that answer. The fifth is the portfolio cadence: one pilot per quarter, maximum, each scoped to a single workflow with a named owner, a rhythm consistent with the conservative-majority posture Gartner's polling shows most organizations have sensibly adopted (Gartner, 2025). Together these flip the information asymmetry: the vendor now proves capability under your conditions, at their risk, before your capital commits.
Section 6
Solution framework
Compress the diligence into five questions, asked in order, each with a disqualifying answer. One: what does this system decide on its own? If the honest answer enumerates branches a human authored, you are buying automation, possibly worth buying, never worth agent pricing (Gartner, 2025). Two: show me a failure. Ask the vendor to demonstrate the system encountering an unplanned obstacle; real agents degrade visibly and recover or escalate; washed products either cannot run the scenario or fail silently. Three: what does setup actually involve? If configuration means consultants encoding your rules for weeks, the autonomy lives in the consultants. Four: will you contract on outcomes? Refusal is information; hedged acceptance with measurable triggers is the credible middle. Five: who else at my scale runs this in production, and what humans does it still need? Then score the deployment itself against Gartner's three documented killers, cost trajectory, value clarity, risk controls (Gartner, 2025), before signing: a budget that includes integration and supervision labor, a metric the CFO accepts as value, and a written answer to what the agent is permitted to do unsupervised. Firms that institutionalize the five questions report a useful side effect: vendor conversations get shorter. The washed sellers identify themselves early, usually at question two, which returns the scarcest resource a growth-stage founder spends on procurement, attention.
Section 7
Evidence-based action plan
Days 1-30: inventory and educate. List every tool in your stack marketed as agentic, including features inside software you already own, recall that task-specific agents are arriving embedded, predicted in 40% of enterprise apps by end-2026 (Gartner, 2025). Classify each against the assistant-automation-agent hierarchy (Gartner, 2025). Brief whoever buys software on the five-question screen; agent washing succeeds through unbriefed buyers. Days 31-60: structure the pipeline. Select the single workflow where an agent would create measurable value, decision-heavy, repetitive, instrumented. Draft your standard pilot agreement: paid sprint, written metrics, kill date, unscripted demonstration rights. Shortlist three vendors and run the five questions; expect the shortlist to shrink. Days 61-90: run one pilot. Execute against real work with a named internal owner, log the human effort the system actually consumes, and hold the kill discipline, Gartner's 40% cancellation prediction describes projects that lacked one (Gartner, 2025). Decide on the metrics, not the relationship. Quarterly thereafter: revisit the inventory, because embedded agents will keep appearing in tools you already pay for, and each deserves the same classification. The closing posture is Verma's, worth adopting verbatim as policy: pursue enterprise productivity, not task theater, agents where decisions are needed, automation for routine workflows, assistants for retrieval (Gartner, 2025). A firm that buys on that taxonomy is nearly impossible to agent-wash, and positioned to move fast when the roughly 130 real vendors become 500. For adjacent evidence in this pillar, see [Human Relationships in an Agentic Market: The Evidence for Pricing the Human Layer](/blog/growth-human-relationships-agentic-market) and [Agentic Commerce Protocols: What ACP, MCP, and the New Payment Rails Mean for Operators](/blog/growth-agentic-commerce-protocols-acp-mcp-payment-rails).