Skip to content

The delivery gap

Enterprise AI agents have become easy to set up and remain hard to deliver. In Keenan Vision's Data 360 Strategy Sprint — a study of eighteen enterprise Agentforce engagements conducted in late 2025 and early 2026 — 87% of projects that cleared scoping, received budget, and entered build did not reach production. The proximate causes varied from project to project. The underlying cause did not: no artifact carried the business intent of the project from conception through production. Intent was rebuilt by hand at every boundary in the delivery chain, and every rebuild lost something.

What the field data shows

The Sprint comprised structured interviews with enterprise Salesforce customers, chosen as a deliberate cross-section of company scale and industry. The interviews produced 858 coded findings, which clustered into five binding constraints on Agentforce adoption.

Any single stalled project has a plausible local explanation. The Sprint's cases stalled on integration failures, governance escalations, trust reviews, data-lineage disputes, policy rewrites, and change-management delays — the ordinary vocabulary of enterprise project trouble. Read one case and you blame the integration. Read all eighteen together and a different picture emerges. These are not many different failures. They are one failure wearing many faces.

Consider the life of a typical project. The business scopes it in a slide deck. A technical team rebuilds that scope, by hand and from memory, as platform configuration. A trust review then asks questions the build artifacts cannot answer — who authorized this action, over which data, under which policy — so the team reconstructs the answers in interviews. A data-governance board asks whether the agent's use of the data matches the terms it was collected under, and the team reconstructs again. At every boundary, the project's intent is reconstructed — partially, and differently each time. Fidelity is lost at each reconstruction, and at enough reconstructions the project collapses under the weight of its own divergence.

Intent fidelity falling across scoping, build, trust review, and production, rebuilt by hand at every boundary — conceptual, not measured Intent fidelity falling across scoping, build, trust review, and production, rebuilt by hand at every boundary — conceptual, not measured

Read together, four of the Sprint's five binding constraints (the full list) turn out to be versions of exactly this: a stable statement of what the business wanted failing to survive the journey from scoping through build, trust review, and deployment. Where projects did reach production, teams had improvised a carrier for that statement — a wiki page, a shared spreadsheet, a sandbox export — pressed into a role no native artifact exists to fill.

A note on reading the number: the 87% figure is field evidence from a bounded, purposively selected sample, gathered in a strategy engagement rather than a peer-reviewed study. It should be read as a sharp signal from one location in the Salesforce ecosystem, not as a market-wide rate.

The wider pattern

The Sprint's signal does not stand alone. MIT's NANDA initiative reported in August 2025 that 95% of enterprise generative-AI pilots produce no measurable P&L impact — a profitability measure rather than a deployment measure, from a broad industry survey (the report). Gartner forecast in June 2025 that more than 40% of agentic AI projects will be cancelled by the end of 2027 — an analyst projection, citing escalating costs and unclear business value (the press release). And the Architecture as Strategy study — forty-eight interviews with Salesforce customers and partners, by Keenan and colleagues — found that pilots adding AI as an overlay to existing systems dominated the failures, while pilots embedding AI into existing interfaces, data paths, and governance dominated the successes (the study's findings). We read that pattern as the same diagnosis: failures were not scattered randomly across the project lifecycle — they clustered at the seams where intent had to be rebuilt.

These are different instruments measuring different things, and none of them proves the others. But they converge on one picture: agentic systems fail at the seam between what the business wants and what the system does, because nothing durable carries the first across to the second.

The paradox: setup collapsed, delivery didn't

What makes this stall genuinely strange is that it coexists with a real engineering triumph. Standing up an enterprise agent once took months; it now takes minutes. Guided setup, prebuilt templates, and progressive disclosure have made the product side of agent adoption nearly frictionless.

The engagement around that product is unchanged. It still runs on discovery workshops whose findings live in slide decks. It still depends on certified architects who are scarce, expensive, and booked. Its statements of work are still hand-built, one engagement at a time. And its quality still depends, decisively, on which individual happens to show up. Setup is industrial; delivery is artisanal.

That asymmetry explains the deployment ratio. When setup took months, setup was the bottleneck, and the industry fixed it. Now the bottleneck is everything around setup — capturing what the business actually wants, deciding what to build, and keeping that decision intact through governance and into production. Until delivery gets the same treatment setup received, the pattern in the data — projects funded, built, and abandoned short of production — should be expected to continue.


This is the gap Rosetta was built to close — start with why Salesforce should care, or return to the overview.