Skip to content

Field evidence

The delivery gap makes the argument. This page documents the instruments behind it — what each study measured, how, and where its limits lie — and then states the claims the research program has put on the record in a form that future evidence could prove wrong. That is the house discipline: every number travels with its instrument, and every thesis travels with the test that would falsify it.

The Data 360 Strategy Sprint

The program's primary field instrument is the Data 360 Strategy Sprint, conducted between late 2025 and early 2026: eighteen enterprise Salesforce engagements running Agentforce projects, chosen as a deliberate cross-section of company scale and industry and studied through structured interviews. The interviews were analyzed by structured thematic coding, producing 858 coded findings.

Those findings clustered into five binding constraints on Agentforce adoption:

  1. Scoping-to-build translation failure — projects scoped by the business in one artifact (a deck, a wiki page) and rebuilt by hand, partially and divergently, into platform configuration.
  2. Trust-review escalation — late-stage governance reviews asking questions the build artifacts could not answer (who authorized this, over which data, under which policy), forcing reconstruction by interview.
  3. Data-lineage disputes — escalations over whether an agent's use of data was authorized under the policies the data was collected under.
  4. Cross-cloud integration friction — projects spanning multiple Salesforce clouds rebuilding the same intent separately for each cloud's configuration, review, and handoff.
  5. Model behavior variability — the inherent non-determinism of the underlying models, which the papers treat as a distinct problem that better intent-carrying enables managing but does not solve.

The headline result is the 87% figure: of projects that cleared scoping, received named budget, and began configuration work, 87% never reached production. The definitions matter — the working paper operationalizes "entered build" and "reached production" explicitly rather than leaving them to impression.

The limits matter just as much. The Sprint was a strategy engagement, not a peer-reviewed study. Its eighteen engagements were selected purposively rather than sampled randomly; coding was single-rater with spot-checks rather than fully double-coded. These constraints are appropriate to the design, and they set the number's reach: 87% is precise about what happened inside this sample, and silent, on its own, about the wider market.

The Architecture as Strategy study

A second, independent instrument: the Architecture as Strategy working paper — forty-eight interviews with Salesforce customers and partners, conducted by Keenan and colleagues in work associated with UC Berkeley's Haas School of Business. Where the Sprint measured whether funded builds reached production, this study asked where in the lifecycle failures concentrate. Its central finding is architectural: pilots that added AI as an overlay — new interfaces, new data paths, new governance sitting parallel to existing systems — dominated the failures, while pilots that embedded AI into existing interfaces, data paths, and governance dominated the successes. Its limits: a qualitative interview study, strongest on pattern and mechanism rather than magnitude.

The bake-off — an independent head-to-head

In July 2026, Salesforce's own Data 360 product team designed and ran an independent evaluation of the distilled-skills lane: three identically tasked AI agents working against a live Data 360 org — one with the platform's standard MCP tool surface, one with the Rosetta-distilled skills available but not required, and one directed to use them. Only the agent using the Rosetta skills reached the correct architectural outcome. The evaluators' own takeaway: skills are a cheap and effective distribution mechanism for governed knowledge.

The counterweights, straight from the evaluation: the middle agent — skills available, never opened them — is its own finding (availability is not adoption); the winning agent, though architecturally right, looped enough that the evaluator had to nudge it back on course, while the losing agents finished unaided — just wrong; and wherever a knowledge assertion had not been verified against the live platform, that is exactly where it failed. The evaluation write-up's sharpest line deserves quoting: "curation without a test loop just produces prettier wrong answers." That lesson is now load-bearing in the pipeline itself — every distilled skill wears a draft/UNVERIFIED banner that only a passing live smoke test can lift.

The limits, as always: one scenario, one org, three runs, an evaluator in the loop — a directional engineering evaluation, not a controlled study. What makes it weigh more than its size is who ran it: the platform owner's own team, on their own tooling, with every incentive to find the standard path sufficient.

One platform-neutral lesson from the same run is worth the record: for agents, error messages are as load-bearing as APIs. A misleading error sends an agent down a wrong path as confidently as a wrong document sends a human — which is one more argument for governed, verified context on both sides of the tool call.

Corroborating signals

Two wider readings run the same way, and each sits on a different rung of the instrument ladder. The MIT NANDA figure — 95% of enterprise generative-AI pilots showing no measurable P&L impact, reported in August 2025 (the report) — comes from a broad industry survey, and it bounds profitability, a different quantity from the deployment ratio the Sprint measured. The Gartner figure — more than 40% of agentic AI projects cancelled by the end of 2027, forecast in June 2025 (the press release) — is an analyst projection: a professional judgment about the future, not yet a measurement of anything. Neither corroborates 87% arithmetically. What the three unlike instruments corroborate is a direction: enterprise agentic AI stalling short of realized value.

The program makes testable claims

The working papers do not stop at diagnosis; they commit to predictions that could fail.

The Spec Manifest paper states two. First: projects whose intent travels as a first-class, machine-readable artifact from scoping through build and trust review will show measurably fewer late-stage trust-review escalations than comparable projects carrying intent in wikis, threads, and requirements documents — because reviewers can query the artifact instead of re-interrogating the team. Second: agent data access authorized at the level of that artifact will produce measurably fewer data-lineage disputes than access authorized only by user identity — because the artifact carries the policy basis the data was collected under. Both are offered as targets for empirical work, not forecasts of magnitude.

The Governed Intent Compilation paper adds a falsifiable conjecture with a planned field-pilot design: that compiling intent into a governed working environment — through a durable intermediate artifact that carries its own provenance, with individual units of work compiled just-in-time against the live state of that environment — yields less rework-under-governance and more complete audit trails than translating a specification straight into implementation in one pass. The claim fails two ways. It fails if a straight one-pass translation matches the compiled approach on rework and audit completeness under the same governance obligations. And it fails if operator-agnosticism does not hold: the planned pilot executes the same compiled work environment once with a human operator and once with an AI agent, and if the two turn out to require materially different compilation or governance, the conjecture is refuted.

Stating claims this way is the point. A research program that can be wrong in public — that names its instruments, publishes its limits, and specifies its own refutation conditions — is one you can trust in private.


The papers these claims live in are described in the working papers. The argument built on this evidence is The delivery gap.