← All reports

October 4, 2026

Agents in Production: What the 5% Actually Do Differently

The 95% pilot failure stat, what it really measured, and the three habits that separate teams whose agents reach production from those whose don't.

Contents

What changed

The "95% of AI pilots fail" line comes from MIT Project NANDA's report The GenAI Divide: State of AI in Business 2025, dated July 2025. It has been quoted for a year, usually without the sentence underneath it. The report's own funnel is more useful than the headline: enterprise-grade GenAI tools, custom-built or vendor-sold, were evaluated by 60% of organizations, piloted by 20%, and reached production at 5%. The 95% figure is the inverse of that last step, measured against a specific bar: a deployment that gets past the pilot with measurable KPIs, assessed six months after the pilot ends. MIT attributes roughly $30–40 billion of enterprise spend to that period with no measurable P&L movement.

The split that matters for anyone deciding how to build sits in the interview sample. Externally partnered tools, customized and able to learn from feedback, reached deployment about 67% of the time. Internally built tools reached deployment about 33% of the time, and employee usage of the external tools ran roughly double. MIT prints its own caveat: the split is self-reported and rests on 52 interviewed organizations, not the survey of 153 senior leaders or the review of 300-plus public initiatives. Correlation, not proof, and possibly a proxy for organizational capability rather than for buying versus building.

Not everyone reads the study the same way. GoML, an AI delivery firm, argues the report is a Trojan horse for MIT's own agent-internet initiative, and objects to a sub-six-month window for P&L impact and to failure data sourced from project sponsors rather than implementation teams. GoML also describes the sample as roughly 400 large organizations, which does not match the 52-organization figure attached to the 67/33 split. We flag the conflict rather than pick a side.

A year of newer data points the same direction on returns while moving the opposite way on adoption. McKinsey's State of AI in 2026: On the road to ROI, published 25 August 2026 with 1,719 respondents across 97 countries, reports 40% of organizations above $1 billion in revenue now scaling AI agents in at least one business function, up from 27% a year earlier. Companies below that line sat flat at 22%. Enterprise-wide AI scaling rose from 38% to 44%. The share reporting any EBIT impact from AI did not move: 37%, two years running. About 6% clear McKinsey's high-performer bar of at least 5% of EBIT attributable to AI, also unchanged.

Gartner, in a June 2025 press release, predicted over 40% of agentic AI projects would be cancelled by the end of 2027, citing cost, unclear value, and immature use cases. The same release said 19% of surveyed organizations had made significant agentic AI investments, 42% conservative ones, 8% none, and 31% were waiting. Gartner also claimed that of the thousands of products marketed as AI agents, only about 130 had genuine agentic capability, and coined "agent washing" for the rest.

Why this matters to a brand spending $50K+/month on paid social

At that spend level the agent pitch lands on a few jobs: support and returns triage, feed and catalogue hygiene, creative briefs and variant naming, reporting and anomaly alerts, and increasingly buying-side actions like budget shifts. Each touches either money or the customer directly, which changes what "good enough" means: a demo that works once tells you almost nothing about a tool that has to work on every ticket.

The clearest public evidence on that distinction is Sierra's tau-bench (written τ-bench), released in 2024 and now extended into further domains. Version 1 contains 165 tasks, 115 retail and 50 airline. Each is a multi-turn conversation with an LLM-simulated customer, domain APIs that actually read and write a database, and a written policy the agent must not break. Scoring is on the final database state, not on how good the conversation sounded. The metric that matters is pass^k: the probability the agent succeeds on all k independent attempts at the same task, not the average.

Independent write-ups report that strong function-calling models land well under 70% pass^1 on retail and roughly 35–46% on airline, and that pass^8 falls far below those figures for every model tested. We could not open a current per-model leaderboard, so treat the percentages as directional. The shape transfers: an agent that works 70% of the time is not 70% reliable across a repeated workflow. On 4,000 SKUs or 900 tickets a week, that is an operational cost, not a rounding error.

What the 5% appear to do differently

Three sources, measuring different things, converge on one short list. None is about model choice.

They rebuild one workflow instead of bolting agents onto all of them. McKinsey's high performers are far more likely to have fundamentally redesigned workflows around AI rather than inserted AI into existing ones — one analysis puts it at roughly three-quarters of high performers against about a quarter of everyone else — and about twice as likely to have a defined process for measuring impact. Civic's read on the same survey is that the large-versus-small gap is implementation capacity, not software cost: about one in five respondents say AI operating costs including tokens constrain usage, but for agents specifically only about one in ten cite cost. The constraint is people who can configure and maintain the thing.

They scope narrowly and define the human handoff before launch. Salesforce surveyed 2,025 agentic AI decision-makers for its State of Agentic AI in the Enterprise report, fielded across 20 countries in May 2026. Its three most predictive success factors were clean, accessible data at the moment the agent acts (36%), a narrowly scoped use case (36%), and human escalation paths defined before launch (35%) — all ahead of model quality, platform choice, and orchestration layer at around 30%. Salesforce reports 30% of respondents running agents in production, meaningful ROI in about eight months on average, 53% employee adoption, and a 29% lift in customer satisfaction. These are Salesforce's own figures, including a claim that retailers running AI agents grew online sales at four times the rate of those that did not — a comparison with no published control for retailer size or category.

They do not wait for perfect data. Only 31% of Salesforce's deployers had fully unified their data before launching agents. Those who did reached meaningful ROI in 7.3 months against 8.8 months for those who deployed first and fixed the data later. That is a real gap and a small one — evidence for sequencing, not for a two-year data project. Being early to deploy also did not predict being early to return: professional and business services were among the slowest to deploy and fastest to ROI at 6.5 months, while high tech was among the biggest deployers and one of the slowest at 10.1 months.

What to do this month

Three actions, each finishable inside 30 days by a team with an analyst and an engineer.

1. Pick one workflow and write its pass^k test before you automate it

Choose the highest-volume repetitive task in your ops queue — returns triage, feed error resolution, creative tagging, weekly anomaly reports. Write down 20 to 30 real historical cases with the correct outcome for each, including the messy ones. Run any candidate agent against that set eight times and record how often it gets all eight right, not the average. That is your pass^8, and it predicts how many exceptions land back on a human. If it is low, the fix is usually a narrower scope, not a different model.

2. Define the escalation path and the blast radius in writing

Before anything runs unattended, specify the actions the agent may take, the ones that require a human, and the hard limits: maximum budget change per run, maximum refund value, maximum records modified per hour, and what happens on tool failure. Salesforce's data puts pre-defined escalation paths among the top three predictors of success, and Gartner's cancellation forecast is driven partly by projects that discovered cost and complexity only at scale. Write the limits down.

3. Instrument the workflow before you change it

Baseline current cost and cycle time: touches per ticket, minutes per feed fix, hours per reporting cycle, error rate. EBIT impact flat at 37% for two years while adoption climbed is at least partly a measurement failure — without a baseline there is no way to tell a real saving from a reallocated one. If you cannot measure the workflow today, you cannot prove the agent did anything, and that is how a working pilot dies at renewal.

What we would watch next

  • Whether the EBIT line moves in the 2027 survey. Two flat years at 37% while scaling rose is the most informative number in this space. A third flat year would say the constraint is structural.
  • Gartner's 2027 cancellation forecast against actual cancellations. The 40% prediction was made in mid-2025; by late 2027 it becomes checkable rather than quotable.
  • pass^k reporting by vendors. Almost no commercial agent vendor publishes pass^k on its own domain. The first that do — stating agent model, user-simulator model, task release, and trial count — will be worth taking seriously.
  • Buying-side agents in ad platforms. The reliability argument sharpens when an agent can move budget. Watch for published guardrail defaults and per-run spend caps, not capability announcements.
  • Whether the 67/33 build-versus-buy split reproduces on a larger sample. It rests on 52 self-reporting organizations.

What we could not verify

  • Primary copies of the MIT NANDA and McKinsey documents. Both mckinsey.com and gartner.com blocked our fetches (Akamai and Cloudflare challenges). The McKinsey figures come from two independent write-ups citing the same 25 August 2026 publication that agree with each other; the MIT figures come from a secondary analysis quoting the report's methodology directly.
  • The MIT sample size. Our main analysis cites 52 interviewed organizations behind the 67/33 split, 153 surveyed senior leaders, and 300-plus public initiatives; GoML's critique describes roughly 400 large organizations. We could not reconcile these publicly.
  • Current tau-bench per-model scores. The pass^1 ranges quoted (under 70% retail, 35–46% airline) come from benchmark write-ups rather than a live leaderboard. Scores are only comparable when domain, task release, agent and user-simulator models, prompting strategy, and trial count are all specified.
  • Salesforce's retail claim. "Retailers running AI agents grew online sales at four times the rate" is Salesforce platform data with no published methodology or control. Attribute it, do not plan against it.
  • Gartner's "only 130 genuinely agentic tools" count. Gartner has not published the criteria or the list, and we read the claim through press coverage rather than the original release.
  • Whether any of these figures hold for ecommerce specifically. None of the surveys break out DTC or ecommerce ops as a segment. Two searches for ecommerce-specific agent deployment data returned nothing usable.

Sources