The short answer
Enterprise AI agents that reach production report an average of roughly 171% ROI, with U.S. enterprises reporting 192%, in self-reported 2025-2026 deployment surveys. But that number is survivorship-biased: 95% of generative AI pilots deliver zero measurable P&L impact (MIT Project NANDA, July 2025). The ROI is real. The failure rate is also real. The gap between the two is not a model problem — it is a production problem. This post gives you the hard numbers for both sides of that gap, and the operating pattern that separates the 5% that ship from the 95% that stall.
Key Takeaways
- ~171% average ROI (self-reported) on agentic AI deployments that reach production; ~192% for U.S. enterprises; a majority of executives report positive ROI inside the first year. (AI Monk aggregated survey data, 2025-2026)
- 95% of GenAI pilots return nothing — $30-40B invested, no measurable business return — per MIT’s The GenAI Divide (July 2025). The differentiator is not the model; it’s whether the agent reaches production under real governance.
- Named proof: Klarna’s assistant did the work of ~700 agents in month one, cut resolution from 11 min to 2, and was projected at ~$40M profit improvement (Klarna). Best-in-class customer-service agents hit 80%+ resolution (not deflection).
- Buy/partner beats internal build ~2:1: vendor and forward-deployed partnerships succeed roughly twice as often (a 2:1 ratio) (MIT NANDA).
- Back-office is underfunded relative to its return. Over half of GenAI budgets go to sales and marketing; MIT found the largest ROI in back-office automation.
- The FDE signal: forward-deployed engineer postings grew ~700% YoY into 2026 — the market is pricing in embedded delivery, not shipped software alone.
What “ROI” actually means for an AI agent in 2026
For an enterprise decision-maker, agent ROI reduces to one formula:
Value released = (hours automated × loaded headcount cost) + error/rework avoided + revenue protected — (agent build + run + governance cost).
The reason the headline 171% figure is compelling and misleading at the same time: it is self-reported and measured on agents already in production. The 95% that never cleared pilot never entered the denominator. So the honest framing for a board is two questions, not one:
- What does an agent return once it’s live? (Answer: 150-192% is the current benchmark band.)
- What’s my probability of getting one live? (Answer: low the way most pilots are run — but buying or partnering succeeds roughly twice as often as building internally, per MIT NANDA.)
2026 ROI benchmarks by number
| Benchmark | 2026 figure | Source |
|---|---|---|
| Average agentic AI ROI (self-reported) | ~171% (U.S.: ~192%) | AI Monk survey data, 2025-2026 |
| Executives seeing ROI in year one (self-reported) | majority | AI Monk |
| GenAI pilots with zero P&L impact | 95% | MIT Project NANDA, GenAI Divide, Jul 2025 |
| Enterprise GenAI investment with no return | $30-40B | MIT NANDA |
| Buy/partner success rate vs. internal build | roughly twice as often (a 2:1 ratio) | MIT NANDA |
| Best-in-class CX agent resolution rate (vendor-reported) | 80%+ | Fin.ai, 2026 |
| Typical resolution, independent benchmark (reality check) | often well below headline claims | Lorikeet, 2026 |
| FDE job posting growth (delivery-model signal) | ~700% YoY into 2026 | AOL / hiring data |
Information gain — the number vendors hide: resolution ≠ deflection. Deflection counts conversations no human touched; resolution counts problems actually solved. Independent benchmarks consistently find real-world resolution runs well below headline vendor claims. When you underwrite ROI, underwrite the resolution number, not the deflection number.
Proof, not promises: named production agents
The winning enterprise-AI content format is no longer thought leadership — it’s the case study with hard numbers. Here’s the current proof band:
- Klarna (customer service): in its first month live, the OpenAI-powered assistant handled 2.3M conversations — the workload of roughly 700 full-time agents — cut average resolution from 11 minutes to 2, and was projected to drive ~$40M in profit improvement in 2024. (Klarna press release; OpenAI)
- Back-office automation: MIT NANDA identifies back-office workflows — invoice processing, reconciliation, contract review — as the highest-ROI, most underwriteable category, even though most budgets chase front-office use cases. (MIT NANDA, 2025)
Why 95% stall: the pilot-to-production gap
MIT’s finding is not that the models are weak. It’s that static pilots can’t retain context, run under governance, or touch real data — so they never move the P&L. The five failure modes we see repeatedly:
- Demo data, not production data. The agent works on a curated set and breaks on the real one.
- No governance/identity model. It can’t get access to the systems where the value lives, so it stays a toy.
- No owner in the repo. Nobody is embedded to close the last-mile integration gaps.
- Deflection theater. Success is reported as deflection; the board later discovers real resolution is half that.
- Wrong budget allocation. Over half of GenAI spend chases sales/marketing flash while the durable ROI sits in the back office.
The pattern that beats these — the one MIT’s buy/partner cohort is implicitly buying — is embedded delivery: a forward-deployed engineer who sits in your repo, standups, and Slack, and ships production code running under real governance, on real data, in weeks.
The Nucleo Labs position: forward-deployed, production-first
The ~700% surge in forward-deployed engineer demand is the market repricing how AI gets delivered. Shipped software is table stakes; shipped outcomes under governance are the moat. Nucleo Labs embeds forward-deployed engineers to move agents from perpetual pilot to production — model-agnostic (benchmark, route, and optimize across leading LLMs rather than betting on one), governed, and measured against the resolution number, not the deflection number.
That is the difference between the 171% you read in a benchmark and the 171% that shows up in your own P&L. If your pilots keep stalling before that point, start with why 95% of enterprise AI pilots fail, then see the industries we deploy in.
FAQ
What is a realistic ROI for an enterprise AI agent in 2026? For agents that reach production, roughly 150-192% is the self-reported benchmark band (~171% average, ~192% for U.S. enterprises), with a majority of executives reporting positive ROI within the first year. The caveat: that band is measured only on live agents, and 95% of pilots never get there.
Why do 95% of AI pilots fail to deliver ROI? Per MIT’s July 2025 GenAI Divide report, most pilots use demo data, lack a governance/identity model, and can’t retain context or act autonomously — so they never touch the P&L. Buying from specialized vendors or partnering succeeds about twice as often as internal builds (a 2:1 ratio).
Which use cases have the best ROI? MIT found the largest returns in back-office automation (BPO elimination, agency-cost reduction, process streamlining), even though most budgets chase sales and marketing. The five canonical production use cases are customer service, contract review, supply chain, code modernization, and fraud detection.
What resolution rate should I expect from a customer-service agent? Vendor-reported best-in-class agents reach 80%+ resolution, but independent benchmarks consistently find real-world resolution runs well below those headline claims. Insist on resolution (problems actually solved), not deflection (conversations merely avoided), and validate it on your own traffic.
Should we build agents internally or partner? The data favors partnering roughly 2:1. The differentiator is embedded, forward-deployed delivery — engineers who ship production code on your real data and governance, which is why FDE demand grew ~700% year over year into 2026.
Related reading
- Why 95% of enterprise AI pilots fail (and the 5% that don’t)
- What is a forward-deployed engineer — and why enterprise AI needs one
- Industries we deploy in
Nucleo Labs commits to the outcomes up front and gets measured on them. Book a strategy call and we’ll come back with the ROI we can defend.
Sources: MIT Project NANDA, The GenAI Divide (2025); AI Monk enterprise ROI case studies (2025-2026); Fin.ai CX ROI benchmarks (2026); Lorikeet resolution-rate benchmarks (2026); Klarna press release.