← All posts

95% of enterprise AI pilots fail, and it's not the models

August 01, 2026 · Rta Labs · thesisindustry

The most useful number in enterprise AI this year is not a benchmark score. It's this: roughly 95% of enterprise AI pilots produce no measurable P&L impact. When researchers trace the failures, they land overwhelmingly on implementation, not model capability (MIT NANDA, via TechCrunch).

The models are good. The value dies in deployment: scoping the workflow correctly, wiring it into real systems, governing it, and getting the people who do the work to actually adopt it.

Why is every AI company suddenly hiring forward-deployed engineers?

Watch what AI companies are doing, not what they're saying. Job postings for "forward-deployed engineers" (engineers who embed inside a customer to build and ship the deployment) grew several hundred percent year over year, and the share of companies planning to hire them jumped from under 10% to around 70% in half a year (Perspective AI). Compensation clusters between $300k and $550k. Of roughly 17,000 people in the role in the US, analysts estimate only about 2,000 reliably deliver ROI (TechCrunch). In the first days of July alone, two hyperscalers committed a combined multi-billion dollar sum to standing up forward-deployment organizations of their own.

When an entire industry starts paying half a million dollars per person to solve the same problem, that is the market agreeing, with its budget, on where the bottleneck sits: the deployment layer.

Who gets left out of that math?

Almost everyone. A $500k engineer pencils out against nine-figure strategic accounts and nothing else. Nobody staffs one against a mid-market document workflow, let alone a small business. In AI-agent deployment specifically, buyer guides now describe the high-touch vendors as shipping "a managed implementation, not software," with 8 to 16 weeks to a first live agent (Cresta).

Our analysis: most of the deployment job is software waiting to be written

We broke the forward-deployed role into its parts, using the hiring companies' own job descriptions and engineering blogs as the source material. The job divides cleanly:

The pattern is hard to miss once it's laid out. The middle of the job is repeatable engineering, which means it can be product. The two ends, knowing what to build and getting humans to adopt it, are where the 95% failure rate actually lives, and where scarce human judgment belongs. The industry is paying frontier-lab salaries for a role whose middle should not be manual.

What we built

OMNI is that middle, productized. Upload your documents. The platform auto-labels them against your task schema, trains a compact specialist on your conventions, and deploys it behind guardrails with tiered autonomy. Routine requests resolve locally at near-zero marginal cost. Uncertain ones escalate to a frontier model for review.

We're precise about the boundary: OMNI doesn't replace the judgment of knowing what to build, and it doesn't do change management. What it does is compress the weeks of wiring and agent construction, so scarce human effort goes where it's irreplaceable, and so the economics finally work at every company size rather than just the Fortune 500.

That's the gap the market has spent a year documenting, and it's the one we built OMNI to close, with software economics. As of this week you can procure it through Google Cloud Marketplace.

In the next part of this series we'll expand on the other half of the story, the inference cost crisis: why enterprise AI bills keep rising even as per-token prices collapse, and what the industry's own cost data says the fix looks like.

Benchmarks and methodology ยท contact@rtalabs.org