The 95% Problem: How to Buy Automation That Actually Delivers ROI
The single most-quoted number in business automation right now is a discouraging one. In its report The GenAI Divide: State of AI in Business 2025, the MIT NANDA initiative found that roughly 95% of enterprise generative-AI pilots produced no measurable impact on profit and loss — despite an estimated $30-40 billion poured into them. It is the kind of statistic that makes a cautious buyer sit on their hands. But the more useful story is not the 95% that failed; it is the 5% that worked, and what they did differently. Read closely, MIT's finding is not a verdict on the technology at all. It is a verdict on how automation gets bought. This guide translates that research into a practical buying discipline, so your next project lands in the 5% instead of the pile.
What the data actually says
It is worth being precise about the study, because the headline has travelled further than the detail. MIT's team drew on 52 structured executive interviews, a survey of 153 leaders, and an analysis of 300 publicly disclosed AI deployments. Their central finding was a stark split they named the "GenAI Divide": adoption is nearly universal, but transformation is rare. Tools are everywhere; returns are not.
Other 2026 datasets rhyme with it. McKinsey's State of AI work reported that around 88% of organisations now use AI regularly in at least one function, yet only about 6% qualify as high performers who can attribute more than 5% of earnings to it. IDC has projected that close to half of AI-driven use cases will miss their ROI targets in 2026, citing unclear business goals, weak human-machine hand-offs and shaky data foundations. Different researchers, different methods, same shape: a wide base of activity sitting on a thin sliver of value.
The temptation is to read this as "AI is overhyped, wait it out." That is the wrong lesson, and an expensive one, because the same reports show a minority of buyers getting real, repeatable returns today. The question a buyer should ask is not whether automation pays, but what the payers have in common. MIT's answer is uncomfortable for anyone hoping the model does the work for them: the difference is almost entirely in how the tool was chosen, scoped and embedded.
Finding #1: buying beats building — by a lot
The most actionable number in the whole report is easy to miss. When MIT split deployments by where the tool came from, automation sourced from specialised external vendors and partners reached successful deployment roughly 67% of the time. Tools built in-house succeeded at about half that rate. Same ambition, same models available to both camps, and the buyers roughly doubled the builders' hit rate.
That is counterintuitive to a lot of capable engineering teams, who reasonably assume that no one understands their process better than they do. But understanding the process is not where pilots die. They die in the unglamorous middle: the edge cases nobody documented, the integration that turned out to be three integrations, the second and third iterations that a busy internal team never gets back to once the demo is signed off. A specialist vendor has usually already paid that tax on someone else's project. Buying is, in effect, buying past the most common failure points rather than rediscovering them.
This does not mean custom builds are dead — far from it. Retool's 2026 build-versus-buy research found that 35% of enterprises had already replaced at least one SaaS tool with a custom build, and most expected to build more. The reconciliation is straightforward: buy the commodity majority of your automation, where a vendor has solved a problem thousands of times, and reserve scarce internal engineering for the genuinely differentiating slice that no one else can build for you. Deciding where that line sits is exactly the trade-off our guide to no-code versus custom automation works through in detail.
Finding #2: the money is in the back office, not the storefront
MIT found a consistent misallocation of budget. Respondents steered the large majority of their hypothetical GenAI spending — on the order of 70% — toward sales and marketing use cases, where the results are visible and easy to show a board. Yet the stronger, more measurable returns were showing up in back-office automation: document and invoice processing, customer-service triage, reconciliation, the quiet plumbing between systems. That work is underfunded precisely because it is boring, and it is where the report saw organisations cut millions in outsourced-process and support spend.
There is a clean logic to why back-office work pays first. The volume is high and repetitive, so a small saving per item compounds fast. The output is checkable against a clear right answer, so you can actually measure quality. And the baseline cost is already known, because someone is doing the work today. Marketing copy generated by AI is pleasant to have, but the value is diffuse and hard to attribute; an invoice-processing pipeline that clears 4,000 documents a month at a known cost per document is a spreadsheet you can defend.
| Where budgets go | Why it feels attractive | Why ROI stays shallow or hard to prove |
|---|---|---|
| Sales & marketing content and outreach | Visible, board-friendly, easy to demo | Diffuse value, hard to attribute, quality is subjective |
| Customer-facing chat and "assistants" | Feels transformational, high engagement | Success depends on messy real conversations; deflection is easy to overstate |
| Back-office document & data processing | Unglamorous, so it gets skipped | Actually where returns are strongest — high volume, checkable, known baseline |
| Internal reconciliation & system-to-system data | Invisible to leadership | Underfunded, but reliable savings and low regulatory exposure |
For a buyer, this is a gift, because the highest-ROI opportunities are the ones your competitors are ignoring. If you are unsure where your own back-office volume concentrates, our walkthrough on which business processes to automate first gives a repeatable way to rank candidates by volume, checkability and cost.
Finding #3: the "learning gap" is the real failure mode
MIT's own explanation for the 95% is a phrase worth internalising: a "learning gap." The stalled pilots did not fail because the model could not draft an email or summarise a document. They failed because the tools did not learn — they did not retain feedback, adapt to the specifics of the business, or get better with use. A system that makes the same mistake on Friday that it made on Monday, no matter how many times a human corrected it, never earns the trust required to become a real process. It stays a demo forever.
This reframes what a buyer should be inspecting. The impressive part of any automation demo is the first output. The part that predicts ROI is the hundredth, after a month of corrections. When you evaluate a tool, the questions that matter are less "how good is the answer" and more "what happens when the answer is wrong, and does the system get better because I told it so?" A tool that closes the learning gap is one you can hand a genuine process; a tool that cannot is a novelty you will quietly stop using.
This is also why so many pilots look successful and then evaporate — the enthusiasm of the launch fades into the reality of a tool that never improves. If your own past automation projects delivered less than the brochure promised, the mechanics behind that gap are worth understanding directly; our analysis of why automation ROI comes in lower than expected covers the recurring reasons a promising pilot underdelivers.
A buyer's checklist for landing in the 5%
Put the three findings together and they become a purchasing discipline. Before you commit budget to any automation or AI-agent purchase, work through these questions. None of them is about the model; all of them are about how the tool will live inside your business.
- Is there a specialist who already sells this outcome? If yes, the base rate strongly favours buying over building. Make the internal-build case only where the process is a real differentiator.
- Is this a high-volume, checkable process? Favour work where you can count the items and verify the output against a right answer. That is where returns are measurable and where the 5% concentrate.
- What is the baseline, in numbers? Write down today's cost — hours, cost per item, error rate — before the pilot starts. You cannot prove ROI against a baseline you never recorded.
- Does the tool learn from correction? Ask specifically what happens after a wrong output. A system that adapts to feedback closes the learning gap; one that does not will stall.
- How deep is the integration, really? Confirm the tool connects to the systems the process actually touches. Integration debt, not model quality, is where most pilots quietly die.
- Who owns iteration after go-live? A named owner and a maintenance path separate the automations that keep paying from the ones that decay. Buying the build without buying the upkeep is a common, expensive mistake.
- What number would justify rolling it out? Agree the success threshold before you start, so the decision to scale is a fact, not a feeling.
How to run a pilot that can actually prove itself
A striking share of the failures MIT catalogued were not disasters — they were pilots that simply never resolved into a yes or a no. They ran, produced some output, generated some enthusiasm, and then drifted because nobody had defined what winning looked like. A pilot designed to be conclusive is a different thing from a pilot designed to be impressive.
- Pick one process, not a theme. "Automate invoice intake for the France entity" is a pilot. "Explore AI for finance" is a budget line that will never close.
- Set the baseline first. Measure the current cost and error rate for two weeks before the tool touches anything, so the comparison is real rather than remembered.
- Time-box it to six to eight weeks. Long enough to get past the honeymoon and into the corrections, short enough that an inconclusive result is cheap.
- Define the go/no-go number up front. Decide what saving or accuracy level would justify rolling out, and write it down where you cannot quietly move it later.
- Instrument the corrections. Track how often a human has to intervene, and whether that rate falls over the weeks. A falling intervention rate is the single best leading indicator that the tool is closing the learning gap.
- Decide, then either scale or kill. The one outcome to avoid is the zombie pilot that neither ships nor dies. Both a clear yes and a clear no are wins; ambiguity is the only real loss.
Buyers who run pilots this way get a second, subtler benefit: the discipline forces the vendor conversation onto outcomes rather than features. When you know your baseline and your target number, you can also structure the commercial deal around results, which is where the market is heading anyway. Our buyer's guide to outcome-based automation pricing covers how to tie what you pay to what you actually get.
The contrarian read: 95% failure is a buyer's opportunity
It would be easy to close the MIT report and conclude that automation is a coin flip you should sit out. The 5% tells the opposite story. The failures cluster around a handful of avoidable choices — building what you could have bought, funding the visible over the valuable, and running pilots with no baseline and no target. None of those is a technology problem, which means none of them requires waiting for better models. They require buying with discipline, and that is available to you today.
There is a competitive edge hiding in the gloom, too. When the majority of the market is pouring money into marketing demos that cannot prove their worth, the buyer who quietly automates high-volume back-office work, from a specialist, against a measured baseline, is compounding a real advantage while everyone else runs zombie pilots. The GenAI Divide is not only a description of who is failing; it is a map of where the unclaimed value is sitting. The 95% is a warning. The 5% is a playbook — and it is a buying playbook, not a technical one.
Buy automation that lands in the 5%
Browse ready-made workflows and vetted specialists who ship measurable, back-office-first automation — the pattern the data says actually pays off.
Explore the FlowMarket marketplaceFAQ
What is the 95% figure everyone is citing?
It comes from MIT's report The GenAI Divide: State of AI in Business 2025, published in August 2025 by the MIT NANDA initiative. Across 52 executive interviews, a survey of 153 leaders and an analysis of 300 public AI deployments, it found that about 95% of enterprise generative-AI pilots produced no measurable profit-and-loss impact, while only around 5% delivered real value.
Does 95% failure mean the technology does not work?
No. MIT was explicit that the failures were not caused by weak models but by a "learning gap" — poor integration into real workflows, tools that do not remember or adapt, and misplaced budgets. The same report found a minority of deployments delivering strong returns, so the gap is about how organisations buy and embed automation, not whether the technology can work.
Why do purchased tools succeed more often than internal builds?
In MIT's sample, automation sourced from specialised external vendors and partners reached successful deployment roughly 67% of the time — about twice the rate of tools built in-house. Specialist vendors have already solved the integration, edge cases and iteration that internal teams tend to underestimate, so buying skips the most common ways a pilot stalls.
Which processes should a buyer automate first for ROI?
The data points to unglamorous back-office work — document processing, customer-service triage, reconciliation, data entry between systems — where the volume is high, the output is checkable and the savings are measurable. MIT noted that budgets skew toward sales and marketing where returns are visible but shallow, while back-office automation, where ROI is stronger, stays underfunded.
What is the "learning gap" MIT describes?
It is the gap between a tool that can demo impressively and one that improves inside your business. Most stalled pilots used systems that did not retain feedback, adapt to context or get better with use, so they never moved from novelty to dependable process. Closing the gap means choosing tools that learn from corrections and wiring them into a workflow people actually run.
How do I write a pilot so it can actually prove ROI?
Pick one high-volume process, define the baseline cost and a single success metric before you start, scope the pilot to six to eight weeks, and agree in advance what number would justify rolling it out. A pilot with no baseline and no target cannot land in the 5% because there is nothing to measure it against, however well the tool performs.
Is building custom automation ever the right call?
Yes, when the process is a genuine competitive differentiator or no vendor covers it. Retool's 2026 build-versus-buy research found 35% of enterprises had already replaced at least one SaaS tool with a custom build. The point is not never to build, but to buy the commodity 80% and reserve scarce engineering effort for the part that is truly yours.
How long before automation shows a return?
For a well-scoped back-office automation bought from a specialist, teams typically see measurable savings within one to three months, because the work is high-frequency and the baseline is easy to compare against. Broad, open-ended "transformation" programmes are the ones that drag on for quarters without a clear number, which is exactly the pattern the 95% share.