peekgenticBook a call

Analysis · AI in business

Why 95% of AI pilots fail, and why custom AI is the exception

The research behind the headline AI failure numbers (MIT, Gartner, S&P Global, BCG, IBM, RAND) says generic AI tools bolted onto a business stall. AI built around one real workflow is where the returns show up.

By Miguel GutierrezCo-founder, Peekgentic · Dallas
· 11 min read
Two people reviewing a printed document at an office desk
Most of the value the studies found sits in ordinary back-office work: orders, invoices and the documents behind them.Photo: ThisisEngineering / Unsplash

In brief

The famous AI failure numbers are real, but they are mostly about generic tools that were never fitted to a real job. MIT's 95% figure counts pilots with no measurable profit impact, and the same report found the wins in back-office work, built with specialized partners, using systems that adapt to the workflow. Gartner, BCG and McKinsey point the same way: the projects that pay off are built into the way a team already works. That is what custom AI is.

  • MIT's 95% counts pilots with no measurable P&L impact, not broken software.
  • The same research found the wins in back-office work, built with specialized partners, fitted to the workflow.
  • Gartner, BCG and McKinsey agree: AI that pays off is built into how a team already works.
  • A chatbot subscription for everyone is a tool, not a workflow. Non-technical teams need the workflow designed for them.
  • Start with one desk, one measurable cost, and a person approving every step.

The headline everyone repeats

If you run a company in 2026, you have heard some version of this: 95% of AI projects fail. It shows up in board decks, LinkedIn posts and sales calls from people who want you to buy something else. Usually it is said with a shrug, as if the lesson is that AI is a fad and the smart move is to wait.

The number is real. So are a handful of others like it. But almost nobody reads past the headline, and the research behind it says something more useful than "AI doesn't work." It says a certain kind of AI project doesn't work, and it describes, fairly clearly, the kind that does.

We build custom software for manufacturers and distributors in Dallas-Fort Worth, so we have a stake in this question. That is exactly why we went back to the sources instead of the summaries. Here is what each study actually measured, what it found about the projects that worked, and how that lines up with what we see at real desks.

There is a second story running alongside it, mostly on YouTube and podcasts: the AI companies themselves are burning money. Critics like Cory Doctorow argue the industry has spent on the order of a trillion dollars against roughly $50 billion in revenue. Whether or not that turns into a bubble, it is a question about the economics of the companies selling AI. It says very little about whether a specific piece of software can take retyping off your order desk this quarter. Those are two different bets, and only one of them is yours.

What the 95% actually measured

The 95% comes from MIT's Project NANDA report, The GenAI Divide: State of AI in Business 2025. The researchers looked at roughly 300 public AI deployments, interviewed leaders at companies running them, and surveyed employees. Fortune's coverage put the headline plainly: about 95% of generative AI pilots at companies were failing to produce rapid revenue or measurable profit impact.

Three details matter more than the headline:

  1. "Failure" meant no measurable profit-and-loss impact. It did not mean the software broke. A pilot that people liked, but that never moved a number someone could point to in the P&L, counted as a failure. That is a fair standard for a business, and it is the right one. It is also a different claim from "the technology doesn't work."
  2. Most of the money went to the wrong place. According to the report, more than half of generative AI budgets went to sales and marketing tools, while the clearest returns came from back-office automation: cutting outside processing costs, agency spend and manual operations work.
  3. How the project was built mattered a lot. Projects bought from specialized vendors or built with outside partners succeeded about 67% of the time. Projects companies tried to build entirely in-house succeeded only about one-third as often.

The report's own explanation is what the authors call a learning gap. General-purpose chat tools are great for one person drafting an email. They stall inside a business because they don't remember context, don't learn from corrections, and don't fit into the steps a team actually runs every day. The divide, in the authors' framing, came down to approach, not to which AI model a company used.

It is also worth being honest about the study's limits. It is a preliminary research report, not a peer-reviewed paper, and the sample is modest. Treat the exact percentages as directional. The direction, though, matches everything else below.

The other numbers, read the same way

How many AI efforts actually paid off

Share that delivered value, by study. Each study measures something different, so read the bars as a range, not a race.

  1. MIT NANDA 2025Gen AI pilots with measurable profit impact5%
  2. IBM CEO Study 2025AI initiatives that delivered the expected ROI25%
  3. BCG 2024Companies showing tangible value from AI26%
  4. Gartner 2026IT operations AI use cases that fully met ROI goals28%
  5. BCG 2026Companies now generating value, after reshaping workflows48.5%

Sources: MIT Project NANDA via Fortune (Aug. 2025); IBM Institute for Business Value (May 2025); BCG (Oct. 2024; Sept. 2026: 7.5% "future-built" plus 41% "scaling"); Gartner (April 2026).

S&P Global: companies are walking away from AI projects. In S&P Global Market Intelligence's survey of more than 1,000 companies in North America and Europe, 42% said they had abandoned most of their AI initiatives, up from 17% a year earlier. The average company scrapped 46% of its proofs of concept before they reached production. The top reasons given were cost, data privacy and security risk. Those are the reasons you give when a project was started because AI was available, not because a specific job needed it.

Gartner: the pilots that don't make it out of the lab. In mid-2024 Gartner predicted that at least 30% of generative AI projects would be abandoned after proof of concept by the end of 2025, citing poor data quality, weak risk controls, rising costs and unclear business value. A year later it predicted that over 40% of "agentic" AI projects would be canceled by the end of 2027, and warned about "agent washing": vendors relabeling old chatbots and scripts as AI agents. By Gartner's estimate only about 130 of the thousands of vendors claiming agentic AI are the real thing.

Gartner's most recent data is the most useful. In a survey of 782 infrastructure and operations leaders published in April 2026, only 28% of AI use cases fully succeeded and met their ROI goals, and 20% failed outright. The failures were mostly projects that tried to do too much, too fast. The successes were credited mainly to integrating AI into existing workflows and systems, with real support from business leaders.

BCG: it is a people-and-process problem, not an algorithm problem. BCG's 2024 study of 1,000 executives found that 74% of companies had yet to show tangible value from AI. BCG estimated that about 70% of the challenge sits in people and process, 20% in technology, and only 10% in the AI algorithms themselves, and that the leaders got most of their value from core business processes. BCG's September 2026 follow-up is more encouraging: nearly half of companies now report real value from AI. What separates them is that they reshape workflows end to end and measure the impact directly in their P&L.

IBM: CEOs are still waiting on the return. IBM's 2025 survey of 2,000 CEOs found that only 25% of AI initiatives had delivered the expected ROI over the previous few years, and only 16% had been scaled across the enterprise.

McKinsey: the few who profit redesigned the work. McKinsey's State of AI surveys keep finding the same thing: most companies use AI somewhere, fewer than four in ten can tie any of their earnings to it, and the small group of high performers is far more likely to have fundamentally redesigned the workflows AI touches.

RAND: why projects fail, in the words of the people who build them. RAND's 2024 study interviewed 65 experienced data scientists and engineers. It is often quoted for an "80% of AI projects fail" figure, which RAND presents as an outside estimate, not something it measured. The real value is the list of root causes: leaders misunderstanding the problem the project should solve, missing or poor data, chasing the newest technology instead of a real user's problem, weak infrastructure, and pointing AI at problems it can't yet solve.

Put the studies side by side and one pattern shows up

Failing projects vs. projects that paid off

What the failing projects had in commonWhat the projects that paid off had in common
A general tool looking for a useOne specific job with a cost you can measure
Sales and marketing experimentsBack-office work: orders, invoices, documents, operations
Bolted on next to how people already workBuilt into the existing workflow and systems
Does the same thing every time, never learnsKeeps context, learns from corrections
Built entirely in-house, often by people new to itBuilt with a specialized partner
"Hours saved" guessesImpact measured in the P&L
Tries to change everything at onceScoped small, then expanded

The right-hand column is a fair description of custom AI. Not "custom" in the sense of a giant multi-year build, but in the sense of software made for one team's actual work: their documents, their customers, their system, their exceptions.

Why one-size-fits-all AI software stalls

Most AI products on the market are built to be sold to thousands of companies at once. That is a reasonable business for the vendor, and it is the root of the problem for the buyer. A mass-market tool is set up for an average company that doesn't exist. It doesn't know your part numbers, your customers' habits, the way your ERP wants an order entered, or which exceptions your best person catches without thinking.

A mass-market tool is set up for an average company that doesn't exist.

So it gets close and then stops. It reads most of a document correctly and leaves someone to fix the rest, every time, forever. Nobody inside the company has the time or the know-how to tune it, and the vendor has no reason to tune it for one customer. That is the learning gap MIT described, and it is why so many pilots look good in a demo and stall at the desk.

The exception proves the rule. Off-the-shelf AI works well when the job is narrow and the same everywhere: transcribing a meeting, cleaning up an email, reading a standard W-9. MIT's finding that specialized vendors succeed far more often than generic tools fits that too: a narrow tool built for one job is already halfway to custom. For everything that is specific to how your company runs, someone has to fit the software to the company: teach it your documents, write in your rules, decide what counts as "doesn't match," and keep adjusting as it meets real work. Without that person, the project is set up to fail.

A chatbot login is not a workflow

The most common AI rollout we hear about goes like this: the company buys everyone a ChatGPT or Claude subscription, sends a memo saying "use it," and waits for productivity to go up. Often the employees were already using AI anyway. MIT's researchers found that only about 40% of companies had bought an official AI subscription, while workers at more than 90% of them were using personal AI tools for their jobs.

A subscription is a powerful tool for one person drafting an email or summarizing a document. Keep it. But it is not a workflow. Someone in purchasing who is not technical has no way to know how to turn a chat window into a process that reads every supplier confirmation, checks it against the PO, updates the ERP and flags the ones that don't match. That takes knowing which steps to automate, where the data lives, how to connect the systems, and where a person has to approve. Without that, every employee invents their own prompts, the results don't add up across the team, and nobody can measure whether anything got faster.

The training numbers show the gap. In BCG's surveys of workers, only about a third say they have been properly trained on AI, and most are given no guidance on what to do with the time it saves. People who got at least five hours of training were much more likely to become regular users. And feeling faster is not the same as being faster: in a 2025 randomized trial, experienced software developers took 19% longer on real tasks when they used AI tools, yet believed afterward that AI had sped them up by about 20%. If experts misjudge it, an office team with no technical help will too.

This is the job custom AI does. Someone designs the workflow around the desk, builds it into the systems the team already uses, and measures it, so the people at the desk get the benefit without having to become engineers.

What we see at real desks

We spend our time inside the offices of manufacturers and distributors, watching how orders, purchases and invoices actually move. A few things show up almost every time, and they line up with the research.

A crowded order-log spreadsheet with color-coded cells and #REF errorsThe same orders on a board where each one runs through steps and waits for a person where needed
Before and after at one order desk: the shared order log most offices know, and the same orders running as steps a person approves. Peekgentic demo, sample data

The value is in the boring work. The desk that pays back is rarely the flashy one. It is the person reading a customer's PO out of an email, retyping it into Sage or the ERP, then typing it again for the job and the paperwork. That is exactly the back-office category MIT found the clearest returns in. Our first live build does this for a food-equipment manufacturer that supplies national grocery chains: the software reads each order and fills it in, and their team approves it.

The hard part is the exceptions, not the AI. Reading a clean PO is easy. The work is in the orders where a customer uses an old part number, a price doesn't match, or a ship-to address is new. Generic tools guess. A system built for that desk is told what "doesn't match" means for that company, and it stops and asks a person instead of guessing. This is the learning gap MIT described, closed on purpose.

The rules live in someone's head. Every office has a person who just knows that one customer always means the left-hand model, or that another wants partial shipments. Off-the-shelf software can't know that. Custom software can, because we write those rules down with that person and build them in. The knowledge stops walking out the door at 5 p.m.

People adopt what fits how they already work. When the software works inside the email, spreadsheet and accounting system the team already uses, and a person approves every step, it gets used. When it asks a team to change how they work to fit a new platform, it becomes one more pilot that stalls. That is Gartner's top success factor, seen from the desk.

Small and measured beats big and vague. We start with one desk, time it for a week, build, and time it again. That is the "measure it in the P&L" habit BCG found in the companies that are winning, scaled down to a 40 to 400 person company.

Where custom AI is not the answer

To be fair to the other side: custom isn't always right. If a good off-the-shelf product already does the job the way your team works, buy it. If a process happens a few times a month, automating it rarely pays. And if nobody can say what a win looks like in hours, dollars or orders per person, no AI project, custom or not, should start yet. RAND's first root cause, misunderstanding the problem, applies to everyone.

A short checklist before you spend on AI

  1. Name one job. "Order entry from emailed POs," not "use AI in operations."
  2. Write down today's cost. Minutes per item, items per week, and who does it.
  3. Keep the system you have. The work should land in your existing ERP, accounting system or spreadsheets.
  4. Keep a person approving. Every step logged, every action reversible, until the numbers prove it out.
  5. Plan for the exceptions. Ask how the software behaves when something doesn't match. "It stops and asks" is the right answer.
  6. Measure it like money. Compare the before and after on that one desk, then decide on the next one.

Common questions

Is it true that 95% of AI projects fail?

MIT's 2025 GenAI Divide report found about 95% of enterprise generative AI pilots had no measurable profit-and-loss impact. That measures business impact, not whether the software worked, and the same report found much better results for back-office automation built with specialized partners.

Why do most AI pilots fail?

Across MIT, Gartner, BCG and RAND the same causes repeat: a general tool with no specific job, projects bolted on next to the real workflow, tools that don't learn from corrections, poor data, and no clear measure of value.

Does custom AI work better than off-the-shelf AI?

The research points that way for business processes. Gartner credits integration into existing workflows and systems for successful AI use cases, and MIT found specialized, workflow-fitted tools built with partners succeeded far more often than generic or purely in-house projects.

Where does AI pay off first for a manufacturer or distributor?

Back-office work with a clear cost per item: order entry from emailed POs, purchasing, invoice matching, quotes and dispatch. MIT found the clearest returns in back-office automation.

How should a small or mid-size company start with AI?

Pick one desk, time the work for a week, automate the typing while a person approves every step, and measure again before expanding.

We gave everyone a ChatGPT or Claude subscription. Isn't that enough?

It helps individuals with drafting and summarizing, but it is not a workflow. Most employees are not technical and have no way to turn a chat window into a process connected to your systems. BCG finds only about a third of workers feel properly trained on AI. The returns come when someone designs the workflow, builds it into your systems and measures it.

Who builds custom AI for companies in Dallas-Fort Worth?

Peekgentic is a Dallas firm that builds custom AI automation for manufacturers and distributors, around the systems and workflows they already use.

Sources

  1. MIT Project NANDA, The GenAI Divide: State of AI in Business 2025, as reported by Fortune, Aug. 18, 2025
  2. S&P Global Market Intelligence survey, via CIO Dive, 2025
  3. Gartner, “30% of Generative AI Projects Will Be Abandoned After Proof of Concept by End of 2025,” July 29, 2024
  4. Gartner, “Over 40% of Agentic AI Projects Will Be Canceled by End of 2027,” June 25, 2025
  5. Gartner, “AI Projects in Infrastructure and Operations Stall Ahead of Meaningful ROI Returns,” April 7, 2026
  6. BCG, “AI Adoption in 2024: 74% of Companies Struggle to Achieve and Scale Value,” Oct. 24, 2024
  7. BCG, “AI Is Starting to Pay Off. Almost 50% of Companies Now Generate Value with It,” Sept. 30, 2026
  8. IBM Institute for Business Value, 2025 CEO Study, May 6, 2025
  9. McKinsey, The State of AI
  10. RAND, The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed, 2024
  11. Fortune, “MIT report: the shadow AI economy,” Aug. 19, 2025
  12. BCG, AI at Work 2025: Momentum Builds, but Gaps Remain, June 2025
  13. BCG, AI at Work: Why Strategy Matters More Than Tools, 2026
  14. METR, “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity,” July 10, 2025
  15. PPC Land, “Cory Doctorow puts AI revenue at $50bn against $1tn of spending,” 2026

More from Peekgentic

Published October 9, 2026 · Written by Miguel Gutierrez, Co-Founder, Peekgentic (Dallas, TX). Every statistic links its source.