AI insights

Enterprise AI failure has identifiable causes

The organizations that succeed are more disciplined about four specific things that most organizations skip, rush, or delegate to the wrong people.

By Jacob AndraTalbot West08-01-202610 root causes / 4 severity tiers / 1 diagnostic

Most of the AI work we are asked to assess was designed by someone else: an internal team, a platform vendor, a prior consultant. We see it after it has stalled, or while a buyer is deciding whether to trust it.

The postmortem rarely turns up an exotic technical problem. It turns up a decision made early, by the wrong person, on the wrong basis, and never revisited. Four of those decisions account for most of what we find. All four are made before anyone writes code, which is why we settle them in the first weeks of an engagement rather than finding them in an audit.

And it is not only us

What we see one engagement at a time, four independent research programs have now measured at industry scale.

The measurements
95%of enterprise generative AI pilots produce zero return.MIT NANDA
80%of AI projects fail outright.RAND
50%at least half of generative AI projects are abandoned after proof of concept.Gartner
6%qualify as high performers capturing meaningful financial returns. 88% of organizations now use AI; only 39% see any impact on enterprise EBIT.McKinsey

Those numbers come from MIT, RAND, Gartner, and McKinsey, not from vendors with an agenda. They describe an industry-wide pattern, and the pattern has identifiable causes.

The organizations that succeed are not luckier or better-funded. They don't buy better AI models. Instead, they avoid the single greatest trap in enterprise AI: broad-brush painting. Rather than throwing a generic large language model at a massive, ambiguous workflow, they decompose problems into fundamental primitives, apply the right specialized tool to each, and rigorously scope the integration.

Takeaways
  1. 01The single most consequential decision is choosing what to work on; wrong use case selection produces failure regardless of execution quality.
  2. 02Most organizations apply large language models to problems that require entirely different AI technologies.
  3. 03AI pilots succeed in controlled environments and fail in production because the pilot was designed to avoid the dependencies that production exposes.
  4. 04When the AI roadmap is designed by whoever sells the solution, the recommendations serve the vendor's platform, not the organization's needs.
  5. 05The 5% that succeed treat AI transformation as an organizational change project that happens to involve technology.
01 / Root cause

Solving the wrong problem well

When we reconstruct a failed initiative, the decision that sank it is almost never technical. It is the sentence that named the project.

RAND identified misaligned use cases as the primary root cause of AI project failure. Gartner's analysis of hundreds of GenAI implementations reached the same conclusion from different data: "lack of business value" is the most fundamental failure mode.

Organizations have too many potential AI use cases, not too few. But more critically, they lack the expertise to scope them with specificity. Without a systematic prioritization methodology, selection defaults to broad-brush painting: whoever has the most political capital buys an "AI solution" to "automate customer service" or "streamline operations."

These are not use cases; they are categories.

Success requires decomposing a category into highly specific, solvable primitives: extracting data from a PDF, routing a ticket, evaluating a logic rule. Those primitives are then recomposed into a workflow. Vendor demos and conference hype rarely teach this scoping discipline.

Fig 01 / APEX / five dimensions, three rounds of elimination

We developed a prioritization methodology called APEX specifically to prevent this failure mode. It evaluates opportunities across five dimensions through three rounds of competitive elimination.

Stakeholder urgencyRevenue impactTechnical feasibilityImplementation complexityStrategic alignment

The first round is the most important: it kills the initiatives that feel exciting but score poorly on feasibility or strategic alignment. Most organizations never perform this triage, and the consequences are predictable.

MIT's NANDA report found that the biggest ROI in enterprise AI comes from back-office automation: eliminating outsourcing costs, cutting agency spend, streamlining operations.

Budget follows excitement rather than value. That misallocation is a use case selection failure at the portfolio level.

Where the money goes> 50%More than half of GenAI budgets go to sales and marketing tools.
02 / Root cause

Forcing every problem into the same tool

The most common thing we are handed to review is a large language model pointed at a problem that never called for one.

The AI landscape of 2024-2026 has become fixated on large language models.

Fig 02 / What share of enterprise AI work an LLM actually fits
10%LLM fits
20%Hybrid
70%Requires something else

LLMs handle roughly 10% of enterprise AI work well: text generation, summarization, conversational interfaces.

Another 20% requires hybrid architectures combining LLMs with specialized models.

The remaining 70% requires technologies most organizations never evaluate.

Computer visionPredictive MLOptimization algorithmsTime-series forecastingKnowledge graphsReinforcement learning

When an organization deploys an LLM-based chatbot to solve a problem that actually requires a predictive model, the solution looks impressive in demos and fails in production. The LLM generates plausible-sounding answers that are statistically unreliable for the task. The organization blames "AI." The actual failure point was technology selection.

Case / NYC MyCity chatbot

New York City's MyCity chatbot is the canonical example. The city deployed a generative LLM to answer questions about business regulations: a task requiring deterministic, legally accurate responses. The chatbot told business owners they could legally take workers' tips and serve food contaminated by rodents.

As we have written elsewhere, an LLM generates probabilistic text. Legal compliance requires deterministic accuracy. The technology was wrong for the task. Without a symbolic reasoning layer to enforce rigid rules over the neural network's language capabilities, failure was inevitable.

We use an architectural approach called CHAI (Cognitive Hive AI) that matches specific AI technologies to specific task requirements. A document review workflow might combine computer vision for layout parsing, an LLM for semantic understanding, and a rule engine for compliance checking. Forcing all three tasks onto a single LLM is the wrong tool for two of the three jobs.

03 / Root cause

The pilot worked; reality didn't

We often encounter the AI pilots launched by others that failed when pushed to production. In most cases, the production path was inadequately scoped, and foreseeable dependencies were not mapped.

IBM's CEO Study found that only 16% of organizations achieve scale beyond the pilot stage.

The other 84%

Build something that works in a controlled environment and then discover that productionizing it requires 3-5x the pilot budget, integrations nobody scoped, and organizational changes nobody planned for.

Production reality

A controlled demo with curated data, cooperative users, and a dedicated engineering team will almost always look impressive. Production usually involves messy data, resistant users, edge cases, integration requirements, security constraints, the accumulated friction of organizational complexity.

This pilot-to-production chasm is exactly what separates shallow tooling from deep architectural value. As we outlined in our analysis of the McKinsey ladder and the Talbot West onion, buying an AI feature sits on the outer layer of an organization. It's easy to pilot but offers low leverage. Moving to the center of the onion, where AI is deeply integrated into core workflows and data systems, drives compounding EBITDA value, but requires actual systems engineering. Because the pilot was designed to avoid these deep dependencies, pushing it into production without a true architectural build almost guarantees failure.

Fig 03 / From our engagements / professional services firm, exit preparation

A professional services firm preparing for exit discovered during our technical audit that the AI capabilities they'd piloted couldn't survive integration with their production ERP. The pilot had used sample data in a sandboxed environment. Production required real-time queries against a 20-year-old system with undocumented customizations. The gap between those two realities was months of remediation work that nobody had budgeted for.

We caught it before it became a due diligence problem. Most organizations discover it later.

04 / Root cause

The vendor is designing your roadmap

By the time we are brought in, a roadmap usually exists. The question we ask first is who wrote it and what they were selling.

Two patterns dominate vendor-driven failure.

Pattern 01

A vendor sells their AI platform and then helps the organization find problems for it, inverting the correct sequence of diagnosis before prescription.

Pattern 02

Consulting firms expand engagements beyond the point of diminishing returns because their revenue model rewards duration and headcount, not outcomes.

A significant portion of that waste traces back to solutions designed around the vendor's capabilities rather than the organization's actual needs.

Vendor-driven strategy compounds other failure modes. An organization without internal AI expertise (a common condition) delegates its roadmap to whoever is selling. The vendor prescribes their own platform. The organization cannot evaluate whether the recommendation serves its interests. The result looks like bad luck. It is a predictable consequence of outsourced judgment.

S&P Global, 202546%of AI proof-of-concepts were scrapped before reaching production, at the average organization.

We maintain no vendor partnerships, take no referral revenue, and hold no platform allegiances for exactly this reason. Every architecture recommendation we make reflects what integrates best for the client, not what generates revenue for a platform partner. That structural independence is the prerequisite for honest diagnosis.

05 / The discipline

What the 5% do differently

The organizations that succeed at AI transformation are more disciplined about four things. These are the four we hold an engagement to.

Discipline 01

Choose before building

They invest disproportionate effort in choosing what to build before investing in how to build it. Systematic prioritization across multiple dimensions prevents the most expensive mistake: solving the wrong problem well.

Discipline 02

Match technology to problem

They match AI technologies to business problems rather than forcing every problem into the same technological frame. This requires genuine breadth of technical expertise, which most organizations lack internally and which most vendors cannot provide objectively.

Discipline 03

Prerequisites first

They treat change management, data readiness, and dependency resolution as prerequisites, not afterthoughts. Many AI initiatives fail because of underlying organizational dependencies, not technical limitations. Gartner predicts that 60% of AI projects unsupported by AI-ready data will be abandoned. The AI model works fine. The data it needs doesn't exist, isn't clean, isn't accessible, or isn't governed.

Discipline 04

Design for compounding

They design AI initiatives as components of a system, not isolated tools. Each deployment is scoped to deliver standalone value and create infrastructure that makes the next deployment more powerful. Organizations that deploy AI as features leave the compounding value on the table.

Features are additive / the sum of their individual valuesArchitecture is multiplicative / each component makes the others more valuable

The unglamorous work is diagnosis before prescription.

Understanding what kind of organization you are, what problems actually matter, what dependencies stand in the way, and what architecture will turn individual initiatives into compounding intelligence.

Jacob AndraJacob Andra
About the author

Jacob Andra is the CEO of Talbot West. He hosts The Applied AI Podcast and spends his time pushing the limits of what AI can accomplish in real-world applications. Jacob speaks, writes, and publishes extensively on digital transformation, AI integration, and business process improvement. His expertise spans multiple disciplines, including business strategy, systems integration, digital transformation, and applied artificial intelligence. He's the co-developer of Cognitive Hive AI (CHAI), a modular, composable ensemble framework, and the developer of the Talbot West AI Prioritization and EXecution (APEX) methodology for mapping business opportunities and surfacing the best opportunities for applied AI.

Addendum A1

Deep dive: the full diagnostic framework

The main body above covers the four root causes we see most frequently. The full taxonomy includes ten, mapped against four severity tiers. This section provides the complete framework for readers who want the full picture.

A1-01

The severity spectrum

Some failures are spectacular. Most are quiet. A few masquerade as success.

TierNet impactDefining featureTypical signal
CatastrophicNegative beyond investmentCollateral damage (reputation, regulatory, workforce)Makes the news
CostlyNegative, contained to investmentIntensity of investment with zero returnGets quietly shelved
FizzleNear-zeroAbsence of momentumStops mattering
Mild successPositive but below potentialLocal value without compoundingLooks like success, leaves most value on the table
Catastrophic

Failures extend past the project budget into reputational damage, regulatory exposure, or workforce trauma. Zillow's iBuying algorithm overvalued homes so aggressively that the company lost over $500 million and laid off 25% of its workforce. The institutional scar tissue from that event made data-driven decision-making harder for years afterward.

Costly

Failures consume significant resources with zero meaningful return but keep the damage contained. S&P Global's 2025 survey found that the average organization scrapped 46% of AI proof-of-concepts before production. Real engineering hours, real consulting fees, real opportunity cost.

Fizzles

Never achieve escape velocity. A pilot impresses the steering committee but nobody can figure out how to operationalize it. A tool gets deployed but adoption stalls at 12% of the target user base. The project doesn't die; it stops mattering. RAND found that over 80% of AI projects fail, and most of those failures are fizzles.

Mild successes

Deliver something measurable but capture a fraction of the available value. A company deploys an AI document search tool that saves employees 20 minutes per day. An integrated knowledge management system connecting document search with process automation, institutional memory, and predictive analytics would have delivered multiples of that value. The organization settled for a point solution and declared success.

A1-02

Complete root cause taxonomy

Beyond the four root causes covered in the main body, six additional patterns recur across failed AI initiatives.

05

Insufficient organizational buy-in

Bain's research shows that 88% of business transformations fail to achieve their original ambitions. The technology usually works. The organization doesn't adapt. AI transformation is an organizational change project that happens to involve technology. When leadership treats it as something the IT department handles, adoption stalls.

06

Unaddressed prerequisites

The AI model works fine. The data it needs doesn't exist, isn't clean, or isn't accessible. An AI-powered scheduling system requires standardized process definitions. An intelligent document routing system requires a taxonomy. A predictive maintenance system requires sensor infrastructure. These prerequisites are unsexy, slow, and essential.

07

Fragmented execution

Multiple AI initiatives launch successfully but in isolation. A customer service chatbot here, a demand forecasting model there, an automated document classifier in another department. Each solves its narrow problem. None shares data, context, or capability with the others. The organization has AI tools but not AI intelligence.

08

Governance vacuum

No clear owner of AI outcomes in production. No monitoring for drift, bias, or degradation. No escalation path when the system misbehaves. Apple's credit card algorithm drew complaints and a regulatory review after it reportedly gave women lower credit limits than men, with no explainability or audit trail behind the decision. Amazon's recruiting tool penalized women's resumes for years before anyone tested for bias. Both systems had organizational support. Neither had someone watching the output.

09

Talent gaps

Distinct from buy-in. Buy-in is willingness; this is ability. VW's Cariad division hired 6,000 employees within months, many without automotive software experience. Over $7.5 billion in operating losses followed. The ambition was real. The capability wasn't. Talent gaps compound other root causes because the organization cannot evaluate whether it has selected the right use case, chosen the right technology, or measured outcomes accurately.

10

Measurement failure

Wrong metrics (measuring deployment activity instead of business outcomes), no metrics, or metrics that conflate correlation with causation. Google Flu Trends predicted outbreaks from search query volumes and initially appeared accurate. It then overestimated flu prevalence by up to 140%. The system confused correlation with causation, and the validation methodology could not catch the drift.

A1-03

How severity and root causes interact

Root causeCatastrophicCostlyFizzleMild success
Wrong use case
Wrong technology
Insufficient buy-in
Unaddressed dependencies
Pilot-to-production chasm
Fragmented execution
Governance vacuum
Talent gaps
Vendor-driven strategy
Measurement failure

Multiple root causes reinforcing each other produce the worst outcomes. An organization that selects the wrong use case, applies the wrong technology, and lacks organizational buy-in is heading for catastrophic failure. One that selects good use cases but executes them in isolation with insufficient change management will accumulate mild successes that never reach their potential.

A1-04

Compounding failure patterns

Wrong use case + insufficient buy-in

Leadership champions a high-visibility AI project without securing genuine organizational support. When the project underperforms, it validates the skeptics and poisons the well for subsequent, better-scoped initiatives.

Unaddressed dependencies + pilot-to-production chasm

The pilot succeeds precisely because it sidesteps the dependencies. Clean sample data, controlled environment, small user group. Productionization exposes every shortcut simultaneously.

Governance vacuum + wrong technology

An LLM deployed for a task requiring factual accuracy, with no output monitoring. The system generates plausible but incorrect information. Nobody catches it. Errors propagate until external parties discover them.

Talent gaps + vendor-driven strategy

An organization without AI expertise delegates its roadmap to the vendor selling the solution. The vendor prescribes their own platform. The organization cannot evaluate whether the recommendation serves its interests.

A1-05

Run the diagnostic on your own initiative

The taxonomy is only useful if you can locate yourself in it. Below is one question per root cause. Answer them about a single named initiative rather than about your organization in general, and answer them with evidence rather than impression.

AskA “no” or “not sure” tells you
01
Can you name the business metric this initiative moves, and the size of the move?Wrong use case. Selection is running on enthusiasm or politics, and execution quality will not rescue it.
02
Did anyone evaluate non-LLM approaches before this one was chosen?Wrong technology. The demo will hold and production will not.
03
Can you name the executive whose own numbers change if this works?Insufficient buy-in. Expect a fizzle regardless of how well the thing is built.
04
Does the data this system needs already exist, clean, accessible, and governed?Unaddressed prerequisites. You are funding a model that will starve.
05
Does the pilot run on production data, production systems, and users who did not volunteer?Pilot-to-production chasm. Your budget is understated by a factor of three to five.
06
Does this initiative share data or capability with anything already deployed?Fragmented execution. You are buying tools instead of building intelligence.
07
Is there a named owner of this system's output in production, with monitoring and an escalation path?Governance vacuum. This is the failure mode that turns costly into catastrophic.
08
Could your team tell you if the vendor's recommended architecture were wrong?Talent gap, and it compounds every other answer on this list.
09
Did the roadmap come from someone with no financial interest in the solution?Vendor-driven strategy. Diagnosis is being performed by the party writing the prescription.
10
Are you measuring a business outcome, or a proxy for one (usage, deployment count, sentiment)?Measurement failure. You will not be able to tell success from a mild success, or a mild success from a fizzle.

If you answered no or not sure more than twice, do not try to fix all of it at once. Fix in the order the questions are listed. Use case and technology selection are upstream of everything else; a governance fix on a project solving the wrong problem is wasted motion.

The diagnostic is not a scorecard to feel good or bad about. It is a sequencing tool: it tells you which of your unresolved problems to fix first.

References
  1. MIT NANDA. The GenAI Divide: State of AI in Business 2025.
  2. RAND Corporation. The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed.
  3. Gartner. Why 50% of GenAI Projects Fail — And How to Beat the Odds.
  4. Gartner. Lack of AI-Ready Data Puts AI Projects at Risk, 2025.
  5. McKinsey & Company. The State of AI in 2025: Agents, Innovation, and Transformation.
  6. IBM Institute for Business Value. CEO Study, 2025.
  7. S&P Global Market Intelligence. Generative AI Shows Rapid Growth but Yields Mixed Results, 2025.
  8. Bain & Company. 88% of Business Transformations Fail to Achieve Their Original Ambitions, 2024.
  9. Associated Press coverage of New York City's MyCity chatbot, 2024.
  10. Reporting on the Zillow Offers shutdown and associated workforce reduction, 2021.
  11. Reporting on Apple Card credit limit algorithm complaints and the subsequent regulatory review.
  12. Reporting on Amazon's discontinued AI recruiting tool.
  13. Reporting on Volkswagen Cariad software division losses, 2020-2024.
  14. Lazer, D., Kennedy, R., King, G., and Vespignani, A. The Parable of Google Flu: Traps in Big Data Analysis. Science, 2014.

Figures cited above reflect the source's most recently published data at the time of writing and are attributed to the organization that produced them rather than to Talbot West.