The numbers have settled into a depressing consensus. Of every ten AI pilots a mid-sized organization launches, two reach production at meaningful scale, three live on as small embedded use cases, and five die quietly. The death is rarely announced; the pilot just stops being mentioned in the steering committee deck, and a year later no one is quite sure whether it ever shipped.
We have run the post-mortem on more than sixty of these pilots in the last two years, across banking, public sector, energy, and consumer goods. The cause of death is almost never the model. It is almost always the operating system the model was meant to plug into. Mid-sized organizations — say, $200m to $5bn in revenue — sit in the worst position for this: large enough that change management is real work, small enough that there is no dedicated MLOps team to make it look easy.
This article is for the mid-sized leader who is now on their second or third pilot, has seen the first stall, and is starting to wonder whether the technology was oversold. It was not oversold. The pilot was just structured to fail.
Why pilots fail to scale
Pilots fail in a predictable funnel. Of the typical hundred candidate ideas a mid-sized organization considers in a year, roughly forty get approved for proof-of-concept. Half of those produce a working demo. A third of those make it past procurement, security review, and integration. About one in five reach a production deployment. And of those, perhaps two-thirds operate at the scale originally promised in the business case. The compound attrition is brutal — but each stage of attrition has a specific operational cause.
The four operating-model gaps that kill pilots
The pilots that die share four characteristics, and they are remarkably consistent across sectors.
1. Ownership ends at the demo
The most common version of pilot-death: a project sponsor commissions a proof-of-concept, the vendor builds a model, the demo works, everyone applauds — and then no one in the operating business has been asked to commit to running it. The sponsor goes back to their day job. The model sits on a server, technically functional, but with no production owner, no SLA, and no budget line. After six months, the server is decommissioned in a routine cleanup.
The fix is structural, not technical. Before the POC starts, a named operating-business owner must commit to running the system if it works — and that commitment must include a P&L line for ongoing costs, a named accountable executive, and a defined service level. If no one will sign that commitment, do not run the pilot. The result will be the same as not running it, but you will have spent the money.
2. The data foundation was assumed, not audited
Every pilot business case quietly assumes that the production data will look like the training data. In about 70% of the pilots we reviewed, it didn't. Sometimes the training data was a sample that excluded the messy 30% of cases. Sometimes the production data lived in a different system with different field semantics. Sometimes — and this is the painful one — the production data simply did not exist at the volume or quality the model needed.
A serious pilot includes a data audit before the model is built, not after. It costs a few weeks. It saves the eight months that follow when production data turns out to be different from the training set.
3. The change-management work was deferred
The third pattern: the model is technically correct, the data is fine, and the production owner is committed — but the people who are supposed to use the model do not. They have not been involved in its design. They do not trust its outputs. They have workarounds that are 80% as good and that they understand. The model becomes a system everyone references in board presentations and no one actually relies on for decisions.
Most AI pilots are won or lost in the eighteen months before the model is built. By the time the data scientists arrive, the operating-model choices that will determine the outcome have already been made.
4. Procurement and security caught up at the wrong time
Pilots tend to skate past procurement and security in the POC phase — the spend is small, the data is synthetic or anonymized, and no one wants to slow the experiment. Then production looms, the spend is real, the data is real, and procurement is meeting the project for the first time. We have seen pilots add six to nine months at this stage. Two of them died there.
The fix: bring procurement and security to the design table early. Their constraints are not obstacles to engineer around; they are inputs to the architecture. A model that cannot pass security review is not a model.
What the 20% do differently
Across the pilots that made it to production at promised scale, three patterns recur. They are not surprising. They are deliberately mundane.
They pick boring problems with clear ROI. The pilots that scale are not the moonshots. They are the ones that automate a well-defined, high-volume, repetitive task with a measurable improvement in cycle time or accuracy. Document classification, case triage, demand forecasting, and contract review dominate the list. The exciting use cases — generative content for customers, fully autonomous decisioning — almost never appear among the survivors.
They build for the operating model, not the model. The 20% organizations spend roughly 60% of total program cost on data foundations, integration, change management, and governance — and 40% on the model itself. The failing pilots inverted this ratio. The model is a small fraction of the work. Pretending otherwise is the most common single mistake.
They define what they will stop doing. A surprising pattern across the pilots that scaled: the production owner had publicly committed, before launch, to retiring a competing manual process. They did not run "AI alongside the existing workflow" — they replaced it. This forced the model to be good enough, and it forced the organization to actually depend on it.
- Pick boring problems. Document classification and case triage scale. Generative customer content and autonomous decisioning rarely do.
- Budget 60% for everything-but-the-model. Data, integration, change, governance. The model is the cheap part.
- Name the production owner before the POC starts. If no one will commit to running it, do not build it.
A capability map for where to start
For an organization on its second or third pilot, the most useful diagnostic question is not "which use case next?" but "what is our actual maturity on the supporting capabilities?" A model with no integration team, no data foundation, and no change-management capacity will not produce production value — even if it is technically perfect.
We map mid-sized organizations on two axes: their capability to deploy AI (engineering, data, integration, governance) and their organizational readiness (clear sponsorship, change capacity, decision rights). The intersection tells you what kind of pilot will actually work for you — and what kind will burn cash.
The most common position — by a long way — is bottom-left: stall zone. That is not a failure. It is the honest starting point for most mid-sized organizations. The mistake is to act as if you are top-right, buying an enterprise AI platform when what you needed was a single use case, one engineer, and six months of operating-model work.
A 90-day playbook
For an organization sitting in the stall zone and trying to break out, we recommend a short, structured ninety days before any new model is built. The work is unglamorous — but it determines whether the next pilot is a success or another quiet death.
Weeks 1–3: Diagnose. Map your previous pilots: where did each die, and on which of the four operating-model gaps? Map your data foundations: where does production data actually live, in what state, governed by whom? Map your change capacity: how many concurrent transformation programs is the organization already running?
Weeks 4–8: Choose one boring problem. Pick a single, high-volume, well-defined use case from the document-classification or triage family. Name the production owner. Get them to commit, in writing, to retiring a specific competing manual process if the model works.
Weeks 9–13: Build the operating model first. Procurement, security, governance, integration, change management. The model can wait. By the end of this phase, you should be able to explain — to a sceptical board — exactly who will run the system, what they will stop doing, and what the failure modes are.
The hardest sentence to say in an AI program is "we will stop doing this." It is also the sentence that determines whether the program works.
What this means for leaders
The temptation, for a leader looking at the failure rate on AI pilots, is to conclude that the technology is not ready. It is. The pilots are failing on operating-model questions — ownership, data, change, procurement — that have nothing to do with the technology and would defeat any other capital project too.
The mid-sized organizations that get this right do not look like Silicon Valley success stories. They look like steady, disciplined operators that picked one boring problem, built the supporting machinery before they built the model, and committed publicly to retiring a competing process. They scale not because they have a 10x model but because they have a 1x operating model that the model can actually run inside.
For everyone else, the question to ask before the next pilot is not "what should we build?" It is "what stalled the last one?" The answer will tell you what to build next — and what to fix first.