Why AI pilots run aground before reaching production
Experimenting is nearly universal. Reaching production is rare. RAND's research on 65 data scientists explains why, and what a pilot has to prove first.
More than 80 per cent of AI projects fail, twice the failure rate of ordinary IT projects. That is the conclusion of RAND Corporation, based on structured interviews with sixty-five experienced data scientists and engineers, published in August 2024. The main finding is not the technology. It is the people at the top.
This piece is about the gap between experimenting and production, why that gap is nearly universal, and what it takes to cross it.
Why experimenting is easy and production is not
Running a pilot is relatively cheap and low-risk. A small team, a bounded dataset, a few weeks, and at the end a demo that impresses. None of that process tests what happens when the system has to handle real customers, real exceptions and real pressure.
Production is a different world. The system has to work when the data is messy, when a colleague uses it wrong, when there is a peak day nobody foresaw. A pilot proves something can work. Production proves it keeps working, and that is a much higher bar.
A line graph. A pilot proves something can work; production proves it keeps working.
The real causes, according to the research
RAND identifies causes that are overwhelmingly organisational, not technical. Four recur again and again.
No shared definition of success. If the board understands "it works" differently from the team building it, every discussion after the pilot becomes a stalemate with no clear answer.
A weak data foundation. A pilot often runs on a clean, manually assembled dataset. Production runs on the messy data a company has been collecting for years, and those two are rarely the same.
Gaps in infrastructure and integration. A system running standalone is easy to build. A system that has to talk to five existing systems is a completely different project, and that project is rarely budgeted upfront.
Fading executive sponsorship. A sponsor who starts enthusiastic and has other priorities three months later leaves a team with no mandate to solve obstacles outside their own team.
Why this maps exactly onto what we see
These four causes are the same patterns we describe in our own pieces, from a different angle. Why an AI scan is not an IT project is about exactly the ownership RAND names as cause number four. The hidden costs of AI are about the integration and data costs nobody budgets. And the five signals you are stalling describe what happens if nobody addresses these problems structurally.
It is reassuring and sobering at once: this is not a problem specific to smaller Dutch companies. It is the problem of the entire market, documented by one of the largest research institutes in the world.
An empty office. Fading executive sponsorship is one of the main causes of failure.
What a pilot has to prove to bridge the gap
Four questions a pilot has to answer for production, and that most pilots skip because they make the demo less impressive.
Does it work on messy data, not just the clean test set? Test with real, unfiltered data from your own systems, including the exceptions and errors always present in it.
What happens when someone uses it wrong? A pilot tests use as intended. Production tests what happens when someone uploads the wrong file or asks a question the system does not understand.
Who is responsible if it goes wrong after the pilot? One name, not a team. Without clear ownership after launch, maintenance sinks to nobody's priority.
Is there budget for the part after the pilot? Integration, maintenance and adjustment often cost more than the pilot itself. A business case budgeting only the pilot misses half the real work.
Why this is also a story about expectations
There is a second layer under the RAND research worth naming: the definition of failure is itself part of the problem. A pilot never meant to scale, but judged as if it were, counts in the 80 per cent even though nothing was wrong with the project itself. The expectation was the problem, not the execution.
That is exactly why the first question in a project should not be about technology, but about what success means and for whom. A pilot that tests a hypothesis and rejects it succeeded as a learning exercise and failed as a product. Both labels can be true at once, and it is up to leadership to decide upfront which label counts.
What makes a second project easier than the first
The four causes RAND names get smaller on a second project, not bigger, provided the first pilot was properly debriefed. A team that knows why the first project got stuck on integration budgets for it the second time. That is the same curve we described in building capacity without your own AI team: not every project has to repeat the same mistakes, provided someone recorded the previous one instead of forgetting it.
What to do before starting a pilot
Answer the four questions above before you start, not after the pilot succeeds. A pilot that works on the clean test set and where nobody knows who is responsible after launch is almost certainly headed towards the 80 per cent RAND documented.
About this page
The figure of more than 80 per cent and the four main causes come from the RAND report "The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed", published in August 2024, based on interviews with sixty-five data scientists and engineers. This is the state of play on 26 August 2026.
Frequently asked questions
Sources
Tell us what you need.
We respond within 24 hours, from a real human.
Get in touch



