From AI POC to production: why it stalls, and how to get through
The proof of concept (POC) works, the demo won people over, and six months later nothing is in production. The story has become routine. It rarely comes down to the model: it almost always comes down to everything around it.
In many organisations, you will count more than a dozen AI POCs. Some were launched by IT, others by a business unit, a few by a motivated team with a subscription to an assistant. Most produced a successful demo. Very few run every day, on real data, for users who were not part of the project.
People often quote spectacular figures for this failure rate. I prefer to stick to what I see: the move from POC to production is where most AI projects stop, and the reasons look much the same from one organisation to the next.
A POC proves one thing; production demands ten
A POC answers a single question: can the model do the task, on a chosen sample, in good conditions? That is useful. But it is a small part of what a production system has to guarantee.
What a POC almost never tests:
- Real data, with its gaps, inconsistent formats and in-house vocabulary.
- Integration with existing tools: CRM, ticketing system, reference data, access rights.
- Error handling: what happens when the system gets it wrong, and who takes over?
- Running costs at real volume, not on a hundred test queries.
- The change in day-to-day work for the teams who will live with the tool.
- Compliance: GDPR, and now the AI Act, depending on the use case.
- Measurement: a baseline before, tracking after.
Any one of these can block go-live on its own. Together, they explain why a successful demo tells you almost nothing about what comes next.
The 70 / 30 ratio
Successful AI transformations are 70% an organisational transformation and 30% a technical project.
That conviction shapes my work. When the ratio is inverted, teams spend weeks comparing models and tuning prompts while the real blocker is elsewhere: nobody has decided who owns the system, or how the teams' work changes on Monday morning.
The five blockers I see most often
- No business owner. The POC belongs to the team that built it. Nobody on the operations side is accountable for the outcome or for the trade-offs.
- Success defined by a demo, not by a metric. “It works well” does not survive the first budget meeting. You need a number, a baseline and a target.
- POC data that looks nothing like production. A hand-cleaned sample gives an accuracy the live flow will never match.
- A process that doesn’t move. AI is bolted on top of the existing setup instead of redesigning the flow: who receives what, who approves, who escalates.
- Risk and compliance arrive at the end. Legal discovers the project at go-live, and everything stops for months.
None of these blockers is solved by a better model.
What to decide before writing a line of code
Scoping is the highest-return part of the project. Before building, I ask teams to answer six questions in writing:
- What operational problem are we solving, expressed in volume, lead time or cost?
- Which metric proves the result, and what is its value today?
- Who is the business owner, with the authority to make trade-offs?
- What is the scope, and above all, what will the system not do?
- What happens when it gets it wrong: escalation, human takeover, traceability?
- What running budget is acceptable at real volume?
If you can’t answer, the project isn’t ready. That is not a failure: it is information, obtained before the build budget has been spent.
Getting through: a bounded, measured scope, then widened
The method that works is nothing spectacular. It means shrinking the scope until it is manageable, then widening it on evidence.
- Operational scoping over a few days: the six questions above, the data actually available, the constraints.
- A pilot in real conditions on a limited segment: one request category, one team, one site.
- Instrumentation from day one: every decision the system makes is logged and comparable with the baseline.
- An improvement loop with frontline teams, who flag errors and enrich the domain vocabulary.
- Widening in stages, each stage gated by thresholds set in advance.
It is the logic of a minimum viable product applied to AI: the aim is to learn fast in the field, not to impress in a demo. And it is the logic of a hypothesis roadmap: each stage validates an explicit hypothesis.
What the Home Partners case shows
At Home Partners of America, a Blackstone subsidiary, the engagement covered the contact centre of a residential portfolio of 30,000 units. The AI system qualifies, routes and resolves tenant requests. It was deployed to production across the whole portfolio, within the operational and regulatory constraints of a real estate operator.
Engagement performed at Home Partners of America, a Blackstone subsidiary, 2022 – 2023. Public data, informed stakeholders, within applicable confidentiality agreements.
These three figures do not measure a model. They measure an operational chain: what gets resolved without an agent, what is correctly understood in the language of the business, and how long a request takes to be settled. That is exactly what production demands, and what a POC does not measure. I explain how to instrument these metrics in measuring an AI support assistant.
Frequently asked questions
How long does it take to go from POC to production?
It depends on the scope and the state of the data, but a realistic plan is measured in months, not weeks. A horizon of 3 to 6 months to put the best candidates in a portfolio into production is a healthy order of magnitude, provided ownership and measurement were settled at the outset.
Should we stop doing POCs?
No. Stop doing POCs without exit criteria. A useful POC states from the start the target metric, the business owner and the decision that will follow from the result: industrialise, redirect or stop.
Who should own an AI project in production?
A business leader, backed by the technical team. If the project is owned only by IT or by an innovation team, it will stay a project. The owner is the person who will live with the result and who can change the process around it.
So where do you start?
Depending on your situation, I work in three bounded formats:
- AI Production Audit (15 to 20 days) for organisations sitting on 5 to 15 POCs, few of which reach production: a portfolio diagnostic and a 3 to 6 month production plan.
- AI Act Readiness (30 to 45 days) when AI touches HR, scoring, biometrics or infrastructure: inventory, Annex III classification and an operational compliance plan.
- Operational AI Transformation (60 to 90 days) to transform a high-volume support or ops chain: design and launch of a production system, scoped and measured.
All of them start with a 5 to 7 day operational framing. For SMEs, ClairAI also offers business applications and a guide to budgeting for a business application.