Move an AI Pilot to Production: Why It Stalls
How to move an AI pilot to production: the five reasons pilots stall, a 60-day path out, and how to decide between finishing and restarting.
Founder, Viithiisys

Why can't teams move an AI pilot to production? The five reasons pilots stall
Pilots stall because the demo was funded and the production work was not. The five gaps are a missing owner, missing evals, unscoped integration, an unbooked security review, and a problem with no budget holder.
Public data says the pattern is common. Gartner predicted in July 2024 that at least 30% of generative AI projects would be abandoned after proof of concept by the end of 2025. The MIT NANDA initiative's 2025 report, "The GenAI Divide", found that most enterprise generative AI pilots produced no measurable P&L impact.
Is this ranked by frequency?
We have not published a frequency count of stall causes across the pilots we have been brought in to finish, so we will not pretend to one. The sections below follow the order in which the gaps tend to surface: ownership first, because it blocks everything else, then quality, then integration, risk and money.
What does an AI pilot purgatory look like in practice?
The signs are consistent. The demo is shown every quarter, the sponsor is asked for a date, and nobody on the team can say who gets paged when the output is wrong.
No owner after the demo
A pilot with no named production owner stalls at handover. Someone must hold the pager, the roadmap and the budget line. If the answer is "the innovation team", the pilot has already stopped moving.
Pilots are usually run by a small team borrowed from other work. When the demo lands, those people return to their day jobs, and the receiving team never agreed to take on a system that behaves probabilistically and needs ongoing tuning.
What does an owner actually do?
The owner triages bad outputs, approves prompt or model changes, watches cost per request, and decides when quality has drifted enough to act. That is a product role with engineering support, not a side project.
How to fix it in week one
Name one accountable person on the business side and one on the engineering side. Write down what each will do in the first 90 days after launch. If nobody will sign that, you have learned the pilot does not yet have a real sponsor.
A worked example: a claims-triage pilot returns a wrong category on a Friday evening. With no owner, the message sits in a shared inbox until Monday, and by then the operations lead has gone back to manual sorting. With a named owner, there is a defined path: the engineering contact checks the logs, the business owner decides whether to pause the automated routing, and the incident is added to the eval set so the same failure is caught next time.
No evals, so nobody can prove it works
Without an evaluation set, "it works" is an opinion. Production approval needs a repeatable measure of quality on representative inputs, run before every change. A pilot judged by a few impressive demos cannot pass that bar.
Both major model providers document this practice. Anthropic's guidance on developing test cases and OpenAI's evals guide describe building task-specific test sets and grading outputs against them.
What a minimum viable eval set looks like
Start with 50 to 100 real inputs pulled from the workflow the pilot serves, including the ugly ones: malformed documents, ambiguous requests, edge cases the demo skipped. Label the expected outcome with the person who does this work today.
The trade-off
Hand-labelled sets are slow to build and go stale. Automated grading with a second model is faster but needs spot-checking against human judgement. Most teams use both, and they re-sample production traffic monthly.
One practical rule: agree the pass threshold before the first run, not after seeing the score. If the team decides afterwards that 78% is "good enough", the eval has become a justification exercise. Write the threshold down per category, with a stricter bar for outputs that trigger an action, such as sending an email or updating a record, than for outputs a human reads first.
Integration was never scoped
Pilots run on exported spreadsheets and a sandbox login. Production needs authenticated access to live systems, error handling, retries and audit trails. If that work was never scoped, it appears as a surprise after the demo.
The model call is usually the smallest part of the system. The larger part is reading from the CRM or ERP, writing results back, handling a downstream outage, and deciding what happens when the model returns something unusable.
Where integration estimates go wrong
Teams underestimate three things: legacy systems with no usable API, permission models that differ per user, and data quality problems that the pilot's clean extract hid. Each can add weeks.
A common case is the permission model. The pilot ran under one service account that could see every record. In production, a sales rep should only see their own region's accounts, so the retrieval layer must filter by the requesting user's rights. Retrofitting that after the demo means reworking how documents are indexed, not just adding a login screen.
What to do about it
List every system the pilot touches in production and mark each as API available, API partial, or none. Our AI integration and MLOps work starts from exactly this inventory, because it determines whether the 60-day path is realistic.
Security review nobody booked: what does it block?
An unbooked security review blocks launch regardless of how well the pilot performs. Review queues are often weeks long, and they cannot start until someone submits architecture, data flows and a risk position.
Reviewers will ask where data goes, whether prompts or documents are retained by a third party, how prompt injection is handled, and who can see outputs. The NIST AI Risk Management Framework gives a shared vocabulary for these conversations.
What to prepare before you submit
Prepare a data-flow diagram, a list of vendors that process data, a retention statement, and a description of what the system is allowed to do without a human. Platform controls help: AWS documents content filtering and denied topics in its Bedrock Guardrails guide.
Book it early
Submit the review request in week two, even with a draft. A review that finds problems in week four is cheap. One that finds them in week ten is a missed quarter.
The pilot solved a problem nobody was paid to fix
If no budget holder feels the problem, no one funds the rollout. A pilot can be technically sound and still die because the team whose pain it relieves has no money, and the team with the money has no pain.
A pilot that cannot name who pays for production on day one is a demo, not a pilot.
How to test for a real budget holder
Ask who will lose a target, a headcount plan or a customer if this does not ship. Ask which line in next year's plan changes if it works. Vague answers mean the pilot is a technology exercise.
What to do if the answer is no one
Re-scope to a workflow where the cost is already measured: hours spent on a manual review, error rates on a document process, ticket handling time. Our broken workflow assessment exists for finding that workflow, and we cover it again at the end.
What does the 60-day path to move an AI pilot to production look like?
The path is four phases of roughly two weeks each, with ownership, evals and the security submission front-loaded. Integration and hardening follow, and launch is staged behind a limited rollout.
Our engineering is based in Mohali, with an office in Markham, Ontario, and we have shipped 500+ projects since 2007 for clients including Paytm, Snapdeal, IKEA, Nestle, Shiprocket and Vikram Solar. The sequencing below is how we plan recovery work, not a measured benchmark.
Weeks 1 to 2: ownership, evals, review request
Name the owners, build the first 50 to 100 labelled cases, inventory systems, and submit the security review with a draft architecture.
Weeks 3 to 4: baseline and scope
Score the pilot against the eval set, decide what to keep, and fix the production scope in writing. Fixed scope matters here, which is the same discipline behind our Moonship MVP approach of a working build in 30 days.
Weeks 5 to 6: integration and guardrails
Build the live connections, add retries, logging and human-approval steps, and respond to review findings.
Weeks 7 to 8: staged launch
Release to a small group, monitor quality and cost against the eval set, and widen access only when the numbers hold. Our AI development teams in the US, UK and Canada run this sequence with client engineers rather than around them.
What does it cost to finish versus restart an AI pilot?
Finishing is usually cheaper when the approach tested well and the gaps are operational. Restarting is justified when the foundation fails: unusable data, no measurable success definition, or a vendor lock-in the business will not accept. We do not quote prices here: scope depends on the integration depth and is set after a discovery call.
| Factor | Finish the pilot | Restart |
|---|---|---|
| Model approach scores well on evals | Keep it | Wasted if discarded |
| Data source usable in production | Keep it | Rebuild only if not |
| Integration scoped | Add the missing work | Same work, plus re-learning |
| Sponsor and budget holder exist | Proceed | Not a reason to restart |
| No success definition | Define it first | Define it first |
| Time to a live system | Shorter | Longer: repeats discovery |
The honest limit
Finishing carries inherited debt: prompts nobody documented, shortcuts taken for the demo. If the eval baseline shows the approach does not work, restarting is the correct call, and saying so early is cheaper than shipping a weak system.
Next step
If your pilot is stalled and the sponsor wants a date, book a pilot recovery call through the broken workflow assessment. We will tell you which of the five gaps you have and whether finishing is realistic. For a general question first, use the contact form.
FAQ
- How long does it take to move an AI pilot to production?
- It depends on integration depth and the security review, but a focused path runs about 60 days when an owner, an evaluation set and a review slot are secured in the first two weeks. Pilots with unscoped integrations or an unstarted security review take longer, and the review is usually the long pole.
- Should we finish a failed AI pilot or rebuild it from scratch?
- Finish it if the model approach holds up against a real evaluation set and the problem has a budget holder. Restart if the pilot was built on a data source you cannot use in production, or if nobody can say what success looks like. Check the evidence before deciding either way.
- Why do so many AI pilots never reach production?
- Pilots are built to demonstrate that something is possible, and production requires an owner, measurable quality, integration with live systems, an approved risk position and a budget line. Most pilots are funded and staffed for the first goal only, so the second set of work has no one assigned.
- What is AI pilot purgatory?
- AI pilot purgatory is the state where a pilot works well enough to keep demonstrating but never gets the owner, integration, security approval or budget needed to go live. It rarely gets cancelled, so it consumes attention quarter after quarter without delivering a production system.